Back to blog

LLMs Reflect Their Creators' Ideology. What the New Study Reveals

Nineteen language models were asked to describe nearly 4,000 political figures. The result wasn't just different, it was predictably different: every LLM thinks through its creators' lens. This raises a question not about bias, but about who controls our view of the world.

Privacy & Governance AIGovernancePrivacy

AI data analytics and segmentation, conceptual neon illustration of data flows

A recent study published in npj Artificial Intelligence by a team led by Maarten Buyl from Ghent University ran a simple and elegant experiment. The researchers took 19 popular language models, from GPT-4o to Gemini and Grok, and asked each to describe 3,991 politically significant figures. No leading questions, no tests for liberalism or conservatism: just describe the person. Then, in a separate step, they asked the same model to evaluate its own description: did it come out positive or negative? The trick was that the model did not know why it was being asked to describe people, and it was not trying to appear impartial. It was simply doing what it does best: generating text.

The results were sobering. Google’s Gemini is systematically warmer toward figures associated with progressive values, civil liberties, and multiculturalism. xAI’s Grok leans toward those tied to national sovereignty, centralized authority, and economic protectionism. The Chinese models diverged: Alibaba’s Qwen aligns with international standards and sits closer to the Western spectrum, while Baidu’s Wenxiaoyan is oriented toward the domestic market with a corresponding ideological profile. Arabic models, Russian models, European models, each with its own pattern. And most interestingly, even the language of the prompt changes the outcome. The same model queried in English and in Russian describes the same people with a different tone. This means ideology sits not only in the data and weights but in the linguistic space the model operates in.

The authors avoid the word ‘bias’ and they do so deliberately. They insist: neutrality is a philosophically problematic concept. You cannot be objective about questions where the very definition of objectivity is a cultural construct. They reference Chantal Mouffe and her idea of agonistic pluralism: democracy is not consensus but a space where different viewpoints compete. The researchers suggest viewing model differences not as a bug but as a feature. If we end up with a single ‘neutral’ model, we will not get objectivity. We will get a monopoly on perspective.

This framework strikes me as both accurate and troubling. Accurate, because it explains what many of us feel intuitively. Anyone who has worked with different LLMs has noticed: they have different characters. One leans toward compromise, another toward provocation, a third toward diplomatic evasion. We tend to attribute this to temperature, system prompts, or fine-tuning, but the study shows it runs deeper: the model absorbs ideology at the data selection stage, long before alignment. Troubling, because concentrating AI power in the hands of a few companies means concentrating ideological perspectives. When millions of people receive information through the same interface every day, the question is not about the quality of answers but about the homogenization of thought.

I work with different models myself: Claude for reflection and long-form text, ChatGPT for analysis, Gemini for quick drafts. And I have long noticed that I choose a model not just by code quality or speed but by how it thinks. This is not a metaphor. Each model genuinely has its own angle, its own sensitivity to context, its own sense of what matters and what can be omitted. Buyl’s study simply confirms instrumentally what practitioners already know at the level of instinct.

Several practical consequences follow from this. Choosing a model is not just a technical decision: latency, cost per token, and code quality. It is a decision about whose lens you place yourself and your users inside. For some tasks this does not matter: summarizing a research paper can be done with any model, the difference will be minimal. For others it is critical: political analysis, news generation, educational content, legal advice. Moreover, we will have to learn to live with the fact that the perfect answer does not exist. There are only answers, each carrying the imprint of its creators, and the skill of the future is not finding the one true answer but consciously choosing between different perspectives depending on the task.

Buyl’s group does not offer ready-made recipes, but it shifts the conversation from ‘which model is smarter’ to ‘which model for what purpose.’ And that is already a fundamentally different discussion.

More thinking