International marketing

How AI Search Engines Cite Sources by Language and Country

How AI Search Engines Cite Sources by Language and Country
Rayne Aguilar
Written by
Rayne Aguilar
Elizabeth Pokorny
Reviewed by
Elizabeth Pokorny
Updated on
July 28, 2026

When someone asks an AI assistant a question, the sources it points to determines what that person learns and which brands they run into. And we realized…how do these assistants choose their sources? Does the behavior shift across languages, countries, topics, and AI systems?

Like with all the other questions we had, we decided to investigate.

We came up with 400 questions spanning 5 topic categories and 3 levels of specificity, wrote each one in English, French, Spanish, and Japanese, and put them to 4 AI systems across 5 countries. That gave us more than 16,000 answers, and we logged every source each system cited.

Several patterns came through clearly. Some are about how these systems work in general, but the one that matters most to any brand entering new markets is, you guessed it: about language.

Let’s get into it.

How We Ran the Study

We wanted the query set to reflect real curiosity rather than your standard SEO keyword list, so we spread the 400 questions across 5 categories (ecommerce, SaaS, general knowledge, global news, and tutorials) and 3 query lengths, from short and general to long and specific.

We then ran each question in 2 versions per market: in English everywhere, and in the local language of each market (French in France, Spanish in Spain, Argentina, and the US, and Japanese in Japan). The 4 systems were GPT 5.5, Claude Haiku 4.5, Gemini Flash 3.5, and Google AI Overviews, except for France, as it still hasn’t rolled out yet.

For every answer, we recorded the source URLs the system returned. The numbers below come from that raw set of over 16,000 responses.

Chatbots and AI Overviews Behave Like Opposite Machines

The biggest divergence in behavior was between the chatbots and Google’s AI Overviews, most notably in how each of them cites their sources.

The chatbots barely reach for sources on short, general questions. But the more the question gets specific, the more they cite. Claude, for example, averages 0.12 sources on a short query and 2.96 on a long one; ChatGPT moves from 0.09 to 2.26. Gemini cites more heavily throughout, climbing from roughly 1.7 sources on short queries to about 6.5 on long ones. If you’re gunning for AI search visibility (which, if you’re reading this, probably means that you are), optimizing your content for specific, long-tail prompts likely increases your chances of getting cited.

However, Google AI Overviews do the reverse. They cite the most on short, general queries (around 6 sources each) and taper off as the query gets longer and more specific (down to roughly 4). That is the behavior of a search results page, not a chatbot: short queries get a spread of links, and the more precise the question, the more the answer narrows.

Here is how the three chatbots compare, measured as the average number of sources cited per query:

ChatbotShort queryMedium queryLong query
ChatGPT0.091.972.26
Claude0.122.122.96
Gemini1.737.146.55

Chatbots answer broad questions from what they already know and only pull in sources when a query gets specific, while AI Overviews attach links the way a search engine does.

Brands Get Named Even When No Link Appears

In many AI interfaces, the reader sees the written answer, not a list of links. So the companies an assistant names inside its answer can matter as much as the URLs it cites. We counted how many distinct companies each chatbot named per query. (Google AI Overviews are left out here, since we could capture their cited links but not their answer text.)

Average Companies Named per Query, by Chatbot and Query Length

ChatbotShort queryMedium queryLong query
ChatGPT2.084.203.00
Claude0.952.131.86
Gemini4.284.383.43

ChatGPT and Gemini name several companies per query even on short, general questions, the same questions where they cite almost no URLs. A brand can turn up in the answer text without a single link pointing to it, which means being the kind of company these models already recognize matters, not only being linkable.

Citations Climb on Commercial and Time-Sensitive Topics

Topic matters as much as query length. Looking at the chatbots, average sources cited per question rose steadily from evergreen subjects to commercial and fast-moving ones.

CategoryAvg. sources cited per query
Tutorials0.94
General knowledge1.62
SaaS3.42
Ecommerce4.09
Global news5.02

Fixed, well-documented topics like tutorials and general knowledge pull the fewest citations; the models answer those largely from training. Commercial and time-sensitive topics like ecommerce, SaaS, and global news pull the most, because the models go looking for current information. For anyone in a commercial category, that is the encouraging half of the finding: these are exactly the queries where being a citable source is up for grabs.

English Still Pulls the Most Sources

Across the three chatbots, English queries pulled more sources than local-language queries in every single market we tested.

Graph showing the average sources cited per query across the three chatbots in Argentina, Spain, France, Japan, and the United States

Ask the same question in English, and these systems reach for more sources than they do in the local language, everywhere. English content still has a structural head start in how much gets cited at all, which makes sense, given that it remains the dominant language on the Internet. What changes with language is not whether sources appear, but which ones.

Ask in a Local Language, and the Sources Become Local

When we took the same questions and asked them in a local language instead of English, the systems shifted noticeably toward local, in-language sources.

In Japan, 26% of all cited sources were on Japanese (.jp) domains when we asked in Japanese, compared with about 1% when we asked the same questions in English. Narrow that to Japanese government and academic sites (.go.jp and .ac.jp) and the gap holds: 4.26% in Japanese versus 0.13% in English.

France showed the same shape. French-language queries sent 16% of citations to .fr domains, against 1.4% for English queries, and French government sites (.gouv.fr) went from near-zero in English to 1.25% in French. Spanish-language queries in Spain sent 7% of citations to .es domains versus about 1% in English, and Argentina followed the same direction more mildly.

Graph showing the local domain citation share comparing local-language vs English series

Same questions, same topics, different query language, and the systems reach for a different set of sources. In a nutshell, content living only in English sits largely outside the pool these systems draw on when they answer in a local language. The effect is also largest in the market furthest from English, Japan, and smallest in the closest ones.

Each System Has Its Own Source Criteria

The systems also differ in the kinds of sources they favor, and that held steady across every language we tested.

Graph showing where each AI system sends its citations

As illustrated in the table, each system extracts its citations from commercial and brand pages, but this behavior diverges when it comes to everything else.

For example, Claude is clearly the biggest fan of commercial and brand pages, with nearly 9 in 10 citations going to these and almost none to media. ChatGPT spreads its weight the widest (but not by much), still largely favoring commercial and brand, with global institutional and media trailing in at 8.1% and 9.9%, respectively.

Gemini and Google AI Overviews lean heavily on community and video sources: YouTube and Reddit alone accounted for 21.5% of Gemini's citations and 29.8% of AI Overviews' citations.

If you are trying to be cited by Gemini or in an AI Overview, a strong YouTube and Reddit presence will make the biggest difference. If ChatGPT and Claude are your targets, your own brand pages and coverage on established publications matter more.

Which Domains and Brands Get Cited Most

Averages describe the shape of the citations. The specific names show who is filling that space. We pulled the domains and companies each chatbot cited most within every topic category, after excluding any company from queries that named it directly, so these reflect what the models reached for on their own. These two tables cover the three chatbots.

Top-cited domains (top 3 per category)

CategoryChatGPTClaudeGemini
Ecommercetomsguide.com (6.6%)
techradar.com (6.1%)
tomshardware.com (3.6%)
alibaba.com (4.3%)
amazon.com (2.5%)
techradar.com (1.4%)
youtube.com (24.7%)
reddit.com (13.8%)
rtings.com (1.6%)
General knowledgeentreprendre.gouv.fr (4.3%)
science.nasa.gov (4.3%)
wwwnc.cdc.gov (4.3%)
arxiv.org (5.0%)
ncbi.nlm.nih.gov (4.1%)
en.wikipedia.org (3.3%)
youtube.com (14.3%)
reddit.com (6.5%)
en.wikipedia.org (3.1%)
Global newsiea.org (6.5%)
digitalstrategy.ec.europa.eu (3.4%)
axios.com (2.5%)
en.wikipedia.org (2.5%)
sciencedirect.com (2.3%)
ncbi.nlm.nih.gov (1.6%)
youtube.com (7.8%)
reddit.com (1.6%)
en.wikipedia.org (1.3%)
SaaStechradar.com (13.1%)
en.wikipedia.org (3.0%)
docs.n8n.io (2.9%)
lovable.dev (1.3%)
sybill.ai (1.0%)
featurebase.app (0.9%)
youtube.com (16.2%)
reddit.com (5.6%)
zapier.com (2.0%)
Tutorialsgraph.facebook.com (12.6%)
developer.mozilla.org (5.1%)
github.com (4.8%)
adstellar.ai (3.6%)
youtube.com (20.3%)
reddit.com (7.2%)
github.com (3.1%)

Here, we can see a clear split.

Gemini concentrates its citations on a small set of platforms, with YouTube and Reddit taking the top spots in almost every category.

ChatGPT and Claude spread theirs across many category-specific authorities instead: hardware review sites like Tom's Guide and TechRadar for ecommerce, and research and reference sources like arXiv, PubMed, and Wikipedia for general knowledge and news.

Ecommerce and SaaS carry the most citations and give the steadiest read; the general-knowledge counts are thin, so treat those positions as directional.

Top-cited companies (top 3 per category)

CategoryChatGPTClaudeGemini
EcommerceLogitech (2.2%)
Razer (1.7%)
Anker (1.6%)
Nike (2.6%)
Logitech (2.2%)
Amazon (2.2%)
Logitech (1.9%)
Razer (1.7%)
Anker (1.7%)
General knowledgeGoogle (2.6%)
Amundi (1.9%)
Google (12.7%)
IBM (4.2%)
Société de crédit agricole (4.2%)
Google (2.8%)
LIGO (1.5%)
NASA (1.4%)
Global newsTesla (1.2%)
McKinsey (1.1%)
BYD (1.0%)
NASA (3.1%)
SpaceX (2.6%)
ESA (2.2%)
Google (1.6%)
Tesla (1.5%)
Microsoft (1.3%)
SaaSHubSpot (1.7%)
Salesforce (1.1%)
Mailchimp (1.1%)
HubSpot (2.7%)
Salesforce (1.7%)
Monday.com (1.5%)
HubSpot (1.6%)
Salesforce (1.4%)
Notion (1.1%)
TutorialsPostgreSQL (3.0%)
MySQL (2.1%)
AWS (1.9%)
Nginx (1.9%)
Express (1.7%)
Traefik (1.5%)
Flask (1.8%)
AWS (1.7%)
PostgreSQL (1.7%)

HubSpot tops SaaS for all three chatbots with Salesforce close behind, a rare level of agreement across models, and ecommerce clusters around consumer hardware brands like Logitech, Razer, and Anker.

Still, this is a modest lead: HubSpot sits under 3% in SaaS, so mentions are spread across a long tail of companies. Being cited is less about unseating one dominant name like you’d want in the traditional top 3 blue links in SEO. Instead, it’s about being in the mix at all.

What This Means for Going Global

The study measures where these systems look, not what any single site should do, so it is worth being precise about the implication. What the data shows is that AI systems localize their sourcing: ask in a language, and they favor sources in that language, from that place.

For a brand expanding abroad, that reframes local-language content as a question of eligibility. Simply put, content in a market’s native language can be cited when a local asks a question. But if it exists only in English, it mostly cannot be cited. It is a pattern we have seen from a different angle in our earlier research on translation and AI visibility, and this study helps explain the mechanism behind it.

Ready to Reach These Markets in Their Own Language?

If you want your content eligible to be cited when people search in their own language, you can start your free trial of Weglot and have a multilingual site running in minutes.

direction icon
Discover Weglot

Good things come to those who wait. International traffic doesn’t.

We’ll get your first languages live. You decide how far you want to go. Try Weglot for free today.

In this article, we're going to look into:
Rocket icon

Ready to get started?

The best way to understand the power of Weglot is to see it for yourself. Test it for free and without any engagement.

A demo website is available in your dashboard if you’re not ready to connect your website yet.

Read articles you may also like

FAQ icon

Common questions

Do AI search engines cite different sources in different languages?

arrow

Yes. In our study, asking the same question in a local language rather than English shifted a large share of citations toward local, in-language domains, most dramatically in Japanese.

Which AI system cites the most sources?

arrow

Among the chatbots, Gemini cited the most overall. Google AI Overviews cited heavily on short, general queries and less on specific ones, the opposite of how the chatbots behave.

Does the topic affect how many sources an AI cites?

arrow

Yes. Evergreen topics like tutorials and general knowledge drew the fewest citations, while commercial and time-sensitive topics like ecommerce, SaaS, and global news drew the most.

Do Gemini and ChatGPT rely on the same kinds of sources?

arrow

No. Gemini and Google AI Overviews leaned heavily on YouTube and Reddit, while ChatGPT and Claude drew far more on brand, editorial, and institutional pages.

Blue arrow

Blue arrow

Blue arrow