Two language streams feeding different AI engines with different citation outcomes

How Query Language and Engine Choice Reshape AI Citations: Evidence From Two GEO Cases

AI citation visibility is not one market-wide score. It changes with the engine, the query language and the source ecosystem available to retrieval. Geolix.ai production monitoring now adds a Chinese-market view to the global evidence: in one controlled Chinese-language case, the same brand and the same ten buyer-intent questions produced mention rates ranging from 10.3% to 81.9% across six engines. The practical lesson for APAC fintech is not that Chinese content always beats English content. It is that English-only measurement cannot reveal how a brand performs in Chinese-language discovery.

Key findings

  • Profound analyzed 3.25 billion citations across seven global engines and 14 countries, showing that query language can materially reshape citation patterns; its study did not include mainland China or Chinese answer engines.
  • In Geolix.ai's low-code SaaS case, 1,233 valid answers to ten Chinese buyer-intent questions produced 15,353 citation records. Brand mention rates ranged from 10.3% on Perplexity to 81.9% on DeepSeek.
  • Chinese-engine and Western-engine source pools overlapped only modestly in that case: 68 shared domains and a 12% Jaccard overlap.
  • Chinese-language sources represented 90.4% to 100% of citations across all six engines in the low-code SaaS case, including ChatGPT, Perplexity and Google AI Mode.
  • These two Geolix.ai cases are complementary evidence, not a matched Chinese-versus-English experiment: they cover different categories, prompts, dates, repetition levels and engine sets.

What global research establishes

Profound’s March 2026 study covered ChatGPT, Claude, Google AI Mode, Google AI Overviews, Gemini, Microsoft Copilot and Perplexity. Social sources accounted for 15.3% of citations in Google AI Overviews and 14.5% in AI Mode, compared with 9.1% in ChatGPT, 3.99% in Claude and 3.6% in Gemini. The useful conclusion is not that one source type always wins; citation behavior varies by engine and market language.

Sources: Profound citation study | Peec AI language study

Peec AI provides a second, independently collected vendor dataset. Across more than ten million prompts and twenty million fan-out searches, it reported that 78% of non-English sessions included English supplementation and that 43% of fan-out searches for non-English prompts were conducted on the English-language web. Peer-reviewed multilingual retrieval research also documents language preference and high-resource-language bias. Together, these studies explain why translation alone cannot guarantee equivalent retrieval.

Sources: ACL 2025 multilingual RAG research | ACL 2026 language-bias research

The low-code SaaS case: the same Chinese questions, six different visibility outcomes

From June 15 to 19, 2026, Geolix.ai tested one low-code SaaS brand using ten Chinese purchase-intent questions across ChatGPT, DeepSeek, Doubao, Google AI Mode, Perplexity and Qwen. Each question was scheduled five times per day per engine. After invalid runs were removed, the dataset contained 1,233 valid answers and 15,353 citation records.

EngineValid answersMention rateTop-1 rateChinese-source share
DeepSeek21081.9%5.2%96.5%
Doubao19971.4%4.5%100.0%
Qwen20670.9%34.5%100.0%
ChatGPT20745.4%3.4%96.3%
Google AI Mode20830.8%n/a90.4%
Perplexity20310.3%1.0%97.1%

The gap is operationally large: mention rate differed by roughly eight times, while measurable Top-1 rate ranged from 1.0% to 34.5%. Google AI Mode’s rank field failed to populate, so its Top-1 result is reported as unavailable rather than zero. At question level, the same content gap could reverse across engines. For example, two questions received zero ChatGPT mentions while DeepSeek recorded 76% and 71%; Qwen recorded 10% and 100%. A single blended GEO score would hide those differences.

Source ecosystems diverged even when the prompt language stayed fixed

The six engines did not simply produce different rankings from one common source pool. DeepSeek drew 62.5% of its citations from cloud and official developer ecosystems. Doubao drew 31.1% from user-generated and video sources, especially Douyin. Qwen drew 72.6% from user-generated and video sources and 25.1% from search self-references. By contrast, ChatGPT’s largest category was vendor websites at 68.6%. These are case-specific observations, not universal rules for every category.

When Geolix.ai grouped DeepSeek, Doubao and Qwen as Chinese engines and ChatGPT, Google AI Mode and Perplexity as Western engines, the two groups shared only 68 citation domains. Their unique-domain Jaccard overlap was 12%. Yet more than 90% of citations in every engine were Chinese-language sources. This is strong evidence that Chinese questions activate a predominantly Chinese evidence environment, while engine architecture still determines which part of that environment is used.

What the English fintech case adds

A separate Geolix.ai case monitored nine English purchase-intent questions for a GEO service brand serving fintech in Singapore and APAC. Between July 20 and 31, 2026, ChatGPT API, Gemini and Perplexity produced 15,495 valid answers and 223,196 citation records. Perplexity mentioned the brand in 75.3% of answers, Gemini in 38.6% and ChatGPT in 6.9%. Most citations were non-Chinese sources: 89.0%, 99.9% and 91.5%, respectively.

The case also shows that citation and brand visibility are different metrics. ChatGPT cited the target brand’s own domain in 22.9% of answers but named the brand in only 6.9%—a 3.3-to-1 gap. Perplexity, meanwhile, cited Chinese-language versions of the brand’s pages 3,149 times, representing 42.5% of its citations to paired pages; ChatGPT and Gemini did not cite those Chinese versions. Even under English prompts, engines can treat localized pages differently.

What APAC fintech teams should do

  • Maintain separate English and Chinese buyer-question libraries based on real discovery, comparison, trust, licensing and market-access intents.
  • Report mention rate, Top-1 and Top-3 recommendation rate, own-domain citation rate and third-party citation rate separately by engine.
  • Map each engine’s cited domains and source types before deciding where to publish or conduct outreach.
  • Keep regulated facts—legal entity, licence, market availability, fees, eligibility and product limits—consistent across language versions and third-party profiles.
  • Treat before-and-after changes as observational unless the test holds prompts, engines, timing and other content actions constant.
  • Add an independent factual-accuracy review; the current Geolix.ai export measures mentions, rankings and citations but does not contain a factual-accuracy field.

Methodology and limitations

Both cases are production-monitoring exports, not randomized experiments. The two cases differ in category, brand, question set, date range, repetition level and engine availability, so their aggregate percentages should not be used to calculate a causal language effect. In the low-code SaaS case, DeepSeek and Doubao region values were not stored. In the English fintech case, Google AI Mode succeeded on only 1 of 450 scheduled runs and Google AI Overviews on 62 of 450; both were excluded. The dataset contains no factual-accuracy field, and negative-sentiment counts were uniformly zero, so neither metric is used as evidence.

Frequently asked questions

Does this prove that Chinese content always performs better for Chinese prompts?

No. It shows that Chinese sources dominated one controlled Chinese-language case and that engines used materially different subsets of those sources. A strict language-effect estimate would require the same brand and same prompts tested in both languages during the same period.

Can ChatGPT performance predict DeepSeek or Qwen performance?

Not safely. In the low-code SaaS case, the same brand and prompts produced large differences in mentions, recommendations and source types across engines.

What is the minimum useful multilingual GEO dashboard?

Prompt-level answers and cited URLs, segmented by engine, language, location and date, with mention, recommendation, own-domain citation and factual-accuracy metrics kept separate.

References

查询语言与引擎选择如何重塑 AI 引用:来自两个 GEO 实测案例的证据

AI 引用可见度不是一个可以覆盖全市场的总分。它会随引擎、查询语言和可供检索的信源生态变化。Geolix.ai 生产监测数据补上了全球研究缺少的中文视角:在一个中文市场案例中,同一品牌、同一组 10 道购买意图题,在 6 个引擎中的提及率从 10.3% 到 81.9% 不等。对亚太金融科技而言,结论不是“中文内容一定优于英文内容”,而是只测英文无法判断品牌在中文 AI 发现链路中的真实表现。

核心结论

  • Profound 分析了 7 个全球引擎、14 个国家的 32.5 亿次引用,发现查询语言会显著改变引用模式;但研究没有覆盖中国大陆和中文答案引擎。
  • Geolix.ai 的低代码 SaaS 案例包含 1,233 条中文有效回答和 15,353 条引用记录;同一品牌的提及率从 Perplexity 的 10.3% 到 DeepSeek 的 81.9%。
  • 该案例中,中文引擎组与西方引擎组仅共享 68 个域名,引用域名 Jaccard 重合度为 12%。
  • 低代码 SaaS 案例覆盖的 6 个引擎中,中文来源占比均在 90.4%—100%,包括 ChatGPT、Perplexity 和 Google AI Mode。
  • 两个 Geolix.ai 案例只能作为互补证据,不能视为严格的中英文同题实验,因为品类、品牌、题目、日期、重复次数和引擎范围均不同。

全球研究已经证明了什么

Profound 的 2026 年 3 月研究覆盖 ChatGPT、Claude、Google AI Mode、Google AI Overviews、Gemini、Microsoft Copilot 和 Perplexity。社交来源在 Google AI Overviews 引用中占 15.3%,在 AI Mode 中占 14.5%,而 ChatGPT 为 9.1%、Claude 为 3.99%、Gemini 为 3.6%。正确解读不是“某一类信源永远有效”,而是引擎和市场语言会改变引用行为。

Sources: Profound 引用研究 | Peec AI 语言研究

Peec AI 提供了另一组独立采集的厂商数据:在超过 1,000 万个问题和 2,000 万次 fan-out 检索中,78% 的非英文会话补充使用了英文检索,非英文问题触发的 fan-out 检索有 43% 发生在英文网络。多语言检索领域的同行评审研究也发现语言偏好和高资源语言偏差。这些证据共同说明,翻译页面并不能保证形成等价的检索结果。

Sources: ACL 2025 多语言 RAG 研究 | ACL 2026 语言偏差研究

低代码 SaaS 案例:同一批中文问题,六种可见度结果

2026 年 6 月 15—19 日,Geolix.ai 对一个低代码 SaaS 品牌进行测试:10 道中文购买意图题,覆盖 ChatGPT、DeepSeek、豆包、Google AI Mode、Perplexity 和通义千问;每题、每天、每引擎计划重复 5 次。剔除无效运行后,共得到 1,233 条有效回答和 15,353 条引用记录。

引擎有效回答提及率Top-1 率中文来源占比
DeepSeek21081.9%5.2%96.5%
豆包19971.4%4.5%100.0%
通义千问20670.9%34.5%100.0%
ChatGPT20745.4%3.4%96.3%
Google AI Mode20830.8%n/a90.4%
Perplexity20310.3%1.0%97.1%

差距具有实际运营意义:提及率约相差 8 倍,可计算的 Top-1 推荐率从 1.0% 到 34.5%。Google AI Mode 的排名字段抽取失败,因此标为 n/a,而不是 0。逐题看,同一内容缺口还会在不同引擎中反转:有两道题在 ChatGPT 的提及率为 0,但 DeepSeek 分别为 76% 和 71%,通义千问分别为 10% 和 100%。用一个平均 GEO 分数会把这些差异全部抹平。

即使题目语言不变,引擎使用的信源生态也不同

DeepSeek 有 62.5% 的引用来自云平台和官方开发者生态;豆包有 31.1% 来自 UGC/视频来源,尤其是抖音;通义千问有 72.6% 来自 UGC/视频,另有 25.1% 为搜索自引。ChatGPT 最大的来源类别则是厂商官网,占 68.6%。这些是本案例的观察结果,不应扩写成适用于所有行业的普遍规律。

将 DeepSeek、豆包、通义千问归为中文引擎组,将 ChatGPT、Google AI Mode、Perplexity 归为西方引擎组后,两组只共享 68 个引用域名,唯一域名 Jaccard 重合度为 12%。与此同时,6 个引擎超过 90% 的引用都是中文来源。这说明中文问题会激活以中文材料为主的证据环境,但不同引擎仍会从中选择完全不同的部分。

英文金融科技案例补充了什么

另一组 Geolix.ai 数据监测了一个面向新加坡/亚太金融科技的 GEO 服务品牌。2026 年 7 月 20—31 日,9 道英文购买意图题在 ChatGPT API、Gemini 和 Perplexity 中产生 15,495 条有效回答和 223,196 条引用记录。Perplexity 的品牌提及率为 75.3%,Gemini 为 38.6%,ChatGPT 为 6.9%;三者非中文来源占比分别为 89.0%、99.9% 和 91.5%。

该案例还证明“被引用”和“被提及”不是同一指标。ChatGPT 在 22.9% 的回答中引用目标品牌官网,却只在 6.9% 的回答中写出品牌名,二者相差 3.3 倍。Perplexity 还引用了 3,149 次中文版本页面,占其同源成对页面引用的 42.5%;ChatGPT 和 Gemini 对这些中文版本的引用均为 0。即使问题是英文,不同引擎对本地化页面的处理也可能不同。

亚太金融科技团队应该怎么做

  • 分别建立英文和中文的发现、比较、信任、牌照与市场准入问题库。
  • 按引擎分别报告提及率、Top-1/Top-3 推荐率、自有域名被引率和第三方被引率。
  • 在决定发布渠道或媒体合作前,先梳理各引擎实际引用的域名与来源类型。
  • 统一不同语言版本及第三方资料中的法律实体、牌照、市场范围、费用、资格和产品限制。
  • 除非控制题目、引擎、时间和其他内容动作,否则前后变化只能写成观察性结果。
  • 另行增加事实准确度审查;当前 Geolix.ai 导出表没有事实准确率字段。

方法与局限

两个案例均来自生产监测,不是随机实验,且在品类、品牌、问题、日期、重复次数和可用引擎上各不相同,因此不能用汇总比例计算语言的因果影响。低代码 SaaS 案例的 DeepSeek/豆包地区字段未落库;英文金融科技案例的 Google AI Mode 计划 450 次仅成功 1 次,Google AI Overviews 仅成功 62 次,均已剔除。数据没有事实准确率字段,负面情绪计数又全部为 0,因此本文不使用这两个指标作为证据。

常见问题

这些数据是否证明中文问题一定更适合中文内容?

不能。它证明一个中文同题案例高度依赖中文来源,而且不同引擎使用的中文来源子集差异很大。若要估算语言本身的影响,仍需对同一品牌、同一问题在同一时间分别使用中英文测试。

ChatGPT 表现能否预测 DeepSeek 或通义千问?

不能安全推断。在低代码 SaaS 案例中,同一品牌和问题在提及、推荐和来源类型上都出现明显跨引擎差异。

最低限度的多语言 GEO 看板应包含什么?

需要保留问题级答案与引用网址,并按引擎、语言、地区、日期拆分;提及、推荐、自有域名引用和事实准确度必须分别计算。

参考资料