Research

How do you measure what AI assistants really recommend? We run the experiments — pre-specified metrics, confidence intervals, thousands of archived answers — and publish what we find, including the parts that are inconvenient. GenSig's own methodology is built on these results.

GenSig Research #1 · July 2026

How stable are AI brand recommendations?
A reproducibility & fidelity study

We asked the same purchase-intent questions 2,500 times across five API configurations and 30 times by hand in the real consumer apps (ChatGPT web, Gemini web). The bare LLM API missed 3 of the 5 brands real sessions consistently recommend; temperature 0 made answers highly repeatable — of the wrong brand set; and each search-grounded API turned out to mirror its own consumer surface.

ρ = 0.23bare API vs what ChatGPT web actually recommends — the naive measurement is nearly uncorrelated with reality
<1%of 500 bare-API runs mention the browser's #3 brand — model memory lags the market
J = 1.0OpenAI's search API reproduced ChatGPT web's top-5 exactly — measure each surface through its own live channel
Read the study →

Every study ships with its raw data and code — download the Research #1 dataset (2,500 API responses + 30 consumer-app runs, prompts, brand dictionary and analysis code).

More studies are on the way — replications across categories and markets. Questions, corrections and replication requests: [email protected].