How do you measure what AI assistants really recommend? We run the experiments — pre-specified metrics, confidence intervals, thousands of archived answers — and publish what we find, including the parts that are inconvenient. GenSig's own methodology is built on these results.
We asked the same purchase-intent questions 2,500 times across five API configurations and 30 times by hand in the real consumer apps (ChatGPT web, Gemini web). The bare LLM API missed 3 of the 5 brands real sessions consistently recommend; temperature 0 made answers highly repeatable — of the wrong brand set; and each search-grounded API turned out to mirror its own consumer surface.
Every study ships with its raw data and code — download the Research #1 dataset (2,500 API responses + 30 consumer-app runs, prompts, brand dictionary and analysis code).
More studies are on the way — replications across categories and markets. Questions, corrections and replication requests: [email protected].