For AI labs, evaluation teams, and GEO platforms

Benchmarks ask made-up questions. We didn't make ours up.

Verso sees the same prompt answered by two models, and it sees which answer the person kept. Not annotators. Not synthetic prompt sets. People doing their actual work, choosing. That preference — paired, organic, and continuous — is the dataset.

Four data products

Everything we sell is derived from the same paired-comparison pipeline. The raw text of conversations never leaves.

Preference rates

Model A wins X% of the time against Model B, on a given type of question. The paired design — same prompt, two models, one choice — controls for prompt variance. A modest sample still resolves a real difference because the pairing does the work.

Brand visibility

Named brands and products as they appear in model responses — not what users ask about, but what models recommend. When someone asks "which laptop should I buy," we record which brands each model names, in what order, and whether that changes across model versions.

Query taxonomy

How people actually phrase questions to AI models — the formulations, the follow-ups, the distribution across verticals. Categorized, not quoted. The taxonomy is a structural map of real demand, not a conversation log.

Longitudinal drift

A fixed panel of prompts, replayed as model versions change. If Model A used to recommend Brand X and now recommends Brand Y, that shows up here. The signal is what moved, and when.

How the data is collected

Verso is a free Chrome extension that sits next to ChatGPT, Claude, Gemini, and Perplexity. When a user asks a question, Verso sends the same prompt to a second model and displays both answers side by side. The user picks one — or calls it a draw. That preference, the two responses, and the prompt category are stored. The methodology is a paired comparison with organic prompts and revealed preference.

Read the full methodology

Coverage

5 verticals, 14 models compared, 14,320 distinct prompts since 2025-02. Cells with sufficient depth support full preference rates; thinner cells support model-vs-model deltas only.

See the full coverage table

Sponsored panels

If your vertical is thin in the coverage table, we don't pretend otherwise. We recruit for it. Consenting users, your category, your timeline — funded by the engagement. You get depth where you need it instead of breadth you can't use.

What we will not sell you

Raw conversations. Not as a policy position — as a design constraint. The data is personal, and a dataset that can't survive scrutiny isn't an asset. What leaves Verso is derived signal: preferences, mentions, distributions. Every deliverable is aggregate. The unit of sale is a rate, a distribution, or a trend — never a row.

Request a sample