Is Your New AI Search Actually Better? How to Prove It

Three weeks after launch, someone in the Monday growth meeting asks the only question that matters: is the new search better than the old one? The room goes quiet. There is a demo that looked great, a vendor deck with a big percentage on slide nine, and a dashboard showing more searches than last month. None of those answers the question.
For most US D2C brands, that silence is the real problem. Buying AI search is easy. Knowing whether it works is not.
Why this question is getting more expensive
The shoppers arriving at your store are changing. Adobe Analytics, analyzing more than 1 trillion visits to US retail sites, reported that AI-referred traffic grew 138% year over year in May 2026 and converted 54% better than non-AI traffic. Since October 2024, that traffic has grown 1,324%, according to the same analysis (as reported by Digital Commerce 360 on June 17, 2026).
Adobe's data describes how shoppers arrive, not what they type into your search box once they land. But it is reasonable to expect that people who have already described what they want to an assistant arrive with specific expectations: a use case, a constraint, a budget. A search box that returns "no results" or ten loosely related products wastes the most valuable traffic you have.
So the stakes of getting search relevance right are rising. The tooling to measure it has not kept pace.
How do you evaluate ecommerce AI search?
Answer-engine block: To evaluate ecommerce AI search, build a golden query set: 150 to 300 real shopper queries from your own logs, each with graded relevance labels from a merchandiser. Score every search change against it using NDCG@10 and zero-result rate, then confirm with an A/B test measuring revenue per search session.
The rest of this article explains each step.
Step 1 — Build a golden query set from your own logs
A golden query set is a fixed list of real queries with known good answers. It is the unit test suite for search.
Sample the queries that actually break things
Do not pick queries that make the demo look good. Pull from your search logs and stratify on purpose:
- Head queries: the highest-volume terms. They usually work, and they protect you from regressions.
- Long-tail queries: specific, low-volume searches where keyword matching is weakest.
- Zero-result and near-zero-result queries: your current failures.
- Conversational queries: "something to wear to an outdoor wedding in October."
- Attribute and constraint queries: size, material, price ceiling, compatibility.
- Exact-match queries: SKUs, model numbers and brand names, where semantic search can quietly get worse.
Two hundred queries is enough to start. Quality of coverage matters more than count.
Label results with the people who know the catalog
For each query, have a merchandiser or category owner grade the top results on a simple scale: 0 (irrelevant), 1 (related), 2 (good), 3 (ideal). This takes a few days, not months. It is also the step teams skip, which is why they cannot tell whether a new model helped.
Step 2 — Choose metrics that can fail
A metric that cannot go down is not a metric. Use a small set:
- NDCG@10: scores whether the best products appear near the top of the first ten results, using your graded labels.
- Zero-result rate: the share of queries returning nothing. Track a second version that flags zero results on products you actually have in stock.
- Search refinement and exit rate: how often a shopper reformulates or leaves right after searching.
- Revenue per search session: the business metric that decides whether the project paid for itself.
Report them by query segment, not as one average. An improvement on conversational queries can hide a regression on SKU lookups, and an average will bury it.
Step 3 — Gate every change, then test live
Treat the golden set like a regression suite. Run it before any change to ranking, embeddings, synonyms or the model. If NDCG@10 drops on any segment beyond a threshold you set in advance, the change does not ship.
Then validate with a live A/B test on revenue per search session. The offline set tells you whether a change is safe; the live test tells you whether it pays. You need both, because offline labels reflect your merchandisers' judgment and live traffic reflects what shoppers actually buy.
Where AI search commonly fails the test
Three patterns are worth testing for specifically, because they are easy to miss in a demo:
- Exact-match regression. Pure semantic retrieval can rank a conceptually similar product above the one with the exact model number. Hybrid approaches that combine keyword and vector retrieval exist for this reason, and your exact-match segment will show whether yours works.
- Stock blindness. A relevant product that is out of stock is a bad result. Include availability in your relevance labels.
- Vocabulary drift. Your catalog's language and your customers' language diverge over time. Re-sample the golden set from fresh logs every quarter.
At MnT Future, we treat the golden query set as the first deliverable of any AI search and recommendations engagement, before choosing a retrieval approach or a model. It gives the brand and the engineers one shared, testable definition of "better." We applied the same thinking to our own semantic search and shopping assistant build in MnT Commerce.
What to do this week
- Export 90 days of search queries and sort by volume and by zero results.
- Draw your 200-query sample across the six segments above.
- Ask your category owners to grade the top five results for each.
- Run your current search against it and record the baseline.
You will learn more from that baseline than from any vendor benchmark, and you will own it.
FAQ
What is a golden query set in ecommerce search?
A fixed list of real shopper queries, each with graded relevance labels, used to score search quality before and after every change.
What is NDCG@10?
A ranking metric that rewards placing the most relevant results at the top of the first ten, using graded relevance labels.
How many queries do I need?
Start with 150 to 300, stratified across head, long-tail, zero-result, conversational, attribute and exact-match queries.
Want a baseline on your own store?
If you are weighing AI search for your store, we will run a free agent-readiness audit and walk through how your search and product data would hold up against AI-driven shoppers. Or book a free strategy session and bring your toughest search queries.
Source: Adobe Analytics via Digital Commerce 360, "Adobe: AI-referred traffic to retail sites doubles in a year," June 17, 2026.
