Verified facts
What happened
Clarify-Then-Search evaluates a workflow in which a clarifier asks questions, a user-answer component supplies only information stated in the intended query, a rewriter reformulates the original query, and WebDancer performs the resulting deep search.
The benchmark contains 518 curated instances based on real-world query data from the Baidu search engine, pairing an intended query with an underspecified query.
The benchmark reports that clarification improved performance over the no-interaction baseline when one question was allowed, with generally greater gains under larger question budgets; GPT-5.2 led at one question, while ERNIE-4.5-Turbo-128K led overall at three questions.
The supplied evidence is limited to the arXiv:2608.20357v1 title and abstract; it does not establish commercial availability, production deployment, or performance on merchant-specific search data.
Business relevance