Verified facts
What happened
The new arXiv paper “Retrieval Needs Multivectors: An Exponential Separation” provides an explicit family of query and document sets with relevance matrices where single-vector embeddings require exponential size to rank all relevant documents above irrelevant ones, while polynomial-size multi-vector embeddings suffice.
The paper introduces the ANDOR retrieval benchmark to represent these difficult cases and reports that state-of-the-art single-vector embedding models perform poorly in zero-shot testing, with only marginal improvement after fine-tuning.
The same summary reports that multi-vector models consistently outperform their single-vector counterparts on ANDOR and improve substantially with fine-tuning.
The available evidence is limited to the arXiv title and summary from a single source; the source body was not retained, so detailed methods, implementation requirements, and production performance cannot be verified here.
Business relevance