Lexical, Dense & Hybrid Information Retrieval
Keyword and meaning-based search found complementary evidence. Combining them improved retrieval on two public benchmarks; a fixed neural reranker then made the strongest system slower and worse.
Central findingHybrid retrieval improved aggregate MAP over the active lexical stage on both public benchmarks, while the tested cross-encoder reranker reduced quality and increased latency.
Evidence
The result in context
- SciFact ΔMAP, hybrid vs active lexical stage
- +0.0431
- 95% CI [0.0194, 0.0667].
- NFCorpus ΔMAP, hybrid vs active lexical stage
- +0.0331
- 95% CI [0.0260, 0.0406].
- Estimated SciFact hybrid end-to-end latency
- 154.8ms
- RRF total assembled from measured direct-system CPU timings.
- Automated extension tests
- 22
On this page
Question
When does modern semantic retrieval genuinely improve on a strong lexical engine, and what does it cost in latency, complexity and interpretability?
A field-aware lexical search engine followed by an independent study of dense retrieval, rank fusion, neural reranking and latency on public benchmarks.
Part I · Team Project
A field-aware lexical search engine
The historical system built a positional index over licensed TREC Disk 4 & 5 / Robust04 coursework data, combining BM25F field weighting, phrase and proximity logic and query expansion across 249 topics.
The historical team was Blazej Olszta, Muhamad Husaam Ateeq, Max Monaghan and Sulaiman Bhatti. Its strongest BM25F + phrase/proximity configuration reached 0.1961 MAP, 0.4040 P@10 and 0.4033 nDCG@10.
Provenance
Historical data is not redistributed
The original evaluation was run locally using the licensed TREC coursework data and associated artefacts, which are not redistributed in the public repository. The historical phase is therefore not independently rebuildable from the public repository alone.
Part II · Independent Extension
Lexical and semantic retrieval were tested together
Plain BM25F scored 0.6374 MAP on SciFact and 0.1443 on NFCorpus. The active lexical ladder stage, BM25F + phrase/proximity, scored 0.6349 and 0.1444. Hybrid reciprocal-rank fusion reached 0.6780 and 0.1776.
The paired-bootstrap comparison uses the active lexical stage: SciFact ΔMAP +0.0431, 95% CI [0.0194, 0.0667]; NFCorpus ΔMAP +0.0331, 95% CI [0.0260, 0.0406].
Retrieval system explorer
Keyword and meaning are complementary
Compare aggregate quality, CPU latency and two licensed public SciFact query examples.
Keyword search
BM25F
- MAP
- 0.6374
- Measured mean CPU query latency
- 29.05ms
Hybrid RRF vs active lexical stage
Benchmark query ID · 575
In domesticated populations of Saccharomyces cerevisiae, whole chromosome aneuploidy is very uncommon.
Direct-system CPU latency is measured. Hybrid RRF latency is estimated end to end; reranked latency is derived by adding measured reranker time. Examples are public SciFact claims licensed CC BY 4.0 and selected by absolute Hybrid RRF-minus-active-lexical ΔAP in the source query analysis. No document text is reproduced. AP is average precision for one query; MAP averages AP across benchmark queries.
Interactive evidence
Hybrid retrieval is evaluated against a controlled system ladder
MAP on SciFact and NFCorpus across six retrieval systems.
Read: Hybrid reciprocal-rank fusion reaches 0.6780 MAP on SciFact and 0.1776 on NFCorpus; gains differ materially by dataset.
Data table · 12 verified rows
| Dataset | System | Map | Mrr | Precision At10 | Ndcg At10 | System Dataset |
|---|---|---|---|---|---|---|
| Scifact | Bm25 | 0.625589 | 0.637026 | 0.086667 | 0.66407 | Bm25 · Scifact |
| Scifact | Bm25f | 0.637397 | 0.649577 | 0.087333 | 0.675244 | Bm25f · Scifact |
| Scifact | Bm25f Phrase Proximity | 0.634901 | 0.647988 | 0.087 | 0.67285 | Bm25f Phrase Proximity · Scifact |
| Scifact | Dense | 0.649167 | 0.662877 | 0.092 | 0.688526 | Dense · Scifact |
| Scifact | Hybrid Rrf | 0.677963 | 0.690002 | 0.094333 | 0.71741 | Hybrid Rrf · Scifact |
| Scifact | Hybrid Rerank | 0.572211 | 0.578673 | 0.084667 | 0.610178 | Hybrid Rerank · Scifact |
| Nfcorpus | Bm25 | 0.143626 | 0.522607 | 0.216409 | 0.308538 | Bm25 · Nfcorpus |
| Nfcorpus | Bm25f | 0.144308 | 0.521384 | 0.216409 | 0.309107 | Bm25f · Nfcorpus |
| Nfcorpus | Bm25f Phrase Proximity | 0.144425 | 0.521003 | 0.21548 | 0.308816 | Bm25f Phrase Proximity · Nfcorpus |
| Nfcorpus | Dense | 0.164291 | 0.526113 | 0.241486 | 0.328166 | Dense · Nfcorpus |
| Nfcorpus | Hybrid Rrf | 0.177567 | 0.560754 | 0.249536 | 0.345822 | Hybrid Rrf · Nfcorpus |
| Nfcorpus | Hybrid Rerank | 0.168683 | 0.529891 | 0.2387 | 0.329391 | Hybrid Rerank · Nfcorpus |
Interactive evidence
The ablation ladder keeps each retrieval addition attributable
SciFact MAP in the explicit control order, from BM25 through neural reranking.
Read: Fusion improves the controlled lexical and dense baselines, while the neural reranker gives back quality in this frozen evaluation.
Data table · 6 verified rows
| Dataset | System | Map | Mrr | Precision At10 | Ndcg At10 |
|---|---|---|---|---|---|
| Scifact | Bm25 | 0.625589 | 0.637026 | 0.086667 | 0.66407 |
| Scifact | Bm25f | 0.637397 | 0.649577 | 0.087333 | 0.675244 |
| Scifact | Bm25f Phrase Proximity | 0.634901 | 0.647988 | 0.087 | 0.67285 |
| Scifact | Dense | 0.649167 | 0.662877 | 0.092 | 0.688526 |
| Scifact | Hybrid Rrf | 0.677963 | 0.690002 | 0.094333 | 0.71741 |
| Scifact | Hybrid Rerank | 0.572211 | 0.578673 | 0.084667 | 0.610178 |
Interactive evidence
Candidate depth reveals where later reranking can and cannot help
Recall at four candidate depths for lexical, dense and hybrid retrieval on both datasets.
Read: The diagnostic separates candidate generation from reranking: documents absent from the candidate set cannot be recovered downstream.
Data table · 24 verified rows
| Dataset | System | Depth | Recall Pct | Configuration |
|---|---|---|---|---|
| Scifact | Bm25f Phrase Proximity | 10 | 79.011111 | Scifact · Bm25f Phrase Proximity · @10 |
| Scifact | Bm25f Phrase Proximity | 50 | 86.605556 | Scifact · Bm25f Phrase Proximity · @50 |
| Scifact | Bm25f Phrase Proximity | 1,000 | 96.5 | Scifact · Bm25f Phrase Proximity · @1000 |
| Scifact | Dense | 50 | 92.1 | Scifact · Dense · @50 |
| Scifact | Hybrid Rrf | 10 | 84.122222 | Scifact · Hybrid Rrf · @10 |
| Scifact | Hybrid Rrf | 100 | 96.833333 | Scifact · Hybrid Rrf · @100 |
| Nfcorpus | Bm25f Phrase Proximity | 50 | 20.947963 | Nfcorpus · Bm25f Phrase Proximity · @50 |
| Nfcorpus | Bm25f Phrase Proximity | 1,000 | 36.870095 | Nfcorpus · Bm25f Phrase Proximity · @1000 |
| Nfcorpus | Dense | 100 | 29.948514 | Nfcorpus · Dense · @100 |
| Nfcorpus | Hybrid Rrf | 10 | 16.95067 | Nfcorpus · Hybrid Rrf · @10 |
| Nfcorpus | Hybrid Rrf | 100 | 30.749552 | Nfcorpus · Hybrid Rrf · @100 |
| Nfcorpus | Hybrid Rrf | 1,000 | 61.373473 | Nfcorpus · Hybrid Rrf · @1000 |
Latency and ablation
The neural reranker was slower and worse
The tested cross-encoder reduced SciFact MAP from 0.6780 to 0.5722 and NFCorpus MAP from 0.1776 to 0.1687. Approximate CPU latency rose from 154.8ms to 553.3ms per SciFact query and from 21.2ms to 415.1ms on NFCorpus.
That negative result is useful: adding a more complex model is not automatically an upgrade.
Interactive evidence
Quality gains must be read alongside end-to-end latency
MAP with measured direct-system CPU latency, estimated Hybrid RRF totals and derived reranked totals.
Read: The neural reranker is slower and worse than hybrid RRF in the frozen study, so added model complexity is not treated as progress by default.
Direct systems are measured. Hybrid RRF is estimated end to end; Hybrid + reranker is derived by adding measured reranker time.
Data table · 12 verified rows
| Dataset | System | Latency Ms | Latency Nature | Map |
|---|---|---|---|---|
| Scifact | Bm25 | 19.408251 | Measured | 0.625589 |
| Scifact | Bm25f | 29.045203 | Measured | 0.637397 |
| Scifact | Bm25f Phrase Proximity | 141.595711 | Measured | 0.634901 |
| Scifact | Dense | 10.671352 | Measured | 0.649167 |
| Scifact | Hybrid Rrf | 154.826741 | Estimated | 0.677963 |
| Scifact | Hybrid Rerank | 553.256321 | Derived | 0.572211 |
| Nfcorpus | Bm25 | 4.902424 | Measured | 0.143626 |
| Nfcorpus | Bm25f | 4.993294 | Measured | 0.144308 |
| Nfcorpus | Bm25f Phrase Proximity | 9.999905 | Measured | 0.144425 |
| Nfcorpus | Dense | 9.750033 | Measured | 0.164291 |
| Nfcorpus | Hybrid Rrf | 21.221645 | Estimated | 0.177567 |
| Nfcorpus | Hybrid Rerank | 415.052197 | Derived | 0.168683 |
Limitations
What this evidence does not establish
- The historical TREC phase cannot be rebuilt from the public repository alone because licensed data and associated artefacts are not redistributed.
- The extension conclusions are specific to SciFact, NFCorpus, the selected encoders and a fixed reranker configuration.
Source and reproducibility
Trace the evidence
Source code, evaluation outputs and supporting material are available in the repository.
View repository- Public repositoryreports/figures/generated_summary.jsonCommit / evidence ID: 1cc2ea72b0d5fa8e94be28ef2c1e69626d0ad22b
- Historical coursework summaryreports/historical_coursework_summary.mdCommit / evidence ID: 1cc2ea72b0d5fa8e94be28ef2c1e69626d0ad22b
- Quality and latency frontierreports/figures/generated_summary.jsonCommit / evidence ID: 1cc2ea72b0d5fa8e94be28ef2c1e69626d0ad22b