The right questions
Full qrels runs use the datasets’ relevance judgments. Different retrieval profiles show how exact terms, meaning and reranking contribute.
Retrieval benchmarks
Public datasets, complete query sets and inspectable retrieval profiles. CogLake evaluates result quality alongside access permissions and index consistency.
Recorded evaluation · 2026-08-20
SciFact: 5,183 documents and 300 queries. LegalQuAD: 200 documents. All profiles ran locally on an Apple M1 Pro with Docker Desktop. The table compares CogLake configurations within these datasets.
| Dataset | Profile | nDCG@10 | Recall@10 |
|---|---|---|---|
| BEIR SciFact | Lexical | 0.3820 | 0.4645 |
| BEIR SciFact | Dense | 0.4774 | 0.6129 |
| BEIR SciFact | Hybrid | 0.5034 | 0.6614 |
| BEIR SciFact | Rerank | 0.5360 | 0.6603 |
| MTEB LegalQuAD | Lexical | 0.2045 | 0.2400 |
| MTEB LegalQuAD | Dense | 0.3982 | 0.5500 |
| MTEB LegalQuAD | Hybrid | 0.2968 | 0.5150 |
| MTEB LegalQuAD | Rerank | 0.3538 | 0.5150 |
Full qrels runs use the datasets’ relevance judgments. Different retrieval profiles show how exact terms, meaning and reranking contribute.
The recorded full runs found no ACL leaks or index inconsistencies. A separate multimodal suite passed all 25 retrieval cases.
Compare with your own documents, questions and roles. The right retrieval profile depends on the content and task.
Datasets and definitions: BEIR / SciFact · MTEB LegalQuAD