Retrieval quality you can measure.

Public datasets, complete query sets and inspectable retrieval profiles. CogLake evaluates result quality alongside access permissions and index consistency.

More than a single headline score.

SciFact: 5,183 documents and 300 queries. LegalQuAD: 200 documents. All profiles ran locally on an Apple M1 Pro with Docker Desktop. The table compares CogLake configurations within these datasets.

Higher is better. nDCG measures ranking quality; recall measures the share of relevant results retrieved.
DatasetProfilenDCG@10Recall@10
BEIR SciFactLexical0.38200.4645
BEIR SciFactDense0.47740.6129
BEIR SciFactHybrid0.50340.6614
BEIR SciFactRerank0.53600.6603
MTEB LegalQuADLexical0.20450.2400
MTEB LegalQuADDense0.39820.5500
MTEB LegalQuADHybrid0.29680.5150
MTEB LegalQuADRerank0.35380.5150
Download results as CSV

The right questions

Full qrels runs use the datasets’ relevance judgments. Different retrieval profiles show how exact terms, meaning and reranking contribute.

Verify access too

The recorded full runs found no ACL leaks or index inconsistencies. A separate multimodal suite passed all 25 retrieval cases.

Evaluate your own knowledge

Compare with your own documents, questions and roles. The right retrieval profile depends on the content and task.

Datasets and definitions: BEIR / SciFact · MTEB LegalQuAD

How well can CogLake find your knowledge?

Discuss an evaluation