Reference dashboard · MTEB RTEB (deu)
Independent preparation tool: compare public MTEB scores for retrieval, clustering and classification — updated daily via GitHub Actions. Not thesis-specific benchmark data.
Force-fetch runs in GitHub Actions (not in the browser). Each button opens the workflow to run manually. Status: ok = healthy, warn = expected gap, fail = action needed.
| # | Model | Avg | Retrieval | Clustering | Classification | Params | Dim | Price/M |
|---|---|---|---|---|---|---|---|---|
| 1 | codefuse-ai/F2LLM-v2-14B | 63.41 | 55.6 | 45.7 | 74.1 | 13.99M | 5120 MTEB API | — |
| 2 | codefuse-ai/F2LLM-v2-8B | 63.18 | 55.6 | 45.5 | 73.8 | 7.568M | 4096 MTEB API | — |
| 3 | codefuse-ai/F2LLM-v2-4B | 62.50 | 54.9 | 43.7 | 73.7 | 4.022M | 2560 MTEB API | — |
| 4 | codefuse-ai/F2LLM-v2-1.7B | 61.79 | 54.9 | 41.7 | 72.5 | 1.721M | 2048 MTEB API | — |
| 5 | codefuse-ai/F2LLM-v2-0.6B | 59.80 | 52.0 | 38.9 | 71.4 | 0.596M | 1024 MTEB API | — |
| 6 | Alibaba-NLP/gte-Qwen2-7B-instruct | 58.69 | 50.2 | 39.9 | 73.3 | 7.069M | 3584 MTEB API | — |
| 7 | Linq-AI-Research/Linq-Embed-Mistral | 58.45 | 54.5 | 40.7 | 67.5 | 7.111M | 4096 MTEB API | — |
| 8 | codefuse-ai/F2LLM-v2-330M | 58.33 | 51.5 | 36.6 | 69.6 | 0.334M | 896 MTEB API | — |
| 9 | Salesforce/SFR-Embedding-Mistral | 57.32 | 54.4 | 41.2 | 64.2 | 7.111M | 4096 MTEB API | — |
| 10 | | 56.90 | 52.7 | 40.9 | 63.8 | 7.111M | 4096 MTEB API | — |
| 11 | Alibaba-NLP/gte-Qwen2-1.5B-instruct | 56.72 | 53.6 | 36.6 | 67.0 | 1.543M | 8960 MTEB API | — |
| 12 | Alibaba-NLP/gte-Qwen1.5-7B-instruct | 56.43 | 49.6 | 39.6 | 68.4 | 7.099M | 4096 MTEB API | — |
| 13 | | 56.32 | 46.4 | 36.8 | 66.8 | 0.572M | 1024 MTEB API | — |
| 14 | | 55.84 | 49.3 | 40.2 | 61.3 | 0.56M | 1024 MTEB API | — |
| 15 | Salesforce/SFR-Embedding-2_R | 55.74 | 53.4 | 40.1 | 65.7 | 7.111M | 4096 MTEB API | — |
| 16 | Snowflake/snowflake-arctic-embed-l-v2.0 | 55.68 | 55.7 | 32.8 | 60.6 | 0.568M | 1024 MTEB API | $0.0700 LiteLLM |
| 17 | Lajavaness/bilingual-embedding-large | 55.29 | 47.9 | 34.9 | 62.6 | 0.56M | 1024 MTEB API | — |
| 18 | | 55.04 | 51.8 | 33.5 | 60.8 | 0.56M | 1024 MTEB API | — |
| 19 | OrdalieTech/Solon-embeddings-large-0.1 | 54.67 | 50.7 | 33.2 | 61.0 | 0.56M | 1024 MTEB API | — |
| 20 | Lajavaness/bilingual-embedding-base | 53.33 | 47.3 | 34.1 | 59.8 | 0.278M | 768 MTEB API | — |
| Compare | Model | Avg | Retrieval | Clustering | Class. | Dim | Price |
|---|---|---|---|---|---|---|---|
| codefuse-ai/F2LLM-v2-14B | 63.41 | 55.6 | 45.7 | 74.1 | 5120 MTEB API | — | |
| codefuse-ai/F2LLM-v2-8B | 63.18 | 55.6 | 45.5 | 73.8 | 4096 MTEB API | — | |
| codefuse-ai/F2LLM-v2-4B | 62.50 | 54.9 | 43.7 | 73.7 | 2560 MTEB API | — | |
| codefuse-ai/F2LLM-v2-1.7B | 61.79 | 54.9 | 41.7 | 72.5 | 2048 MTEB API | — | |
| codefuse-ai/F2LLM-v2-0.6B | 59.80 | 52.0 | 38.9 | 71.4 | 1024 MTEB API | — | |
| Alibaba-NLP/gte-Qwen2-7B-instruct | 58.69 | 50.2 | 39.9 | 73.3 | 3584 MTEB API | — | |
| Linq-AI-Research/Linq-Embed-Mistral | 58.45 | 54.5 | 40.7 | 67.5 | 4096 MTEB API | — | |
| codefuse-ai/F2LLM-v2-330M | 58.33 | 51.5 | 36.6 | 69.6 | 896 MTEB API | — | |
| Salesforce/SFR-Embedding-Mistral | 57.32 | 54.4 | 41.2 | 64.2 | 4096 MTEB API | — | |
| 56.90 | 52.7 | 40.9 | 63.8 | 4096 MTEB API | — | ||
| Alibaba-NLP/gte-Qwen2-1.5B-instruct | 56.72 | 53.6 | 36.6 | 67.0 | 8960 MTEB API | — | |
| Alibaba-NLP/gte-Qwen1.5-7B-instruct | 56.43 | 49.6 | 39.6 | 68.4 | 4096 MTEB API | — | |
| 56.32 | 46.4 | 36.8 | 66.8 | 1024 MTEB API | — | ||
| 55.84 | 49.3 | 40.2 | 61.3 | 1024 MTEB API | — | ||
| Salesforce/SFR-Embedding-2_R | 55.74 | 53.4 | 40.1 | 65.7 | 4096 MTEB API | — | |
| Snowflake/snowflake-arctic-embed-l-v2.0 | 55.68 | 55.7 | 32.8 | 60.6 | 1024 MTEB API | $0.0700 LiteLLM | |
| Lajavaness/bilingual-embedding-large | 55.29 | 47.9 | 34.9 | 62.6 | 1024 MTEB API | — | |
| 55.04 | 51.8 | 33.5 | 60.8 | 1024 MTEB API | — | ||
| OrdalieTech/Solon-embeddings-large-0.1 | 54.67 | 50.7 | 33.2 | 61.0 | 1024 MTEB API | — | |
| Lajavaness/bilingual-embedding-base | 53.33 | 47.3 | 34.1 | 59.8 | 768 MTEB API | — |
Strong multilingual retrieval with instruct tuning; good default for German RAG.
Starkes multilinguales Retrieval mit Instruct-Tuning; solider Default fuer Deutsch-RAG.
Smaller footprint while keeping competitive retrieval scores.
Kleineres Modell bei weiterhin konkurrenzfaehigen Retrieval-Werten.
German-focused encoder; strong on monolingual clustering benchmarks.
Deutsch-spezifischer Encoder; stark bei monolingualen Clustering-Benchmarks.
Lightweight multilingual baseline for grouping German text.
Leichtes multilinguales Basismodell zum Gruppieren deutscher Texte.
Balanced task coverage across MTEB classification suites.
Ausgewogene Task-Abdeckung ueber MTEB-Klassifikations-Suites.
Efficient option when latency and VRAM matter more than peak accuracy.
Effiziente Option wenn Latenz und VRAM wichtiger sind als Spitzen-Accuracy.
MTEB aggregates task-specific scores. Retrieval uses nDCG-style ranking metrics; clustering measures group quality on German datasets. Scores are normalized for comparison on the public leaderboard. Dimension and pricing data may come from HF Hub or secondary aggregators (LiteLLM, OpenRouter) and are labeled accordingly.