E5-Mistral-7B Embedding (Microsoft 2023年)(イーファイブミストラル7ビー)
2023年Microsoft (Wang et al.)発表E5-Mistral-7B・Industry-leading Mistral-7B decoder embedding + Industry-leading 4096-dim + Industry-leading 32K context + Industry-leading MTEB top + Industry-leading instruction-tuned embedding。
概要
E5-Mistral-7B Embedding は、2023年Microsoft (Wang et al.)発表E5-Mistral-7B・Industry-leading Mistral-7B decoder embedding 2023年 + Industry-leading MTEB top position確立。E5-Mistral-7B specifications = Industry-leading Mistral-7B decoder embedding (Industry-leading E5-Mistral-7B 2023 + Industry-leading Mistral-7B decoder embedding + Industry-leading E5-Mistral flagship) + Industry-leading 4096-dim (Industry-leading 4096-dim Mistral-7B embedding) + Industry-leading 32K context (Industry-leading 32K tokens max context) + Industry-leading MTEB top (Industry-leading MTEB benchmark top 2023) + Industry-leading instruction-tuned embedding (Industry-leading instruction-tuned embedding)。
主な特徴・仕組み
- Org: Industry-leading Microsoft (Wang Liang, Yang Nan, Huang Xiaolong, Furu Wei et al.)
- Year: 2023年 (arXiv 2023年12月発表)
- Paper: Industry-leading "Improving Text Embeddings with Large Language Models" arXiv 2401.00368
- Base Model: Industry-leading Mistral-7B (7B params decoder-only LLM)
- Embedding Dim: Industry-leading 4096-dim embedding
- Max Context: Industry-leading 32K tokens max context (RoPE base extended)
- Architecture: Industry-leading decoder-based embedding (vs BERT encoder-based)
- Training: Industry-leading synthetic data instruction-tuned + GPT-4 generated tasks
- MTEB Benchmark: Industry-leading MTEB benchmark top 2023 (English)
- Synthetic Data: Industry-leading GPT-4 synthetic data 500K samples 93 languages
- HuggingFace: Industry-leading intfloat/e5-mistral-7b-instruct
- License: Industry-leading MIT License open source
- Industry-Leading: Industry-leading Mistral-7B decoder embedding instruction-tuned
スペック比較表
| LLM Embedding Model (2023-2024) | Dim | Org | License | Industry Position |
|---|---|---|---|---|
| E5-Mistral-7B | 4096-dim Mistral-7B | Microsoft | MIT open source | Industry-leading 7B Mistral embedding |
| BGE Large | 1024-dim BERT-large | BAAI | MIT open source | Industry-leading open-source MTEB top |
| Nomic Embed v1 | 768-dim Nomic-BERT | Nomic AI | Apache 2.0 fully open | Industry-leading fully open + reproducible |
| OpenAI text-embedding-3-large | 3072-dim | OpenAI | Closed API | Industry-leading OpenAI closed API |
| Voyage-3 | 1024-dim | Voyage AI | Closed API | Industry-leading Voyage closed API |
具体例・対応製品
- E5-Mistral-7B (2023年, Microsoft): Industry-leading Mistral-7B decoder embedding
- Industry-leading 4096-dim Mistral-7B 7B params: Industry-leading 4096-dim Mistral-7B
- Industry-leading 32K tokens max context: Industry-leading 32K context
- Industry-leading synthetic data instruction-tuned GPT-4: Industry-leading synthetic GPT-4
- Industry-leading decoder-based embedding vs BERT encoder: Industry-leading decoder-based vs BERT
- 競合 BGE Large + Nomic Embed + OpenAI + Voyage: Industry-leading embedding competitors
自作PCでの選び方・注意点
E5-Mistral-7B Embedding は「Industry-leading Mistral-7B decoder embedding + Industry-leading 4096-dim」「Industry-leading 32K context + Industry-leading MTEB top + Industry-leading instruction-tuned」用途のIndustry-leading Mistral-7B decoder embedding 2023年Microsoft発表product。Industry-leading 4096-dim Mistral-7B (Industry-leading 4096-dim Mistral-7B embedding + Industry-leading Mistral-7B 7B params decoder-only LLM + Industry-leading E5-Mistral 4096-dim signature) で Industry-leading 4096-dim Mistral-7B + Industry-leading Mistral-7B decoder + Industry-leading E5-Mistral 4096-dim signature。Industry-leading 32K context (Industry-leading 32K tokens max context + Industry-leading RoPE base extended + Industry-leading E5-Mistral long context signature) で Industry-leading 32K + Industry-leading RoPE extended + Industry-leading long context signature。Industry-leading decoder-based (Industry-leading decoder-based embedding vs BERT encoder-based + Industry-leading E5-Mistral decoder signature + Industry-leading decoder vs encoder paradigm shift) で Industry-leading decoder-based + Industry-leading E5-Mistral decoder + Industry-leading decoder vs encoder shift。Industry-leading synthetic GPT-4 (Industry-leading GPT-4 synthetic data 500K samples 93 languages + Industry-leading synthetic data instruction-tuned + Industry-leading E5-Mistral synthetic GPT-4 signature) で Industry-leading 500K GPT-4 synthetic + Industry-leading instruction-tuned + Industry-leading E5-Mistral synthetic signature。Industry-leading MTEB top + MIT (Industry-leading MTEB benchmark top 2023 English + Industry-leading MIT License open source + Industry-leading intfloat/e5-mistral-7b-instruct HuggingFace + Industry-leading Wang Liang + Yang Nan + Furu Wei Microsoft authors) で Industry-leading MTEB top + Industry-leading MIT License + Industry-leading intfloat HuggingFace + Industry-leading Microsoft authors。但しIndustry-leading BGE Large + Nomic + OpenAI + Voyage competition (Industry-leading BGE Large BAAI 1024-dim BERT-large 2023 + Nomic Embed v1 768-dim Apache 2024 + OpenAI text-embedding-3-large 3072-dim closed 2024 + Voyage-3 1024-dim closed 2024 vs E5-Mistral-7B Microsoft 4096-dim Mistral-7B 32K MTEB top 2023 trade-off) で Industry-leading 4-embedding model competitors vs E5-Mistral-7B + Industry-leading 4096-dim Mistral-7B decoder + 32K context + GPT-4 synthetic 500K instruction-tuned + MTEB top + MIT open source + Microsoft 2023 unique advantage adoption alignment必須。
関連用語との違い
- vs BGE Large (BAAI 2023): E5-Mistralは4096-dim Mistral-7B decoder + 32K + synthetic GPT-4・BGEは1024-dim BERT-large encoder + 512 tokens + 2-stage
- vs Nomic Embed v1 (Nomic 2024): E5-Mistralは7B decoder + 32K・Nomicは768-dim BERT + Matryoshka + 8K
- vs OpenAI/Voyage (Closed 2024): E5-Mistralはopen-source MIT 7B・OpenAI/Voyageはclosed API only
よくある質問(FAQ)
Q1: E5-Mistral-7B vs BGE Large 違いは? A: E5-Mistral-7B (Industry-leading 4096-dim Mistral-7B embedding + Mistral-7B 7B params decoder-only + 32K tokens max context + RoPE base extended + decoder-based embedding + synthetic data instruction-tuned GPT-4 + 500K samples 93 languages + MTEB top + Wang Liang + Furu Wei Microsoft 2023) vs BGE Large (Industry-leading 1024-dim BERT-large encoder + BERT-large 335M params + 512 tokens max + 2-stage training pre-training + fine-tuning with hard negatives + MTEB top + bilingual EN+ZH + BAAI 2023)・Industry-leading 4096-dim Mistral-7B + decoder + 32K + GPT-4 synthetic + Microsoft = E5-Mistral + Industry-leading 1024-dim BERT-large + encoder + 2-stage + bilingual + BAAI = BGE preference judgment。
Q2: Industry-leading 4096-dim Mistral-7B decoder value は? A: Industry-leading 4096-dim Mistral-7B (Industry-leading 4096-dim Mistral-7B embedding + Industry-leading Mistral-7B 7B params decoder-only LLM + Industry-leading decoder-based embedding vs BERT encoder-based + Industry-leading E5-Mistral 4096-dim signature + Industry-leading decoder vs encoder paradigm shift)。
Q3: Industry-leading 32K + GPT-4 synthetic value は? A: Industry-leading 32K + synthetic (Industry-leading 32K tokens max context + Industry-leading RoPE base extended + Industry-leading GPT-4 synthetic data 500K samples 93 languages + Industry-leading synthetic data instruction-tuned + Industry-leading E5-Mistral synthetic GPT-4 signature)。
まとめ
E5-Mistral-7B = 2023年Microsoft発表のMistral-7B decoder embedding。Industry-leading 4096-dim Mistral-7B embedding + Industry-leading Mistral-7B 7B params decoder-only LLM + Industry-leading 32K tokens max context + Industry-leading RoPE base extended + Industry-leading decoder-based embedding vs BERT encoder-based + Industry-leading synthetic data instruction-tuned + Industry-leading GPT-4 synthetic data 500K samples 93 languages + Industry-leading MTEB benchmark top 2023 + Industry-leading Wang Liang + Yang Nan + Furu Wei Microsoft authors + Industry-leading intfloat/e5-mistral-7b-instruct HuggingFace + Industry-leading MIT License open source + Industry-leading Mistral-7B decoder embedding instruction-tuned Microsoft 2023 position確立。