Embed 5 is a Large Language Models (LLMs) tool. State-of-the-art embedding models for enterprise search, RAG, and document retrieval. Key features include Dual-Tier Architecture, Multimodal Document Processing, and Extended Context Window. Best for data scientists and analysts, software developers and engineers and financial advisors and analysts.
About Embed 5
Embed 5 is Cohere's most powerful embedding family, delivering frontier retrieval quality across multimodal, multilingual, and financial documents. Available in Pro and Fast tiers that share one embedding space for flexible enterprise workflows.
Key Features
Dual-Tier Architecture.
Multimodal Document Processing.
Extended Context Window.
Multilingual Capabilities.
Financial Document Excellence.
Flexible Output Formats.
Frequently Asked Questions
Embed 5 is a family of embedding models from Cohere that converts text, images, and documents into searchable vector representations. It powers enterprise search, RAG systems, and AI agents by finding relevant information before it reaches generative models. The Pro tier focuses on quality while the Fast tier optimizes for speed and cost.
Embed 5 Pro costs $0.12 per million text tokens and Embed 5 Fast costs $0.08 per million text tokens. Image inputs cost $0.40 per million tokens for either model. You can index documents with the more expensive Pro model and handle queries with the cheaper Fast model to optimize costs.
Yes. Both models share the same embedding space, which means you can build your index once with Pro and then query it with Fast without regenerating vectors. Cohere's testing shows this combination achieves 98.4% of Pro-only performance while reducing query costs by 33%.
Embed 5 is available through the Cohere API, Model Vault, Microsoft Foundry, Amazon SageMaker, and directly in North. You can deploy it via shared API access or use single-tenant deployment through Model Vault for private VPC or on-premises installations.







