Google DeepMind releases EmbeddingGemma 2, an open on-device model for multimodal search
The 740-million-parameter EmbeddingGemma 2 maps text, code, images, video and audio into a single embedding space, enabling private, offline search and retrieval on devices such as smartphones.

Google DeepMind on October 6 released EmbeddingGemma 2, an open, lightweight multimodal embedding model designed to run on consumer devices, according to an official Google blog post. The model natively maps combinations of text, code, images, video and audio into a unified embedding space, so that, for example, a voice memo can be used to find a specific video clip, or a text query can search hours of audio recordings.
Built on the Gemma 4 architecture and released under the commercially permissive Apache 2.0 license, EmbeddingGemma 2 has 740 million parameters. Google said the first EmbeddingGemma, a text-only model, has been downloaded more than 20 million times, with developers using it for on-device search tools and privacy-first retrieval augmented generation (RAG) pipelines.
The model is modular: text-only workloads need as little as 270 million parameters, with optional vision (170 million) and audio (300 million) encoders for full multimodal support. Using Matryoshka Representation Learning, developers can truncate output vectors from 768 dimensions to 512, 256 or 128, cutting storage by up to six times. With quantization, Google said, it needs about 191MB of active RAM for text-only weights and about 567MB for the full multimodal model on a Pixel 11 Pro. Its context window is 8,000 tokens, four times that of the first version.
According to Google, EmbeddingGemma 2 matches its predecessor’s multilingual text performance while improving MTEB Code scores by 9.92 points, from 68.76 to 78.68, and achieves leading scores among sub-1B multimodal embedding models on benchmarks such as MTEB Code and MAEB. These results are company-reported.
Weights are available on Hugging Face and Kaggle, with support in tools including transformers, sentence-transformers, MLX, vLLM, llama.cpp and Ollama; Google also published guidance on on-device deployment in a Google Developers blog post. Generating embeddings locally, Google says, helps keep data private, reduces latency and allows cross-modal search to work entirely offline.
Related articles

OpenAI drops another batch of mathematical breakthroughs
OpenAI has revealed solutions to a number of long-standing mathematics problems produced by an unreleased frontier model in a batch of 722 manuscripts, covering 372 result families that group related…

Tiny light-measuring chip helps stabilize 10× more combs, could shrink atomic clocks
Researchers have demonstrated a chip-based optical frequency comb that matched bulky tabletop systems in key...

OpenAI’s largest mathematics release tackles 4,000 problems with Lean-checked proofs
OpenAI has released a large collection of mathematical research produced by an internal frontier model,...

85-kW MARVEL microreactor gets liquid-metal coolant ahead of criticality test
A small nuclear reactor at Idaho National Laboratory is approaching its first criticality test after...


Nvidia-backed Lambda seeks up to $4 billion in final private round before planned 2027 IPO
AI cloud provider Lambda is raising up to $4 billion at a $14.5 billion pre-money valuation, led by Blackstone and Coatue, as its backlog jumped to $50 billion in September from $15 billion in June, The Wall Street Journal reported.