New AI release ·

EmbeddingGemma 2: Features, Specs and How to Use It

EmbeddingGemma 2 is a lightweight, open multimodal embedding model by Google DeepMind that maps text, code, images, audio, and video into a single space.

In short
  • EmbeddingGemma 2 is an open model released under the permissive Apache 2.0 license.
  • It unifies text, code, images, audio, and video in a shared embedding space.
  • The model runs locally on devices, requiring as little as 191 megabytes of RAM.
  • Developers can truncate output vectors from 768 down to 128 dimensions to save space.
In this guide
  1. What is EmbeddingGemma 2?
  2. What are the key features of EmbeddingGemma 2?
  3. How does EmbeddingGemma 2 perform on benchmarks?
  4. What are the best use cases for EmbeddingGemma 2?
  5. How to get and use EmbeddingGemma 2?
  6. Frequently asked questions

What is EmbeddingGemma 2?

EmbeddingGemma 2 is an open-source, lightweight multimodal embedding model developed by Google DeepMind. Released on 6 October 2026, it maps combinations of text, images, audio, video, and code into a single, unified embedding space.

This model builds on the success of the original text-only model, which received more than 20 million downloads. It is designed to run efficiently on consumer hardware, enabling developers to build fast, offline search engines and privacy-focused AI applications.

EmbeddingGemma 2 At-a-Glance
Developer
Google DeepMind
Release Date
6 October 2026
License
Apache 2.0
Total Parameters
740 Million
Context Window
8,000 tokens
Base Architecture
Gemma 4

What are the key features of EmbeddingGemma 2?

EmbeddingGemma 2 delivers powerful multimodal capabilities while remaining small enough to run entirely on local devices. It introduces several major efficiency improvements over the first version of the model.

8KToken context window
~191MBRAM needed for text-only weights
~567MBRAM needed for full model
6xMax storage reduction via MRL

The model is modular by design, meaning you do not have to load all encoders at once. It requires 270 million parameters for text-only workloads, with optional vision (170 million) and audio (300 million) encoders.

Using Matryoshka Representation Learning, developers can dynamically truncate output vectors from 768 dimensions down to 512, 256, or 128 dimensions. This flexibility provides up to a 6x reduction in local database storage and active memory usage.

The extended context window allows the model to process up to 5.5 minutes of audio, 29 images, or 58 video frames directly on local hardware. When quantized, it runs smoothly on mobile hardware like the Google Pixel 11 Pro.

How does EmbeddingGemma 2 perform on benchmarks?

EmbeddingGemma 2 achieves top-tier scores among sub-1B parameter models across multiple language, vision, and audio tasks. According to Google DeepMind, the model matches or outperforms several larger specialist models.

The model shows a major improvement in processing programming languages. It delivers a 9.92-point increase on the Massive Text Embedding Benchmark (MTEB) Code dataset compared to its predecessor.

What are the best use cases for EmbeddingGemma 2?

EmbeddingGemma 2 is built specifically for on-device indexing, search, and retrieval pipelines. Processing embeddings locally allows developers to maintain strict data privacy and eliminate network latency.

Because it shares a text tokenizer and audio encoder with Gemma 4, developers can run both models together in a unified pipeline. This combination allows you to build local retrieval augmented generation (RAG) pipelines with a lower memory footprint.

Good for
  • Offline search of media libraries
  • On-device RAG pipelines
  • Local codebase indexing
  • Privacy-first mobile applications
Not ideal for
  • Very long audio files exceeding 5.5 minutes
  • Generating text or code directly (it only outputs embeddings)

How to get and use EmbeddingGemma 2?

Developers can easily access EmbeddingGemma 2 weights to build local cross-platform applications. Google DeepMind has worked with ecosystem partners to ensure immediate compatibility with popular AI frameworks.

What you'll get

By following these setup options, you can implement private, offline multimodal search and indexing directly on consumer devices.

  1. Download the model weights from Hugging Face or Kaggle. For models optimized specifically for local hardware, visit the LiteRT Community on Hugging Face.
  2. Choose an on-device deployment framework. Use Google AI Edge MediaPipe for quick embedding, retrieval, and decision tasks, or use LiteRT for custom model integration.
  3. Serve the model locally using your preferred development tools, such as transformers, sentence-transformers, MLX, vLLM, llama.cpp, SGLang, Ollama, or LMStudio.
  4. Store your generated embedding vectors using a local vector database like Qdrant.
  5. If you need to customize the model for specific business data, follow fine-tuning guides provided by Unsloth.

Frequently asked questions

What is EmbeddingGemma 2?
EmbeddingGemma 2 is an open, lightweight multimodal embedding model developed by Google DeepMind. It maps text, code, images, audio, and video into a single, unified embedding space for local search and retrieval.
What license does EmbeddingGemma 2 use?
EmbeddingGemma 2 is released under the commercially permissive Apache 2.0 license. This license allows developers and businesses to use, modify, and distribute the model for commercial applications.
How much RAM does EmbeddingGemma 2 require?
On a Google Pixel 11 Pro, the quantized version of EmbeddingGemma 2 requires approximately 191 megabytes of active RAM for text-only weights. The full multimodal model requires about 567 megabytes of active RAM.
What is the context window size of EmbeddingGemma 2?
EmbeddingGemma 2 has an 8,000 token context window, which is four times larger than the previous version. This context size allows the model to process up to 5.5 minutes of audio, 29 images, or 58 video frames locally.
Can I run EmbeddingGemma 2 offline?
Yes, EmbeddingGemma 2 is optimized for local, on-device performance on consumer hardware. Generating embeddings locally ensures data privacy, operates offline, and eliminates network latency.

Sources: Google · Google

Want your brand to be the answer AI gives?

See how ready your website is for ChatGPT, Gemini and Perplexity, free, in about ten seconds.

Related: Native app or web app? How to choose in 2026