What does EmbeddingGemma 2 do?

EmbeddingGemma 2 produces vector embeddings, not chat text. Given text (including code), images, video frames, and/or audio, it returns a single vector in a shared 768-dimensional space so you can run semantic search, RAG, classification, clustering, and cross-modal retrieval on consumer hardware. Google’s model card (updated 2026-10-06) lists 740M total parameters: a 270M text backbone plus modular vision (170M) and audio (300M) encoders you can load selectively.

It builds on the Gemma 4 architecture and ships under Apache 2.0. Context is 8,192 tokens shared across modalities. Native output is 768-d; Matryoshka Representation Learning (MRL) lets you truncate to 512, 256, or 128 dimensions (re-normalize after slicing). Official quick start uses SentenceTransformers with `google/embeddinggemma-2` on Hugging Face. Checked against the Google AI for Developers model card and developers.googleblog.com AI Edge post on 2026-10-11.

People also ask “What is Google Gemma?” and “Is Gemini embedding free?”: Gemma is Google’s open model family; EmbeddingGemma 2 is the embedding member for multimodal vectors. Gemini embedding APIs are separate Google Cloud / Gemini products with their own pricing—open weights on Hugging Face are not the same as a billed Gemini embedding endpoint. This guide does not sell keys or host inference.

Search intent around “embeddinggemma 2” today clusters on download, Hugging Face, multimodal capability, GGUF/on-device packs, and vs older EmbeddingGemma (300M text). Ranking pages include ai.google.dev docs, DeepMind pages, Hugging Face model cards, arXiv for the first EmbeddingGemma, and developer blogs. An independent field guide that answers “what it is / how to load / multimodal vs text-only / limits” fills gaps without claiming official status.

Practical next steps: open the model card for architecture and eval tables; use our Hugging Face page for download and SentenceTransformers prompts; use the multimodal page for image/video/audio token budgets and selective encoder loading. Keep this tab for definitions; keep Google and Hugging Face for weights, licenses, and support. Facts above checked 2026-10-11.

Four facts before you download

Official sources

Model card: https://ai.google.dev/gemma/docs/embeddinggemma/model_card_2 · HF: https://huggingface.co/google/embeddinggemma-2 · Blog: https://blog.google/innovation-and-ai/technology/developers-tools/embeddinggemma-2/

Apache 2.0 weights

Commercially permissive per Google’s public posts. Still follow the Gemma Prohibited Use Policy when you deploy.

Selective footprint

Text-only ~270M; text+image ~440M; text+audio ~570M; full multimodal 740M (per model card config_kwargs table).

Limits of this guide

We summarize public docs only. Benchmarks, pack sizes, and library APIs can change. Prefer official cards over third-party mirrors.

Limitations

This site is not affiliated with Google or Google DeepMind. We do not host EmbeddingGemma 2, issue API keys, or provide vendor support. Embedding models are not safety-tuned chat models: Google’s card says developers must add application-level safeguards. Do not run float16 (use bfloat16 or float32). Omitting task instruction prefixes on text lowers quality. 128-d truncation hurts multimodal quality substantially—validate on your data. Language quality is uneven across 100+ languages.

Primary links to keep open

Sources used for this guide (checked 2026-10-11).

Go deeper on one intent

Stay on the same keyword: Hugging Face download, or multimodal setup.

Questions & answers

What does EmbeddingGemma do?

It turns text, and in version 2 also images, video, and audio, into vectors for search, RAG, classification, and clustering. It does not generate chat replies. Source: Google EmbeddingGemma 2 model card, checked 2026-10-11.

What is the best embedding model right now?

There is no single winner. EmbeddingGemma 2 is positioned by Google as strong under 1B for open multimodal on-device use. Hosted Gemini/OpenAI embedding APIs may win on managed scale. Compare on your modalities and latency; we do not publish invented leaderboard ranks.

What is Google Gemma?

Gemma is Google’s open model family. EmbeddingGemma 2 is the embedding model in that family built on Gemma 4 ideas, focused on vectors rather than dialogue. See ai.google.dev/gemma.

Is Gemini embedding free?

Gemini embedding APIs are Google Cloud / Gemini products with their own quotas and pricing—separate from downloading open EmbeddingGemma 2 weights. Open weights still cost you compute. Check current Gemini pricing pages; this site does not sell API access.

Where do I download EmbeddingGemma 2?

Primary public weights: https://huggingface.co/google/embeddinggemma-2. Also see the Google AI model card for library snippets. Community GGUF packs are third-party—verify before production.

How is EmbeddingGemma 2 different from EmbeddingGemma 1?

Version 2 adds native multimodal embeddings (image/video/audio), a smaller text backbone with better code MTEB scores on Google’s card (~14% code improvement claimed), 8K context, and MRL dims 128–768. Older HF ids like embeddinggemma-300m refer to the prior text model.