# EmbeddingGemma 2 in the browser: one search box for photos, sound, video and notes

Source: https://tpiros.dev/blog/embeddinggemma-2-in-the-browser

I typed "cats sleeping" into a search box and got back a photo of 2 cats asleep on a pink blanket. Ordinary enough, except that nobody ever tagged that photo or wrote a caption for it, and the search never looks at filenames. It found the photo by what's in it.

Second place was a tiger. Third was a 1-second WAV of a cat meowing.

![Search results for "cats sleeping": cats.jpg first with a score of 0.754, then tiger.jpg, then cat_meow.wav](https://res.cloudinary.com/tamas/image/upload/f_auto,q_auto,w_900/blog/embeddinggemma-browser/search-cats-sleeping)

The search box belongs to a small app I built, and the whole thing runs in one browser tab. There's no server and no API key. Once the model has downloaded, you can unplug the network cable and it keeps working. The model is Google DeepMind's [EmbeddingGemma 2](https://huggingface.co/google/embeddinggemma-2), running on WebGPU through [Transformers.js](https://huggingface.co/docs/transformers.js). The code is on GitHub at [tpiros/embeddinggemma-browser-search](https://github.com/tpiros/embeddinggemma-browser-search).

This post walks through how it works, starting with what an embedding is, then the model, then the parts of running it in a browser that the model card mentions in a single sentence and that cost me an afternoon each.

## What an embedding is

An embedding model reads something and gives back a list of numbers. EmbeddingGemma 2 gives back 768 of them, whatever you feed it.

What matters is the direction the list points in. Things with similar meaning point in similar directions, so "find me things like this" becomes "find me vectors pointing roughly the same way".

EmbeddingGemma 2 normalises every vector to length 1. That puts every vector on the surface of a sphere, so the only thing that differs between 2 vectors is the angle between them. The similarity score is the cosine of that angle, which for unit vectors is just the dot product. You multiply the 2 lists element by element and add it all up.

768 dimensions are hard to draw, so here's the same idea in 2. Drag the blue query around the circle.

Notice that the clusters mix kinds of file. The cat cluster has a photo and a sound, and the dog cluster has a photo, a sound and a note about the vet. A text-only model can't do that. This one can, and that's most of why it's interesting.

## One model, 4 kinds of input

EmbeddingGemma 2 takes text (code included), images, audio and video, and maps all of them into the same 768-dimensional space. A spoken sentence and a written one about the same thing should land close together. So should a photo of a turtle and the phrase "a turtle swimming in the ocean".

It's 740M parameters in 3 parts:

- a 270M text model (a 130M transformer and a 140M embedding table)
- a 170M vision encoder, which handles photos and video frames
- a 300M audio encoder

The encoders are separate, and you can leave them out. Text-only search only needs the 270M. That matters a lot in a browser, where every parameter is something the visitor downloads.

The context window is 8,192 tokens, and every kind of input pays for its share in tokens. Text pays 1 token per subword. An image costs up to 280, a video frame 140, and a second of audio 25. Those numbers come back later, when they start to hurt.

## Following one file through the model

Here's what happens between dropping a file on the page and having a vector in the database. Pick a file and step through it. The numbers are measured on the real files, not estimated.

The step I find neat is the text model. The encoders produce "soft tokens", vectors that sit in a sequence exactly where text tokens would, and then the same 24-layer text model reads all of it. A photo of cats and the words "two cats asleep" go through identical layers at the end, which is why they come out comparable.

## Running it in a browser tab

The ONNX export lives at [`onnx-community/embeddinggemma-2-ONNX`](https://huggingface.co/onnx-community/embeddinggemma-2-ONNX), and Transformers.js 4.3.1 knows the `embedding_gemma2` architecture out of the box.

The model card's quick start uses the `feature-extraction` pipeline. That pipeline loads every encoder, so I used `AutoModel` and `AutoProcessor` directly, which lets you null out the encoders you don't want:

```js

const MODEL_ID = 'onnx-community/embeddinggemma-2-ONNX';

const config = await AutoConfig.from_pretrained(MODEL_ID);
if (!vision) config.vision_config = null; // skips 109 MB
if (!audio) config.audio_config = null; // skips 340 MB

const [processor, model] = await Promise.all([
  AutoProcessor.from_pretrained(MODEL_ID),
  AutoModel.from_pretrained(MODEL_ID, {
    config,
    device: 'webgpu',
    // 4-bit everywhere except audio, which the card says suffers most from it
    dtype: { model: 'q4', vision_encoder: 'q4', audio_encoder: 'q8' },
  }),
]);
```

With all 3 parts that's 624 MB: 175 for the text model, 109 for vision, 340 for audio. Transformers.js stores the files in the browser's Cache Storage, so the second visit loads in about 3 seconds instead of downloading again.

Calling it is the same for every kind of input. The processor takes `(text, images, audio, videos)`, and the model returns a `sentence_embedding`:

```js
async function run(...inputs) {
  const processed = await processor(...inputs);
  const output = await model(processed);
  // Copy out before the tensors are released
  const data = new Float32Array(output.sentence_embedding.data);
  return Array.from({ length: data.length / 768 }, (_, i) => data.subarray(i * 768, (i + 1) * 768));
}

await run(['task: search result | query: cats sleeping']); // text
await run(null, [[image]]); // an image
await run(null, null, [[pcm]]); // audio, mono 16 kHz
await run(null, null, null, [[video]]); // a RawVideo
```

Text needs a prefix that tells the model what job the embedding is for. Queries and documents get different ones:

```js
const prefixQuery = (text) => `task: search result | query: ${text}`;
const prefixDocument = (text, title) => `title: ${title?.trim() || 'none'} | text: ${text}`;
```

Images, audio and video take no prefix. I kept all of this inside a Web Worker, so the UI never builds a prompt and never blocks while the GPU is busy.

One catch with the worker is that workers have no `AudioContext` and no `<video>` element. Transformers.js's own `load_audio` and `load_video` helpers need both, so they can't run there. The page decodes the media itself and transfers raw buffers to the worker. For audio that's an `OfflineAudioContext`, which resamples and downmixes in one go:

```js
async function decodeAudio(blob) {
  const ctx = new AudioContext();
  const decoded = await ctx.decodeAudioData(await blob.arrayBuffer());
  await ctx.close();

  // 1 channel at 16 kHz: the format the audio encoder expects
  const offline = new OfflineAudioContext(1, Math.ceil(decoded.duration * 16_000), 16_000);
  const source = offline.createBufferSource();
  source.buffer = decoded;
  source.connect(offline.destination);
  source.start();
  return (await offline.startRendering()).getChannelData(0);
}
```

Google's own [reference demo](https://huggingface.co/spaces/webml-community/embeddinggemma-2-webgpu) does the same thing, which is reassuring.

## The nested list trap

This one doesn't throw an error, which is what makes it nasty.

```js
// 2 inputs, 2 vectors
await processor(null, [[cats], [corgi]]);

// 1 input made of both images, 1 vector
await processor(null, [cats, corgi]);
```

A flat list of images is one input containing several images. To embed them separately you need one list per item. Get it wrong and you'll index a single blended "cats-and-corgi" vector, and search will return slightly odd results forever. I checked it with 10 copies of the same photo in one call: 10 vectors came back, all identical, which is what you want.

## The 2,700-token wall

The model card has a warning box that's easy to skim past. On WebGPU, keep each batch under about 2,700 tokens. Above that, some ONNX Runtime WebGPU kernels exceed a GPU dispatch limit and the run fails.

2,700 is a third of the context window, and you hit it quicker than you'd think. Try the presets.

A 60-second video at the default 1 frame per second gets capped at 32 frames by the processor, which is 4,480 tokens and over the line. A 3-minute voice memo is 4,500 tokens, over the line too. So the worker cuts things down before they reach the model:

- audio goes into even chunks of 100 seconds or less (a 180-second file becomes 2 × 90 s rather than 100 s + 80 s)
- video gets 1 segment per minute, with at most 16 frames each
- images go at most 9 per batch

Then batches are packed by token cost. Padding makes a batch cost its longest item times the number of items, so that's what the packer counts:

```js
function pack(items, cost, budget = 2_600) {
  const batches = [];
  let batch = [];
  let max = 0;
  for (const item of items) {
    const nextMax = Math.max(max, cost(item));
    if (batch.length && (batch.length + 1) * nextMax > budget) {
      batches.push(batch);
      batch = [];
      max = 0;
    }
    batch.push(item);
    max = Math.max(max, cost(item));
  }
  if (batch.length) batches.push(batch);
  return batches;
}
```

With that in place, the 60-second video and the 3-minute audio file both went in on WebGPU with nothing in the console.

A long file now has several vectors, so search scores each file by its best segment. The app also remembers which segment won. Open a 3-minute recording from the results and the player starts at 1:30, where the match was.

![The viewer for ted-180s.wav, opened from search results, with the player already at 1:30 of 3:00](https://res.cloudinary.com/tamas/image/upload/f_auto,q_auto,w_900/blog/embeddinggemma-browser/viewer-matched-segment)

## Searching is the easy part

After all that, the search itself is a loop. Every stored vector is in memory, and a query is a dot product against each one:

```js
for (let r = 0; r < rows.length; r++) {
  let score = 0;
  for (let i = 0; i < dim; i++) score += query[i] * matrix[r * dim + i];
  // keep each file's best segment
  const current = best.get(rows[r].itemId);
  if (!current || score > current.score) best.set(rows[r].itemId, { score, segment: rows[r] });
}
```

No vector database and no approximate index. On the 20-file test library the ranking takes under a millisecond, and embedding a text query takes 15 to 40. Brute force should stay fine for a long while: 100,000 vectors at 768 dimensions is 77 million multiply-adds, which I'd expect a laptop to get through in around a tenth of a second (I haven't measured that one).

The fun part is what comes back. I asked for "a talk about the history of technology". First place went to a 34-second clip of Grace Hopper on a talk show. Nothing in its pictures or its filename says "technology". It matched on the soundtrack. Videos go through the vision encoder for their frames, and if they have audible sound the app also runs the soundtrack through the audio encoder and keeps both vectors.

![Results for "a talk about the history of technology": hopper.mp4 first, labelled "Matched on the soundtrack", then a 3-minute talk recording with its best match from 1:30 to 3:00](https://res.cloudinary.com/tamas/image/upload/f_auto,q_auto,w_900/blog/embeddinggemma-browser/search-soundtrack-match)

You can search with a picture too. I gave it a close-up of a ginger cat that wasn't in the library. It put the tiger first, then a monarch butterfly, then a corgi, and the actual cats 4th.

![Searching with a photo of a ginger cat: tiger.jpg, butterfly.jpg and corgi.jpg rank above cats.jpg](https://res.cloudinary.com/tamas/image/upload/f_auto,q_auto,w_900/blog/embeddinggemma-browser/search-by-image)

All 3 winners are orange. I like this result more than a perfect one, because it shows what the model weighs. For a photo query, colour and texture count for a lot, and "orange furry thing" is a fair description of a tiger. A text query of "a cat" doesn't have this problem, since it carries no colour at all.

Voice search goes through `MediaRecorder`, then the same `decodeAudio` as dropped files. I didn't get to test it with a real microphone (the automated browser I used has none), but the clip takes the exact path the audio files do.

## Shorter vectors

EmbeddingGemma 2 is trained with Matryoshka Representation Learning, named after the nesting dolls. The first 128 numbers of the vector are trained to be a usable embedding on their own, and so are the first 256 and the first 512. You can keep a prefix and throw the rest away.

The one rule is to re-normalise after slicing. A unit vector cut in half is no longer length 1, and different vectors lose different amounts. Skip this and the scores still look plausible while the ranking quietly goes wrong:

```js
function truncate(vector, dim) {
  const out = vector.slice(0, dim);
  let norm = 0;
  for (let i = 0; i < dim; i++) norm += out[i] * out[i];
  norm = Math.sqrt(norm);
  for (let i = 0; i < dim; i++) out[i] /= norm;
  return out;
}
```

The app stores the full 768 and truncates both sides at search time, so the dimension is a setting you can flip and compare. These are the real top-2 results for 5 queries at each size.

256 is the sweet spot here: a third of the memory, the same winners, and leads that mostly hold. At 128 the winners survived on this small set, but the gaps got thin enough that I wouldn't trust it on a bigger library of images and audio. The model card says the same: 128 is mostly for text-only workloads.

## Pull the plug

Running locally is half the job. The other half is making sure nothing quietly phones home. 3 things had to be true for the app to work offline.

**The model files come from the cache.** Transformers.js checks Cache Storage before the network, so a model that loaded once loads again without a connection.

**ONNX Runtime's wasm comes from my own origin.** By default Transformers.js points ONNX Runtime at jsDelivr. That works until you're offline, so the worker imports the files from `node_modules` and lets Vite bundle them:

```js

env.backends.onnx.wasm.wasmPaths = {
  mjs: new URL(ortMjs, self.location.href).href,
  wasm: new URL(ortWasm, self.location.href).href,
};
```

**The page itself is cached.** A small Vite plugin writes a service worker at build time, with the hashed file names in its precache list. It covers the HTML, the JS, the font and the 27 MB wasm file.

To check rather than hope, I replaced the worker's fetch with one that throws and reports every call:

```js
env.fetch = async (url) => {
  report(`NETWORK ${url}`);
  throw new TypeError(`offline probe: ${url}`);
};
```

A cached load made zero calls, and search worked. With DevTools set to offline, the production build reloaded from the service worker and searched fine. Your photos and recordings live in IndexedDB in that browser and never leave it, because there's nowhere for them to go.

The first visit is the one moment it needs the network, and the app says so up front, with the size of what you're about to download:

![The first-run screen: "Download the model to start", with checkboxes for text, images and video, and audio, and a total of 624 MB](https://res.cloudinary.com/tamas/image/upload/f_auto,q_auto,w_900/blog/embeddinggemma-browser/first-run)

Once that download finishes, try this yourself. Open DevTools, go to the Network tab, and switch the throttling dropdown from "No throttling" to "Offline". Then keep using the app: search for something, drop in a photo it hasn't seen, record a voice clip. It all still works, because everything it needs is already on your machine.

Honestly, this is my favourite part of the whole project. A 740M-parameter model ranking your photos, sounds and videos with the network switched off still feels a bit like a magic trick.

## The WASM catch

Not every browser has WebGPU, so the app falls back to WASM, which runs on the CPU. The model card's text example even says `device: "webgpu", // or "wasm"`. On WASM, though, `q4` didn't load:

```
Can't create a session. ERROR_CODE: 9, ERROR_MESSAGE: Could not find an implementation
for GatherBlockQuantized(1) node with name '/model/embed_tokens/Gather_Quant'
```

`q8` failed the same way, on a different node. So I pulled the graph files (they're small, since the weights sit in separate `_data` files) and grepped them. Every quantised text and vision variant, `q4`, `q4f16` and `q8` alike, uses a `GatherBlockQuantized` node for its embedding lookup, and the ONNX Runtime WASM build I had (1.31.0-dev) has no kernel for it. The fp32 and fp16 graphs don't use it, and neither does any audio variant.

So on WASM the app uses fp16 for text and vision, and q8 for audio. That's about 1.2 GB instead of 624 MB. It also embeds 1 image at a time, because a batch of 9 at fp16 ran out of wasm memory (`std::bad_alloc`). It works, and the scores land within 0.01 of WebGPU. It's also slow. Loading the model and embedding the same handful of test files took 26 seconds on WebGPU and 403 on WASM.

The settings panel shows both costs before you pick:

![The settings panel: vector size 768, 512, 256 or 128; encoder toggles with their sizes; WebGPU or WASM](https://res.cloudinary.com/tamas/image/upload/f_auto,q_auto,w_900/blog/embeddinggemma-browser/settings)

## What's left

The whole app is about 1,800 lines of plain JavaScript plus a stylesheet, with no framework and no backend. The main bundle is 30 KB and the worker 548 KB. The model is the heavy part.

To run it yourself, clone [the repo](https://github.com/tpiros/embeddinggemma-browser-search), then:

```sh
npm install
npm run dev
```

Open it in a browser with WebGPU and click "Load sample files" for the same 14 files and 3 notes used throughout this post. `NOTES.md` in the repo has the model card comparisons, the dimension results and the full WASM error.

![The app on a phone in dark mode, with "a lake in the mountains" returning the Moraine Lake photo first and a hiking note second](https://res.cloudinary.com/tamas/image/upload/f_auto,q_auto,w_900/blog/embeddinggemma-browser/mobile-dark)

Things I'd look at next: a proper test with a real microphone, the 512-dimension numbers I didn't measure, and whether the fp16 fallback can be avoided once ONNX Runtime's WASM build gets the `GatherBlockQuantized` kernel. I'd also want a bigger library before trusting 128 dimensions, and a less orange set of animals.
