You remember the moment in a video, but not the filename. Someone was speaking, the scene was outdoors, and it might have been recorded months ago. Good luck scrolling through that camera roll.
Google's new EmbeddingGemma 2 is built for that kind of search. Announced on 6 October, it gives developers a way to connect text, pictures, speech and video without sending every search to a cloud service.
The same release brings something rather different: an experimental Mac meeting companion that can turn rough notes into fuller notes using the conversation happening around them.
Neither is a promise that AI suddenly understands everything on your device. The interesting bit is more practical. Search can work with meaning across different kinds of media, and the processing can happen locally.
A search model, rather than another chatbot
An embedding is a numerical representation of content. Pieces with similar meaning can sit near one another in that numerical space, even when they do not use the same words.
Google's developer guide explains how a search query and stored material can be compared this way. The model's job is to help retrieve relevant content. Producing a polished answer is a separate job, potentially handled by a generative model afterwards.
Think of it as finding the right page before writing the reply. A search that retrieves the wrong passage can still lead to an impressively fluent answer about the wrong thing.
That distinction matters for everyday use. If an app finds a moment in a recording, the useful next step is to play that moment and check it. A confident-looking result is a starting point, not proof that the underlying conversation has been interpreted correctly.
Your pictures and your words can meet in the same search
Google's announcement describes a 740-million-parameter model with components for text, vision and audio. Unlike a text-only search tool, it can put these inputs into a shared representation.
The company's mobile demonstrations include Instant Media Search, which searches photos and videos stored on a device. A query can be typed, or supplied through an example image. Video Moments Finder adds another possibility: locating a relevant section within a recording using its visual frames and audio.
Imagine searching for the part of a family video where someone talks about a particular place. That is an example of the kind of task the technology targets, not a claim that we tested it on a Malaysian family's footage.
It also explains why this release is more than a new keyword box. An old filename might tell you very little about the scene inside a clip. Different kinds of evidence can help the search narrow things down.
Those messy meeting notes have a new companion
Google AI Edge Foresight is the experimental Mac application in the release. It uses locally processed system and microphone audio to follow a meeting, including meetings outside a particular video-call platform.
Its notes feature combines a user's brief entries with relevant transcript context. The idea is that a few words written during a busy discussion can become a more useful record without making someone type every sentence.
The app also offers contextual assistance and retrieval from personal files. Google's description pairs EmbeddingGemma 2 with Gemma 4: one helps find relevant material, while the other helps produce language from that context.
9to5Google's report adds that the application can connect meeting-related material such as project files and Google Drive. That is worth distinguishing from local processing itself. A tool that works on-device can still have features involving connected accounts or files supplied from elsewhere.
Offline is useful. It does not make every question disappear
Google presents on-device inference as a way to use these capabilities without relying on an internet connection for every model operation. That could be valuable when connectivity is poor or when a developer wants content to stay on the device during processing.
It does not automatically tell you how every future app stores transcripts, handles account integrations or asks for recording permission. Those decisions belong to the application using the model.
Before trusting a meeting helper with work files, look at what it actually records and accesses. A local model and a complete app are different things. The former does not answer every question about the latter's behaviour.
Nor should a notes tool quietly become the authority on what colleagues agreed. Dates, amounts, names and decisions deserve a check against the relevant recording or document, especially when a short note could have several meanings.
The smaller version has trade-offs
Google says EmbeddingGemma 2 can produce vectors of different sizes. Keeping fewer numbers reduces storage requirements, but it can also change search quality.
Its developer guide reports that 256-dimensional outputs preserved about 95% of full-size performance on certain image, video and speech tests. Moving down to 128 dimensions reduced performance further on those tests. These are Google's reported results, not a universal guarantee for every collection of files.
The model supports a context window of 8,000 tokens. Google's launch material gives examples of input limits including roughly five and a half minutes of audio. That limit describes model input; it should not be mistaken for a promise that an entire hour-long meeting can simply be fed in as one unlimited request.
What about Malaysian languages and real-world messiness?
The public model card describes multilingual support across more than 100 languages. It also identifies limitations involving ambiguity, figurative language and bias in training material.
That makes local testing important. A broad multilingual label does not establish how well a particular app handles a meeting that switches between Bahasa Malaysia and English, includes several accents, or refers to an unfamiliar local name.
Google's earlier, text-only EmbeddingGemma announcement in September 2025 framed retrieval as a way to connect language models to relevant material. This new release extends that search idea across media rather than removing the need for evidence.
The appealing part is easy to picture: less hunting through files, and more time looking at the actual moment you needed. The sensible test remains just as simple. Can the tool bring you back to the right recording, picture or document — and can you verify what it found?
Sources & further reading
Prepared with AI assistance from linked reporting. The cover is an AI-generated editorial illustration. Spotted something we should correct?
