Multimodal Video Search Using Frame Embeddings and Audio Context

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional search systems struggle to accurately understand user intent from text and audio data alone, leading to unsatisfactory search results, often prompting users to seek answers through social media or forums.

Innovation Solution

A multimodal search system that processes video embeddings generated from user device-captured video data, combined with audio data, using machine-learned models to determine search results, incorporating visual cues and contextual information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional text and audio data are used for search queries, then the search system is simple to operate, but the search accuracy and ability to understand user intent deteriorate

Engineering Contradiction:
Improvesearch accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines multiple data modalities (video frames, audio data, text queries) into a unified search system. The machine-learned model processes all these inputs together to generate video embeddings that capture comprehensive user intent, thereby improving search accuracy while managing system complexity through integrated processing.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The search system is designed to handle multiple types of input data (video, audio, text) through a single machine-learned model architecture. This multi-functional approach allows the system to process diverse query types uniformly, improving accuracy without requiring separate specialized systems for each modality.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If video data is processed to generate video embeddings, then the search results become more comprehensive and intuitive, but the processing time and computational resources increase

Engineering Contradiction:
Improvesearch accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The video data processing is segmented into discrete frames that are processed individually to generate image embeddings. These frame-level embeddings are then aggregated into video embeddings, allowing parallel processing of multiple frames and reducing overall processing time while maintaining comprehensive video analysis.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system processes a selected subset of video frames rather than every single frame, using frame selection algorithms to identify key frames that best represent the video content. This partial processing approach maintains search accuracy while significantly reducing computational burden and processing time.

Inventive Principle:
Principle #16Partial or excessive action

3Productivity

If frame selection algorithms are used to process video frames, then the computational efficiency improves, but the complexity of the processing pipeline increases

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidprocessing pipeline complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

Frame selection is performed as a preliminary step before main video processing. By pre-selecting key frames based on simple criteria (temporal spacing, motion detection, or significance metrics), the system reduces the input data volume for subsequent processing stages, improving efficiency without requiring complex processing during the main search operation.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12602429B2Video and audio multimodal searching system
Publication Date: 2026.04.14 GOOGLE LLC
  • US12602429B2 patent drawing
  • US12602429B2 patent drawing
  • US12602429B2 patent drawing

AI summary

A multimodal search system using a video query is described. The system can receive video data captured by a camera of a user device. The video data can have a sequence of image frames. Additionally, the system can receive audio data associated with the video data captured by the user device. Moreover, the system can process, using one or more machine-learned models, the sequence of image frames to generate video embeddings related to the sequence of the image frames. The video embeddings can have a plurality of image embeddings associated with the sequence of image frames. Furthermore, the system can determine one or more video results based on the video embeddings and the audio data. Subsequently, the system can transmit, to the user device, the one or more video results.