Multimodal Video Search Index Construction for Accuracy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current video search technologies have low accuracy due to reliance on textual information, leading to users having to sift through multiple results to find the intended video, degrading user experience.

Innovation Solution

The method involves processing multimodal search data, including text, image, and video data, using pre-configured algorithms to generate semantic labels and vectorized descriptions, which are then used to query pre-constructed indices to accurately retrieve target videos.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If video search is based on textual information only, then the search system is simple to implement, but the search accuracy is low

Engineering Contradiction:
Improvesearch accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines multiple data modalities (textual information from video metadata, image data from video frames, and audio data from video content) into a unified search system. By merging these different types of data and processing them through integrated algorithms, the system achieves higher search accuracy while managing complexity through a coordinated multi-modal approach.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent transitions from single-dimensional textual search to multi-dimensional search by incorporating visual and auditory dimensions. Video frames are converted to image data, audio is processed separately, and all modalities are combined to create a comprehensive search space that operates across multiple dimensions simultaneously, significantly improving retrieval accuracy.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Ease of operation

If multiple video results are retrieved with the same name, then the search system returns comprehensive results, but the user experience degrades due to the need to click through multiple videos

Engineering Contradiction:
Improveuser experienceVSAvoidvideo identification information
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent replaces manual user interaction (clicking through multiple videos to identify the correct one) with automated content-based retrieval. By substituting the mechanical action of user clicking with intelligent algorithmic processing of video content across multiple modalities, the system directly retrieves the intended video, improving ease of operation while preserving accurate video identification.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If video data is processed into metadata and stream data with segmentation, then the processing precision is improved, but the processing time increases

Engineering Contradiction:
Improveprocessing precisionVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary segmentation of video stream data into video frames and extraction of video metadata before the actual search processing. By preparing the data in advance through segmentation and metadata extraction, the system improves processing precision for subsequent analysis while managing time loss through efficient pre-processing that enables faster retrieval during actual search operations.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11782979B2Method and apparatus for video searches and index construction
Publication Date: 2023.10.10 ALIBABA GROUP HOLDING LTD
  • US11782979B2 patent drawing
  • US11782979B2 patent drawing
  • US11782979B2 patent drawing

AI summary

Embodiments of the disclosure provide methods and apparatuses for video searches and methods and apparatuses for index construction. In one embodiment, the method comprises: upon receiving a search request input by a user to search for a target video, processing, based on a pre-configured algorithm, multimodal search data for the target video included in the search request; providing a processing result of the multimodal search data with regard to a corresponding pre-constructed index to search to obtain the target video.