Text-Mapped Video Frame Retrieval for Faster Shot Finding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video retrieval methods, such as manually fast-forwarding through videos to find specific shots, are inefficient and prone to missing the desired content, leading to suboptimal human-computer interaction.

Innovation Solution

A video retrieval method that utilizes a terminal equipped with a processor and memory to process retrieval instructions, employing neural networks for feature extraction and text description generation, constructing a mapping relationship between frame pictures and text sequences to accurately and efficiently locate target frames based on user input.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If manual fast-forwarding is used to search for specific shots in videos, then the user can find the needed content, but the retrieval efficiency is low and the user may miss the desired shot

Engineering Contradiction:
Improveretrieval efficiencyVSAvoidaccuracy of finding desired shot
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system performs preliminary action by pre-processing video content to extract key frames and generate text descriptions before retrieval is needed. This creates a ready-to-use index that enables rapid retrieval without manual traversal, thus improving both efficiency and reliability of finding desired shots

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary mechanism (text description and key frame mapping system) that mediates between the user's retrieval request and the actual video content. Instead of directly searching through entire videos, the system uses text-based intermediate representations to locate and present relevant shots, significantly improving retrieval accuracy and efficiency

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If the user manually traverses all videos to find a specific shot, then the search can be exhaustive, but the process is time-consuming and inefficient

Engineering Contradiction:
Improvecompleteness of searchVSAvoidtime for retrieval process
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system extracts key frames and their text descriptions from videos in advance, separating the essential searchable content from the full video streams. This extraction creates a compact representation that maintains search completeness while dramatically reducing the time required to traverse and search through video content

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

By performing preliminary extraction of key frames and text descriptions before retrieval operations, the system prepares the video data in advance. This allows exhaustive search capability to be maintained through the extracted representations while avoiding the time-consuming process of manually traversing complete video files

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If traditional video retrieval methods are used, then the system is simple, but the human-computer interaction is not intelligent enough

Engineering Contradiction:
Improveintelligence of human-computer interactionVSAvoidcomplexity of retrieval system
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent replaces manual mechanical traversal of video content with an intelligent system that uses text description matching and key frame retrieval. This substitution introduces neural network-based text processing and automated video analysis, significantly enhancing the intelligence of human-computer interaction while accepting increased system complexity

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentEP3796189B1Video retrieval method, and method and apparatus for generating video retrieval mapping relationship
Publication Date: 2025.08.06 CAMBRICON TECH CO LTD
  • EP3796189B1 patent drawingFigure 1~1a
  • EP3796189B1 patent drawingFigure 2
  • EP3796189B1 patent drawingFigure 3

AI summary

The present disclosure relates to a video retrieval method, a method, system and device for generating a video retrieval mapping relationship, and a storage medium. The video retrieval method comprises: acquiring a retrieval instruction, wherein the retrieval instruction carries retrieval information for retrieving a target frame picture; and obtaining the target frame picture according to the retrieval information and a preset mapping relationship. The method for generating a video retrieval mapping relationship comprises: performing a feature extraction operation on each frame picture in a video stream by using a feature extraction model so as to obtain a key feature sequence corresponding to each frame picture; inputting the key feature sequence corresponding to each frame picture into a text sequence extraction model for processing so as to obtain a text description sequence corresponding to each frame picture; and constructing a mapping relationship according to the text description sequence corresponding to each frame picture. By means of the video retrieval method and the method for generating a video retrieval mapping relationship provided in the present application, the efficiency of video retrieval can be improved, and human-machine interaction is made more intelligent.