Video Search Device Using Reference Image Feature Extraction and Movement Tracking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video search technologies do not efficiently allow users to specify the position, orientation, and movement of objects in videos, especially when the object is depicted differently in still images and videos.
Innovation Solution
A video search device and method that receives a still image with reference positions, extracts a reference image, and searches for similar frame images in videos, tracing movement tracks to find target frames where the object appears with specified positions and orientations, using image characteristic amounts and movement analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If traditional video search methods are used, then search speed is improved, but the ability to specify position, orientation and movement of objects is insufficient
Solution Approach 1:
The patent segments the object specification task into multiple independent components: position specification via reference positions, orientation specification via multiple reference positions, and movement specification via target positions. This segmentation allows users to specify each attribute independently, improving ease of operation while maintaining search efficiency.
Solution Approach 2:
The patent transitions from traditional single-image search to multi-frame video search by adding the time dimension. By specifying reference positions in one frame and target positions in another frame, the system enables specification of object movement trajectories, transforming the search from static to dynamic while maintaining operational simplicity.
2Ease of operation
If still image is used as search input, then ease of specifying object appearance is improved, but the ability to search for objects with different positions and orientations is insufficient
Solution Approach 1:
The patent performs preliminary actions by extracting multiple reference images from different positions and orientations within the still image before the actual video search. These pre-extracted reference images are stored and used during video search to match objects regardless of their position or orientation in the video frames, thereby enhancing adaptability while keeping the user interface simple.
Solution Approach 2:
The patent makes the still image search input multi-functional by enabling it to serve multiple purposes: specifying object appearance, specifying object position, specifying object orientation, and specifying object movement. This is achieved by allowing users to designate multiple reference positions within the same still image, which then generates multiple reference images that can match various object states in videos.
3Measurement precision
If multiple reference positions are specified, then orientation specification is improved, but the complexity of input operation increases
Solution Approach 1:
The patent implements self-service by allowing the system to automatically generate multiple reference images from the still image based on the reference positions specified by the user. The system then automatically performs the video search using these generated reference images, eliminating the need for users to manually create multiple search queries or perform complex operations, thus maintaining input simplicity while achieving precise orientation specification.
Data Source
Figure 1A
Figure 1B
Figure 1C
AI summary
The present invention provides a video search device and/or the like for accomplishing video searches in which a user easily specifies the position, orientation and/or the like of an object that should appear in a video. In a video search device (501), a receiver (502) receives input of a still image, two reference positions in the still image and two target positions in a video frame. An extractor (503) extracts a reference image containing the two reference positions from the still image. A searcher (504) searches for similar frame images in which local images similar to the reference image are depicted, from frame images included in the video, traces two movement tracks of two noteworthy pixels depicted at start positions corresponding to the two reference positions in a local image when time advances or regresses from a similar frame image in the video, searches for a target frame image where the two movement tracks approach two target positions and produces as search results videos containing the similar frame image and the target frame image.