Video Summarization Using Facial Recognition and Annotation Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Users face difficulties in efficiently specifying and editing video data captured on electronic devices, as they often need to manually view or edit lengthy videos to focus on specific persons, objects, or themes, which can be time-consuming and cumbersome.
Innovation Solution
A system and method that generates annotation data to select and summarize video segments based on user input, using facial recognition and annotation data to identify and prioritize video segments featuring selected faces or objects, allowing for the creation of focused video summarizations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If users manually view or edit lengthy video data to focus on specific persons or objects, then they can identify relevant segments, but the process becomes time-consuming and cumbersome
Solution Approach 1:
The system performs automatic video summarization without requiring manual user intervention. The processor autonomously identifies video segments containing selected faces using facial recognition technology, ranks them by significance, and generates the summary video automatically, eliminating the need for time-consuming manual editing while maintaining precise identification of relevant segments
Solution Approach 2:
The patent replaces the mechanical manual process of viewing and editing video with an automated computational system. The processor uses facial recognition algorithms and automated ranking mechanisms to identify and select video segments, substituting human manual operations with machine-based automation that is both faster and more precise
2Loss of information
If the system includes all video segments in the summarization, then completeness is maintained, but the summarization becomes lengthy and loses focus on specific persons or themes
Solution Approach 1:
The system selectively includes only the most significant video segments in the summarization rather than all segments. By ranking segments based on their association with selected faces and choosing only those above a threshold or in the top N positions, the system achieves partial inclusion that prioritizes quality and focus over complete coverage, producing a concise summary that highlights specific persons or themes
Solution Approach 2:
The patent applies different selection criteria to different video segments based on their individual characteristics. Each segment is evaluated independently for its relevance to selected faces, and only segments with high local quality (strong facial presence or thematic relevance) are included in the final summary, creating a focused compilation rather than a uniform inclusion of all content
3Measurement precision
If the system performs comprehensive analysis of all video segments, then accurate identification is achieved, but processing complexity and time increase
Solution Approach 1:
The system divides the video data into discrete segments and processes them individually through the facial recognition and ranking pipeline. By segmenting the video and applying automated analysis to each segment separately, the system achieves accurate identification of relevant portions without requiring complex holistic analysis of the entire video, reducing overall processing complexity while maintaining precision
Solution Approach 2:
The patent introduces annotation data as an intermediary layer between the video segments and the final summarization. The processor generates annotation data that captures facial presence and segment characteristics, which then serves as input for the ranking algorithm. This intermediary representation simplifies the processing by abstracting complex video content into structured data that is easier to analyze and rank
Data Source
AI summary
Devices, systems and methods are disclosed for improving a playback of video data and generation of a video summary. For example, annotation data may be generated for individual video frames included in the video data to indicate content present in the individual video frames, such as faces, objects, pets, speech or the like. A video summary may be determined by calculating a priority metric for individual video frames based on the annotation data. In response to input indicating a face and a period of time, a video summary can be generated including video segments focused on the face within the period of time. The video summary may be directed to multiple faces and/or objects based on the annotation data.


