Video Summarization Using Facial Recognition and Annotation Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users face difficulties in efficiently specifying and editing video data captured on electronic devices, as they often need to manually view or edit lengthy videos to focus on specific persons, objects, or themes, which can be time-consuming and cumbersome.

Innovation Solution

A system and method that generates annotation data to select and summarize video segments based on user input, using facial recognition and annotation data to identify and prioritize video segments featuring selected faces or objects, allowing for the creation of focused video summarizations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If users manually view or edit lengthy video data to focus on specific persons or objects, then they can identify relevant segments, but the process becomes time-consuming and cumbersome

Engineering Contradiction:
Improveability to identify relevant video segmentsVSAvoidtime required for manual video editing
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs automatic video summarization without requiring manual user intervention. The processor autonomously identifies video segments containing selected faces using facial recognition technology, ranks them by significance, and generates the summary video automatically, eliminating the need for time-consuming manual editing while maintaining precise identification of relevant segments

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual process of viewing and editing video with an automated computational system. The processor uses facial recognition algorithms and automated ranking mechanisms to identify and select video segments, substituting human manual operations with machine-based automation that is both faster and more precise

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Loss of information

If the system includes all video segments in the summarization, then completeness is maintained, but the summarization becomes lengthy and loses focus on specific persons or themes

Engineering Contradiction:
Improvecompleteness of video contentVSAvoidconciseness and focus of video summary
Core Design Contradiction:
Loss of informationVSProductivity

Solution Approach 1:

The system selectively includes only the most significant video segments in the summarization rather than all segments. By ranking segments based on their association with selected faces and choosing only those above a threshold or in the top N positions, the system achieves partial inclusion that prioritizes quality and focus over complete coverage, producing a concise summary that highlights specific persons or themes

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent applies different selection criteria to different video segments based on their individual characteristics. Each segment is evaluated independently for its relevance to selected faces, and only segments with high local quality (strong facial presence or thematic relevance) are included in the final summary, creating a focused compilation rather than a uniform inclusion of all content

Inventive Principle:
Principle #3Local quality

3Measurement precision

If the system performs comprehensive analysis of all video segments, then accurate identification is achieved, but processing complexity and time increase

Engineering Contradiction:
Improveaccuracy of video segment identificationVSAvoidcomplexity of video processing system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system divides the video data into discrete segments and processes them individually through the facial recognition and ranking pipeline. By segmenting the video and applying automated analysis to each segment separately, the system achieves accurate identification of relevant portions without requiring complex holistic analysis of the entire video, reducing overall processing complexity while maintaining precision

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces annotation data as an intermediary layer between the video segments and the final summarization. The processor generates annotation data that captures facial presence and segment characteristics, which then serves as input for the ranking algorithm. This intermediary representation simplifies the processing by abstracting complex video content into structured data that is easier to analyze and rank

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11393208B2Video summarization using selected characteristics
Publication Date: 2022.07.19 AMAZON TECH INC
  • US11393208B2 patent drawing
  • US11393208B2 patent drawing
  • US11393208B2 patent drawing

AI summary

Devices, systems and methods are disclosed for improving a playback of video data and generation of a video summary. For example, annotation data may be generated for individual video frames included in the video data to indicate content present in the individual video frames, such as faces, objects, pets, speech or the like. A video summary may be determined by calculating a priority metric for individual video frames based on the annotation data. In response to input indicating a face and a period of time, a video summary can be generated including video segments focused on the face within the period of time. The video summary may be directed to multiple faces and/or objects based on the annotation data.