Multimodal Detection System for Highlight Video Generation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for generating highlight videos from personal videos do not effectively capture human-centric activities and emotions, leading to sub-optimal results and a poor user experience, as they lack adaptability to viewer preferences and require human annotations for training.

Innovation Solution

A multimodal detection system that segments video content into clips, tracks human activities and emotions using pose and face analysis, and generates highlight videos based on user-selected actions and emotions, without requiring human annotations, using a combination of activity tracking, autoencoder, class detection, and adaptive filtering systems.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional highlight video generation methods are used, then the system can generate highlight videos, but they fail to capture human-centric activities and emotions effectively, resulting in sub-optimal quality

Engineering Contradiction:
Improvehighlight video qualityVSAvoidhuman-centric activity capture
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The video is segmented into multiple clips based on detected actions and emotions. The system divides the continuous video stream into discrete segments, each associated with specific action classes and emotion labels, enabling precise control over which segments are selected for the highlight video based on user preferences.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Different aspects of the video are analyzed with different levels of detail: pose estimation is performed for action detection, while face analysis is performed for emotion detection. The system applies appropriate processing quality to each local aspect based on its importance for capturing human-centric content.

Inventive Principle:
Principle #3Local quality

2Reliability

If existing methods are used, then highlight videos can be generated, but they require human annotations for training, increasing complexity and time requirements

Engineering Contradiction:
Improvetraining accuracyVSAvoidtraining process complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system performs self-training by automatically detecting actions and emotions from the video data itself. The action detection module and emotion detection module learn from the video content without requiring external human annotations, enabling the system to train itself on the patterns it needs to recognize.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs preliminary detection of actions and emotions across the entire video before generating the final highlight video. This preliminary analysis creates a database of action-emotion associations that can be used to automatically generate highlights without requiring subsequent manual training or annotation.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If the system analyzes all video content in detail, then human-centric activities and emotions can be captured accurately, but the processing time and computational resources increase significantly

Engineering Contradiction:
Improveemotion detection accuracyVSAvoidvideo processing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The video is divided into discrete action segments, and emotion detection is performed only on these segmented clips rather than continuously analyzing the entire video. This segmentation approach reduces the total processing time while maintaining accurate emotion detection for the relevant portions of the video.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs pose estimation and action detection on all video frames, but applies more computationally intensive face analysis and emotion detection only to frames or clips where actions are detected. This partial application of heavy processing only where needed reduces overall computational time while maintaining high accuracy.

Inventive Principle:
Principle #16Partial or excessive action

4Ease of operation

If the system generates highlight videos without user preference input, then the process is simpler, but the videos cannot be tailored to individual viewer preferences

Engineering Contradiction:
Improvevideo generation simplicityVSAvoiduser preference tailoring
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The system provides dynamic filtering options that allow users to adjust the highlight video generation based on their preferences. Users can dynamically select which action classes and emotion types they want to include or exclude, and the system re-generates the highlight video accordingly, making the process adaptable to individual needs.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system serves multiple functions: it can generate highlight videos with default settings for simplicity, or it can provide sophisticated filtering and customization options for users who want to tailor the video to their specific preferences. This multi-functionality allows the same system to serve both casual users and power users.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11574477B2Highlight video generated with adaptable multimodal customization
Publication Date: 2023.02.07 ADOBE INC
  • US11574477B2 patent drawing
  • US11574477B2 patent drawing
  • US11574477B2 patent drawing

AI summary

In implementations for highlight video generated with adaptable multimodal customization, a multimodal detection system tracks activities based on poses and faces of persons depicted in video clips of video content. The system determines a pose highlight score and a face highlight score for each of the video clips that depict at least one person, the highlight scores representing a relative level of the interest in an activity depicted in a video clip. The system also determines pose-based emotion features for each of the video clips. The system can detect actions based on the activities of the persons depicted in the video clips, and detect emotions exhibited by the persons depicted in the video clips. The system can receive input selections of actions and emotions, and filter the video clips based on the selected actions and emotions. The system can then generate a highlight video of ranked and filtered video clips.