Emotion Graph-Based Video Summary Generation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The increasing amount of multimedia data, particularly video content, requires efficient methods to analyze user emotions and provide summary videos that cater to individual user preferences, as traditional video delivery methods are no longer sufficient in a diverse and interest-based viewing environment.
Innovation Solution
A device and method that generate a summary video by analyzing user emotions during a first video and comparing them with emotion graphs of characters, objects, backgrounds, sounds, and lines in a second video to select relevant scenes for inclusion in the summary, using a processor to execute instructions for emotion graph generation and scene selection.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If traditional video delivery methods are used, then video content can be provided to users, but the content does not cater to individual user preferences and interests
Solution Approach 1:
The system performs preliminary emotion analysis on video content to generate emotion graphs and scene emotion scores before user viewing. This pre-processing enables rapid matching with user emotions during actual viewing without real-time analysis delays, resolving the contradiction between personalization capability and system complexity
Solution Approach 2:
The patent introduces emotion graphs and scene emotion scores as intermediary representations between raw video content and user preferences. These intermediaries enable efficient comparison and matching without requiring direct complex analysis of user emotions during viewing, thus reducing system complexity while maintaining adaptability
2Measurement precision
If emotion analysis is performed on user images during video playback, then personalized summary videos can be generated, but processing time and computational resources increase
Solution Approach 1:
The system generates emotion graphs for video scenes in advance, before user viewing occurs. This pre-computation of scene emotion characteristics eliminates the need for real-time emotion analysis during video playback, significantly reducing processing time while maintaining accurate emotion detection through the pre-analyzed emotion graphs
Solution Approach 2:
The patent analyzes only key frames or representative scenes to generate emotion graphs, rather than processing every frame of the video. This partial action approach maintains sufficient emotion detection accuracy while dramatically reducing computational time and resources required
3Manufacturing precision
If multiple emotion graphs (character, object, background, sound, line) are compared with user emotion graph, then accurate scene selection for summary video is achieved, but system complexity and computation increase
Solution Approach 1:
The patent combines multiple emotion graphs (character, object, background, sound, line) into a unified scene emotion score through weighted aggregation. This merging approach maintains high scene selection accuracy by considering multiple emotion dimensions while simplifying the comparison process to a single integrated metric, thus reducing system complexity
Data Source
AI summary
A method for generating a summary video includes generating a user emotion graph of a user watching a first video. The method also includes obtaining a character emotion graph for a second video, by analyzing an emotion of a character in a second video that is a target of summarization. The method further includes obtaining an object emotion graph for an object in the second video, based on an object appearing in the second video. Additionally the method includes obtaining an image emotion graph for the second video, based on the character emotion graph and the object emotion graph. The method also includes selecting at least one first scene in the second video by comparing the user emotion graph with the image emotion graph. The method further includes generating the summary video of the second video, based on the at least one first scene.


