Video Summary Generation Using User Viewpoint Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video summary systems struggle to capture a user's viewpoint effectively, as they rely on geometrical interpretations and limited deep learning methods like LSTM/GRU, which fail to retain past information and produce summaries that are not cost-effective or robust in capturing multiple viewpoints.
Innovation Solution
A method and electronic device that determine a user's viewpoint by analyzing subjective, objective, and physical parameters, using frame selection techniques based on input excitation and excitation parameters to identify key frames and produce video summaries dynamically, incorporating reinforcement learning to capture environmental inputs and user preferences.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If geometrical interpretation and limited deep learning methods (LSTM/GRU) are used for video summary production, then the system is simpler to implement, but the system fails to retain past information and cannot effectively capture user viewpoint
Solution Approach 1:
The patent changes the parameter approach from traditional geometrical interpretations to multi-dimensional parameter analysis including subjective parameters (occupation, age, preference), objective parameters (past history, present goal), and physical parameters (camera angle, location, ambient light). This comprehensive parameter change enables the system to accurately capture user viewpoint while maintaining reasonable system complexity through structured data collection and processing.
2Use of energy by moving object
If traditional deep learning methods are used for video summary production, then the system has lower computational requirements, but the system cannot dynamically adapt to user viewpoints and produces costly summaries
Solution Approach 1:
The patent implements dynamic adaptability by continuously determining user viewpoint based on multi-dimensional parameters and adjusting the video summary generation process in real-time. The system dynamically selects regions of interest and frames based on captured viewpoint information, enabling cost-effective production while maintaining high adaptability to user preferences and viewing contexts.
3Reliability
If neural networks are used to simulate user viewpoint, then the system can capture user perspective, but the neural network cannot retain past information to accurately simulate viewpoint
Solution Approach 1:
The patent applies preliminary action by collecting and storing multi-dimensional user parameters (subjective, objective, and physical) before viewpoint simulation occurs. This pre-collected information serves as a knowledge base that the system references during viewpoint determination, enabling accurate simulation while retaining contextual information about user preferences, history, and current state without relying solely on neural network memory.
Data Source
AI summary
A method of providing a video summary by an electronic device. The method includes receiving, by the electronic device, a video including a plurality of frames; determining, by the electronic device, at least one view point of a user viewing the video; determining, by the electronic device, at least one region of interest (ROI) of the user in at least one frame among the plurality of frames based on the at least one view point of the user; identifying, by the electronic device, a frame set from the plurality of frames including the at least one ROI based on determining the at least one ROI in the at least one frame; providing, by the electronic device, the video summary based on the identified frame set; and displaying the video summary on a display of the electronic device.


