Real-Time Video Stream Navigation Using Semantic Keyframes
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional video conferencing systems lack the ability to navigate previous video frames in real-time, leading to decreased usability and user experience as participants can only see live streams and not review past shared content.
Innovation Solution
A computer-implemented method that identifies semantically significant events in live video streams, associates them with keyframes, and allows users to display these keyframes in addition to or instead of the live stream upon user instruction, enabling efficient navigation back to past events.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If conventional video conferencing systems display only live streams, then the system complexity remains low, but the user experience deteriorates because participants cannot review past shared content
Solution Approach 1:
The patent segments the video stream into multiple keyframes representing different shared content types (screen sharing, webcam, document sharing, etc.). Each keyframe is independently identified and stored with metadata, allowing the system to navigate to specific segments without processing the entire video stream, thus improving user experience while maintaining manageable system complexity
Solution Approach 2:
The patent performs preliminary analysis of the video stream to identify and extract keyframes during the live streaming process. By pre-identifying semantically significant moments and storing them with their timestamps and content types, the system prepares navigation data in advance, enabling fast retrieval and playback of past events without real-time processing delays
2Adaptability or versatility
If the system stores and processes multiple video frames for navigation, then navigation capability is improved, but the loss of time for processing increases
Solution Approach 1:
The patent extracts only the essential keyframes and their metadata from the video stream, rather than storing or processing all video frames. By taking out only the semantically significant moments (screen sharing events, document presentations, etc.), the system achieves navigation capability with minimal processing time and storage requirements
Solution Approach 2:
The patent changes the parameter of video data representation from continuous video frames to discrete keyframe events with metadata. This parameter change allows the system to navigate video content efficiently by working with compressed event representations rather than full video data, reducing processing time while maintaining navigation versatility
3Loss of information
If keyframes are identified and stored for each semantically significant event, then information availability is improved, but the device complexity increases
Solution Approach 1:
The patent creates simplified copies of video content in the form of keyframes with associated metadata (timestamp, content type, duration). These copies capture the essential information of each shared event without requiring the full video data, improving information availability while keeping processing complexity manageable through efficient data representation
Data Source
AI summary
Described are systems and methods that allow a video conference participant to efficiently review semantically meaningful events within the video streams shared by their peers. Because participants are engaged in real-time communication, it is important to provide tools that let them quickly jump back to past events that were shown previously, not requiring them to manipulate a standard video timeline. Our techniques thus find meaningful events in the video stream such as scrolling pages, moving windows, typing text, etc. Using these techniques, a participant can then easily go back to e.g. the last PDF page shown by a peer, while still listening to the live audio.


