Invisible Metadata for Frame-Accurate 360 Video Event Triggering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing digital video systems face challenges in frame-accurate event triggering for 360-degree video playback, as visual timecodes integrated into equirectangular frames cause distortion when mapped to a sphere, and algorithms for hiding timecodes consume processing resources and affect frame rate.
Innovation Solution
A method of inserting non-image data, such as frame identifiers, into predetermined regions of digital video frames that become invisible when texture-mapped to a spherical mesh, allowing for frame-accurate event triggering without distortion or processing overhead.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If visual timecodes are integrated into equirectangular frames for frame identification, then frame-accurate event triggering is enabled, but distortion occurs when mapped to spherical mesh and processing resources are consumed
Solution Approach 1:
The patent extracts the frame identification function from visual timecodes and implements it through invisible metadata embedded in the video stream. This separates the identification data from the visual content, eliminating distortion issues while maintaining frame-accurate identification capability.
Solution Approach 2:
The patent introduces an intermediary metadata structure that mediates between the video frames and the event triggering system. This metadata contains timing information without requiring visual representation, thus avoiding processing overhead and distortion while enabling precise frame identification.
2Measurement precision
If algorithms are used to hide timecodes in video frames, then frame identification is maintained, but processing resources are consumed and frame rate is affected
Solution Approach 1:
The patent removes the need for visual timecode hiding algorithms by extracting frame identification information into separate metadata. This eliminates the processing overhead of manipulating visual pixels while preserving frame identification capability through lightweight metadata parsing.
Solution Approach 2:
The patent replaces the mechanical/visual approach of embedding timecodes in image pixels with a data-based approach using metadata structures. This substitution eliminates computationally intensive image processing operations while maintaining frame identification accuracy.
3Loss of information
If visual timecodes are embedded in equirectangular frames, then frame data is available for event triggering, but distortion occurs when texture-mapped to spherical mesh
Solution Approach 1:
The patent extracts frame identification data from the visual domain into a separate metadata domain. This extraction prevents the identification data from being subjected to spherical mapping transformations, thereby eliminating distortion while maintaining data availability for event triggering.
Solution Approach 2:
The patent introduces metadata as an intermediary layer that carries frame identification information independently of the visual content. This intermediary structure is not affected by texture mapping operations, thus preserving data integrity while allowing the visual content to be freely transformed.
Data Source
AI summary
A computer-implemented method of processing digital video is provided. The method includes determining at least one frame region the contents of which would be rendered substantially invisible, were frames of the digital video to be subjected to a predetermined texture-mapping onto a predetermined geometry; and inserting non-image data into at least one selected frame by modifying contents within at least one determined frame region of the selected frame. Another computer-implemented method of processing digital video is provided. The method includes, for each of a plurality of frames of the digital video: processing contents in one or more predetermined regions of the frame to extract non-image data therefrom; subjecting the frame to a predetermined texture-mapping onto a predetermined geometry, wherein after the texture-mapping the contents of the one or more predetermined regions are rendered substantially invisible; and causing the texture-mapped frame to be displayed. Another computer-implemented method of processing digital video is provided. The method includes for each of a plurality of frames of the digital video: extracting a frame identifier uniquely identifying a respective frame by processing contents in one or more predetermined regions of the frame; and for each of a different plurality of frames of the digital video: estimating the frame identifier based on playback time of the digital video. Another computer-implemented method of processing digital video is provided. The method includes causing frames of the digital video to be displayed; for a period beginning prior to an estimated time of display of an event-triggering frame and ending after the estimated time, causing a frame that is to be displayed prior to the beginning of the period to remain displayed; and after the period, executing at least one event associated with the event-triggering frame and resuming display of subsequent frames of the digital video. Systems and computer-readable media are also provided.


