Audio Video Stream Interstitial Removal via Closed Captioning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio/video stream processing methods, such as those used in digital video recorders (DVRs), require manual user intervention to skip commercial breaks, leading to inaccuracies and inconvenience, as they lack automated detection and removal of interstitial content like commercials.
Innovation Solution
A method for processing an audio/video stream that identifies segment boundaries using closed captioning data with offset information, allowing for automated skipping of interstitials and insertion of substitute content, ensuring accurate temporal alignment and user-defined playback options.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If manual fast forwarding is used to skip commercial breaks, then users can control playback, but the operation is inconvenient and inaccurate
Solution Approach 1:
The system automatically identifies program segments and commercial breaks using closed captioning data and video transition detection, eliminating the need for manual user intervention. The playback device self-determines segment boundaries by analyzing temporal offsets between captioning data and video content, then automatically skips interstitials during playback without requiring user control inputs.
Solution Approach 2:
The system changes the parameter of segment identification from manual temporal positioning to automated detection using closed captioning temporal offset parameters. By utilizing the temporal offset information embedded in closed captioning data and correlating it with video transition points, the system precisely identifies segment boundaries without manual intervention, resolving the contradiction between ease of operation and identification accuracy.
2Measurement precision
If automated detection of commercial breaks is implemented, then skipping accuracy is improved, but the device complexity increases
Solution Approach 1:
The system uses closed captioning data as an intermediary to identify program segments. Instead of directly analyzing video content for segment boundaries, the patent leverages the temporal offset information already present in closed captioning data as a mediator. This intermediary approach simplifies the detection process by using existing metadata rather than requiring complex video analysis algorithms.
Solution Approach 2:
The patent replaces manual mechanical fast-forward operations with an automated electronic detection system. The mechanical user action of pressing fast-forward buttons is substituted with an electronic process that automatically detects segment boundaries using closed captioning temporal offsets and video transition analysis, thereby improving precision while managing complexity through software-based automation.
3Measurement precision
If closed captioning data is used to identify segment locations, then temporal alignment accuracy is improved, but synchronization issues may occur due to temporal shifts
Solution Approach 1:
The system employs feedback by detecting video transitions at the temporal locations identified from closed captioning data. After initially locating potential segment boundaries using captioning temporal offsets, the system verifies these locations by checking for actual video transition points. This feedback mechanism ensures that the identified boundaries are accurate and synchronized with the actual video content, resolving synchronization issues.
Solution Approach 2:
The system performs preliminary identification of segment boundaries using closed captioning temporal offset data before final verification. The closed captioning information provides advance temporal location estimates, which are then refined by detecting actual video transitions at those predicted locations. This preliminary action allows the system to efficiently narrow down potential boundary points before conducting the final verification check.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Locations within an audio/video stream (1 04A) are identified by processing associated text data using autonomous location information referencing the text data. The identified locations within the audio/video stream may be utilized to identify boundaries (208, 210) of segments of content within the stream and interstitials of alternative content. Processing is performed to determine whether the identified boundaries possess specific characteristics, for example, have a substantially black screen and/or have muted audio. If the identified boundaries do not possess the identified characteristics, then additional processing is performed to identify other locations, temporally near the identified boundaries, which correspond with boundaries of the portion of the audio/video stream identified by the autonomous location information. The audio/video stream can then be processed to remove or replace segments or interstitials identified by the boundaries to produce a second audio/video stream (112A). This enables a user to view programming, for example, from which advertisements or other unwanted content has been removed.