Closed Captioning Timing and Position Adjustment via Audio-Video Context
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audiovisual content systems often fail to align closed captioning properly with the audio and video, leading to timing and duration issues that can disrupt the viewer's experience, particularly in situations where the captioning lags behind or is displayed for too short or too long, and placement can obscure important information or details.
Innovation Solution
The system analyzes audiovisual content to adjust the timing and duration of closed captioning by comparing it with a text version of the audio and determining scene context, using artificial neural networks to identify optimal display positions and modify the captioning accordingly, ensuring it aligns with the audio and video elements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If closed captioning is added as a separate data stream to audiovisual content, then the content can provide text representation for hearing impaired viewers and those in noisy environments, but the closed captioning generally does not properly align with the audio of the content
Solution Approach 1:
The system uses audio-to-text conversion to generate a reference text version of the audio, then compares this reference text with the closed captioning text to detect timing offsets. This feedback mechanism allows the system to identify and correct synchronization issues between the audio and closed captioning, improving alignment accuracy without manual intervention.
Solution Approach 2:
The system performs preliminary analysis by converting audio to text and comparing it with closed captioning before final display. This preliminary action identifies timing and duration issues in advance, allowing the system to pre-adjust the closed captioning parameters to ensure proper synchronization with the audio content.
2Reliability
If closed captioning display duration is shortened to match audio duration, then the captioning aligns better with the audio, but the viewer may not have enough time to read all the text
Solution Approach 1:
The system dynamically adjusts the closed captioning display duration based on the specific audio segment being analyzed. Rather than using a fixed duration, the system calculates the optimal display time for each caption based on the corresponding audio duration and text length, ensuring both synchronization and adequate reading time. This dynamic adjustment resolves the contradiction by adapting to varying content requirements.
3Ease of manufacture
If closed captioning is positioned in a predetermined location, then the display position is simple to implement, but the placement may obscure other information or details being displayed
Solution Approach 1:
The system analyzes the video content to identify regions with important visual information, then positions closed captioning in locations that avoid these critical areas. This local quality approach ensures that caption placement is optimized for each specific scene, preventing obscurement of important video details while maintaining readable caption display.
4Ease of operation
If closed captioning lags behind the audio to allow reading time, then viewers have time to read the text, but the captioning no longer aligns with the corresponding audio
Solution Approach 1:
The system changes multiple parameters simultaneously - adjusting both the timing offset and display duration of closed captioning based on the specific audio segment. By modifying these parameters in coordination, the system maintains proper audio-caption alignment while providing adequate reading time, resolving the contradiction between reading comfort and alignment accuracy.
Data Source
AI summary
Embodiments are directed towards the analysis of the audiovisual content to adjust the timing, duration, or positioning of closed captioning so that the closed captioning more closely aligns with the scene being presented. Content that incudes video, audio, and closed captioning is obtained, and the audio is converted to text. A duration and timing for the closed captioning is determined based on a comparison between the closed captioning and the audio text. Scene context is determined for the content based on analysis of the video and the audio, such as by employing trained artificial neural networks. A display position of the closed captioning is determined based on the scene context. The duration and timing of the closed captioning are modified based on the scene context. The video and closed captioning are provided to a content receiver for presentation to a user based on the display position, duration, and timing.


