Live Event Audio-Visual Accessibility Processing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional systems face challenges in accurately generating and displaying subtitles and closed captions for live events, particularly in noisy environments and for users with visual or hearing impairments, due to discrepancies in position markers and phonetic variations across regions.
Innovation Solution
An information processing device that captures audio segments from live events, generates text information for song lyrics and non-song phrases in real-time, and adjusts audio characteristics to improve accessibility for visually and hearing-impaired users, using a combination of speech-to-text conversion and user-selectable display and audio settings.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If subtitles are generated beforehand with position markers, then subtitles can be displayed along with video, but the process becomes tiresome and position markers may not match actual dialogues
Solution Approach 1:
The patent replaces manual subtitle generation with automated speech-to-text processing. The system captures audio from the live event, processes it through speech recognition to generate subtitles in real-time, and displays them synchronized with the video stream, eliminating the need for manual creation and positioning of subtitles beforehand
Solution Approach 2:
The system automatically generates and positions subtitles without human intervention. The speech-to-text processing automatically transcribes the audio dialogue, and the system self-synchronizes the subtitle display with the video timing, making the process self-service rather than requiring manual effort
2Reliability
If closed captions are transcribed by human operator, then closed captions can be generated, but accuracy decreases due to phonetic differences across regions
Solution Approach 1:
The patent replaces human operator transcription with automated speech-to-text processing. The system captures audio from the live event and processes it through speech recognition algorithms that can handle phonetic variations across different regions and languages, generating accurate closed captions without human intervention
Solution Approach 2:
The system adapts the speech-to-text processing parameters to handle different phonetic variations. By adjusting the recognition models and processing parameters based on the detected language and phonetic characteristics of the audio, the system maintains high accuracy across different regions and dialects
3Loss of information
If subtitles are displayed along with video, then users can understand dialogues, but visibility remains unclear for visually-impaired users
Solution Approach 1:
The patent segments the accessibility output into multiple channels: visual subtitles for hearing-impaired users and audio descriptions for visually-impaired users. The system separately processes and delivers these types of information through appropriate channels, ensuring both types of users receive the necessary information without interference
Solution Approach 2:
The system introduces audio description as an intermediary layer for visually-impaired users. Instead of relying solely on visual subtitles, the system provides audio-based descriptions of visual elements and actions, serving as a mediator that bridges the gap between visual content and accessibility for users with visual impairments
4Reliability
If audio characteristics are not adjusted, then original audio is preserved, but accessibility for hearing-impaired users is reduced
Solution Approach 1:
The patent implements dynamic audio adjustment where the system can modify audio characteristics in real-time based on user needs. The audio output can dynamically switch between original audio, enhanced audio with clearer phonetics, or synchronized audio with subtitles, providing adaptability while maintaining clarity for hearing-impaired users
Data Source
AI summary
An information processing device includes circuitry that receives a user-input for selection of one of a visual accessibility feature and an aural accessibility feature and further receives a first audio segment from an audio capturing device at a live event. The first audio segment includes a first audio portion of the audio content and a first audio closed caption (CC) information. The circuitry controls display of first text information for the first audio portion and second text information for the first audio CC information, based on received user-input for the selection of the visual accessibility feature. The circuitry generates a second audio segment from the first audio segment based on a first audio characteristic of the first audio segment. The circuitry controls a playback of the generated second audio segment, based on the received user-input for the selection of the aural accessibility feature.


