Live Event Audio-Visual Accessibility Processing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional systems face challenges in accurately generating and displaying subtitles and closed captions for live events, particularly in noisy environments and for users with visual or hearing impairments, due to discrepancies in position markers and phonetic variations across regions.

Innovation Solution

An information processing device that captures audio segments from live events, generates text information for song lyrics and non-song phrases in real-time, and adjusts audio characteristics to improve accessibility for visually and hearing-impaired users, using a combination of speech-to-text conversion and user-selectable display and audio settings.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If subtitles are generated beforehand with position markers, then subtitles can be displayed along with video, but the process becomes tiresome and position markers may not match actual dialogues

Engineering Contradiction:
Improvesubtitle position accuracyVSAvoidsubtitle generation effort
Core Design Contradiction:
Measurement precisionVSEase of manufacture

Solution Approach 1:

The patent replaces manual subtitle generation with automated speech-to-text processing. The system captures audio from the live event, processes it through speech recognition to generate subtitles in real-time, and displays them synchronized with the video stream, eliminating the need for manual creation and positioning of subtitles beforehand

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system automatically generates and positions subtitles without human intervention. The speech-to-text processing automatically transcribes the audio dialogue, and the system self-synchronizes the subtitle display with the video timing, making the process self-service rather than requiring manual effort

Inventive Principle:
Principle #25Self-service

2Reliability

If closed captions are transcribed by human operator, then closed captions can be generated, but accuracy decreases due to phonetic differences across regions

Engineering Contradiction:
Improveclosed caption accuracyVSAvoidtranscription process complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent replaces human operator transcription with automated speech-to-text processing. The system captures audio from the live event and processes it through speech recognition algorithms that can handle phonetic variations across different regions and languages, generating accurate closed captions without human intervention

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system adapts the speech-to-text processing parameters to handle different phonetic variations. By adjusting the recognition models and processing parameters based on the detected language and phonetic characteristics of the audio, the system maintains high accuracy across different regions and dialects

Inventive Principle:
Principle #35Parameter changes

3Loss of information

If subtitles are displayed along with video, then users can understand dialogues, but visibility remains unclear for visually-impaired users

Engineering Contradiction:
Improveinformation accessibilityVSAvoidvisual impairment barrier
Core Design Contradiction:
Loss of informationVSObject-affected harmful factors

Solution Approach 1:

The patent segments the accessibility output into multiple channels: visual subtitles for hearing-impaired users and audio descriptions for visually-impaired users. The system separately processes and delivers these types of information through appropriate channels, ensuring both types of users receive the necessary information without interference

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system introduces audio description as an intermediary layer for visually-impaired users. Instead of relying solely on visual subtitles, the system provides audio-based descriptions of visual elements and actions, serving as a mediator that bridges the gap between visual content and accessibility for users with visual impairments

Inventive Principle:
Principle #24Intermediary (Mediator)

4Reliability

If audio characteristics are not adjusted, then original audio is preserved, but accessibility for hearing-impaired users is reduced

Engineering Contradiction:
Improveaudio clarity for hearing-impairedVSAvoidaudio format flexibility
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent implements dynamic audio adjustment where the system can modify audio characteristics in real-time based on user needs. The audio output can dynamically switch between original audio, enhanced audio with clearer phonetics, or synchronized audio with subtitles, providing adaptability while maintaining clarity for hearing-impaired users

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS11211074B2Presentation of audio and visual content at live events based on user accessibility
Publication Date: 2021.12.28 SONY GROUP CORP
  • US11211074B2 patent drawing
  • US11211074B2 patent drawing
  • US11211074B2 patent drawing

AI summary

An information processing device includes circuitry that receives a user-input for selection of one of a visual accessibility feature and an aural accessibility feature and further receives a first audio segment from an audio capturing device at a live event. The first audio segment includes a first audio portion of the audio content and a first audio closed caption (CC) information. The circuitry controls display of first text information for the first audio portion and second text information for the first audio CC information, based on received user-input for the selection of the visual accessibility feature. The circuitry generates a second audio segment from the first audio segment based on a first audio characteristic of the first audio segment. The circuitry controls a playback of the generated second audio segment, based on the received user-input for the selection of the aural accessibility feature.