Audio Content Modification Using Closed Captioning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Individuals with hearing loss face difficulties in distinguishing speech from background noise, especially in environments with multiple sound sources, such as during movies or shows, where the combination of sounds can make it hard to follow the story.

Innovation Solution

A system that analyzes content with both audio and video, using closed captioning data to identify spoken dialogue and modify the audio by reducing or eliminating background noise during spoken segments, enhancing the audio experience for those with hearing impairments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If background noise and multiple sounds are combined in content, then the overall viewing experience is improved for some viewers, but individuals with hearing loss find it difficult to distinguish speech from background noise

Engineering Contradiction:
Improveviewing experienceVSAvoiddifficulty distinguishing speech
Core Design Contradiction:
Adaptability or versatilityVSObject-affected harmful factors

Solution Approach 1:

The audio content is segmented into different components: speech signals and background noise. The system processes these segments separately by identifying speech portions through closed captioning data and applying different processing strategies to speech versus non-speech portions, thereby enabling individuals with hearing loss to distinguish speech from background noise while preserving the overall viewing experience

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system extracts speech signals from the mixed audio content by comparing audio data with closed captioning data. Once speech portions are extracted and identified, they can be enhanced or separated from background noise, allowing hearing-impaired individuals to focus on speech while the background noise remains manageable

Inventive Principle:
Principle #2Taking out (Extraction)

2Measurement precision

If audio processing is applied to isolate speech from background noise, then speech clarity is improved for hearing-impaired individuals, but the complexity of the audio processing system increases

Engineering Contradiction:
Improvespeech clarityVSAvoidaudio processing system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

Closed captioning data serves as an intermediary element that bridges the audio signal and the processing system. By comparing audio data with closed captioning data, the system identifies speech portions without requiring complex speech recognition algorithms, thereby improving speech clarity while minimizing system complexity

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system uses closed captioning data as a copy or representation of the speech content. This copy is then compared with the actual audio signal to identify speech portions, avoiding the need for complex real-time speech analysis and simplifying the overall processing system

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20240395251A1Methods, systems, and apparatuses for modifying audio content
Publication Date: 2024.11.28 COMCAST CABLE COMM LLC
  • US20240395251A1 patent drawing
  • US20240395251A1 patent drawing
  • US20240395251A1 patent drawing

AI summary

Systems, methods, and apparatuses may be provided for modifying audio content. A content item that includes audio content and video content may be received. The content item may include or be associated with closed captioning data or other text data. The text data and the audio content for the content item may be evaluated to determine when, within the audio content, spoken words are occurring. While the spoken words are occurring in the audio content, the audio content may be modified to reduce or eliminate background noise and other sounds within the audio content that occur at or around the time that the spoken words occur within the audio content.