Real-Time Audio Transcription Synchronization for Hearing Accessibility

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Voice messages are less useful for hearing-impaired users as they struggle to understand the content by simply listening, and existing transcription services often lack synchronization between audio and text presentation, leading to lagged transcription.

Innovation Solution

A method where a device buffers audio messages and sends them to a transcription system, allowing for real-time transcription and synchronized presentation of both audio and text, ensuring that the text is displayed concurrently with the audio playback.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If voice messages are played back without transcription, then audio playback is simple and fast, but hearing-impaired users cannot understand the content effectively

Engineering Contradiction:
Improveusability for hearing-impaired usersVSAvoidsystem complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

A transcription system is introduced as an intermediary component that converts audio messages into text. This mediator enables hearing-impaired users to access message content through text display while the original audio playback continues, resolving the contradiction between simplicity and accessibility.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system provides multi-functionality by simultaneously supporting both audio playback and text transcription display. This universal approach allows the same system to serve both hearing-capable and hearing-impaired users without requiring separate systems, improving ease of operation while managing complexity through integrated design.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Manufacturing precision

If transcription is provided in real-time, then synchronization with audio is improved, but processing time and system resources increase

Engineering Contradiction:
Improvesynchronization precisionVSAvoidtranscription processing time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system performs preliminary actions by buffering the audio message before transcription begins. This allows the transcription process to start immediately without waiting for the entire audio file to be processed, achieving real-time synchronization while minimizing processing delays through advance preparation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system dynamically adjusts the buffering duration based on network conditions and processing speed. By making the buffering time variable rather than fixed, the system optimizes synchronization precision while adapting to different processing capacities, thereby reducing unnecessary time loss.

Inventive Principle:
Principle #15Dynamics

3Manufacturing precision

If audio is buffered before transcription, then synchronization between text and audio is achieved, but playback delay increases

Engineering Contradiction:
Improvesynchronization accuracyVSAvoidplayback delay
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system applies partial buffering by only buffering the portion of audio that corresponds to the expected transcription delay, rather than buffering the entire message. This excessive action approach ensures sufficient synchronization margin while minimizing unnecessary delay, achieving precise synchronization with minimal playback delay.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11482240B2Presentation of communications
Publication Date: 2022.10.25 SORENSON IP HOLDINGS LLC
  • US11482240B2 patent drawing
  • US11482240B2 patent drawing
  • US11482240B2 patent drawing

AI summary

A method to present communications is provided. The method may include obtaining, at a device, a request from a user to play back a stored message that includes audio. In response to obtaining the request, the method may include directing the audio of the message to a transcription system from the device. In these and other embodiments, the transcription system may be configured to generate text that is a transcription of the audio in real-time. The method may further include obtaining, at the device, the text from the transcription system and presenting, by the device, the text generated by the transcription system in real-time. In response to obtaining the text from the transcription system, the method may also include presenting, by the device, the audio such that the text as presented is substantially aligned with the audio.