Direction-Aware Speech Transcription for Hearing Instruments

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Hearing instrument users face challenges in understanding speech in noisy environments or with unclear pronunciation, especially due to background noise and unfamiliar accents, which existing technologies often fail to adequately address.

Innovation Solution

The method involves converting ambient sound into text data, which is then output as graphical or synthesized speech, with direction and speaker characteristics being determined in real-time to enhance comprehension. This includes varying the output based on the speech's direction and speaker characteristics, using adaptive beamforming and speech recognition technology.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If ambient sound is processed and converted into text data with direction and speaker characteristics, then speech comprehension is enhanced, but device complexity increases

Engineering Contradiction:
Improvespeech comprehensionVSAvoiddevice complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent introduces an external computing device (smartphone, tablet, or cloud server) as an intermediary to perform the complex speech-to-text conversion and analysis tasks. The hearing instrument captures ambient sound and transmits it to the external device, which then processes the audio data to generate text output with direction and speaker characteristics. This mediator approach allows the hearing instrument to benefit from advanced processing capabilities without bearing the full complexity burden.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system divides the speech comprehension support function into separate modules: the hearing instrument handles audio capture and basic signal conditioning, while the external computing device performs speech recognition, text generation, and speaker characterization. This segmentation allows each component to be optimized independently, reducing the complexity of any single device while maintaining overall system effectiveness.

Inventive Principle:
Principle #1Segmentation

2Reliability

If speech is converted into text data and output with direction information, then understanding in noisy environments improves, but loss of time increases due to processing

Engineering Contradiction:
Improvespeech comprehensionVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary actions by continuously monitoring and pre-processing ambient sound even before speech comprehension is explicitly needed. The external computing device maintains ready-state speech recognition models and can quickly generate text output when speech is detected, reducing the effective processing time during actual use.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements continuous audio processing and speech recognition rather than intermittent processing. The system continuously analyzes the audio stream, maintaining an ongoing transcription and speaker tracking process, which eliminates start-up delays and provides near-real-time text output with minimal perceptible lag.

Inventive Principle:
Principle #20Continuity of useful action

3Reliability

If multiple speakers are tracked with direction and characteristics, then comprehension of multiple speakers improves, but device complexity increases

Engineering Contradiction:
Improvecomprehension of multiple speakersVSAvoiddevice complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent uses an external computing device as a mediator to handle the complex task of tracking multiple speakers with direction and characteristics. The hearing instrument captures the audio and passes it to the external device, which employs advanced speech separation and speaker diarization algorithms to identify and track multiple speakers independently, assigning direction and characteristics to each speaker's speech.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system creates separate text copies for each speaker with associated metadata (direction, speaker characteristics) rather than mixing all speech into a single transcript. This copying approach allows the system to maintain distinct representations of multiple speakers, making it easier to track and differentiate them without requiring the entire system to simultaneously process all speaker attributes in a single complex structure.

Inventive Principle:
Principle #26Copying

Data Source

PatentEP4626038A1Supporting the hearing comprehension of a hearing instrument user
Publication Date: 2025.10.01 SIVANTOS PTE LTD
  • EP4626038A1 patent drawingFigure 1
  • EP4626038A1 patent drawingFigure 2
  • EP4626038A1 patent drawingFigure 3

AI summary

A method for supporting the auditory comprehension of a hearing instrument user and an associated hearing system (2) with such a hearing instrument (2) are specified. By means of the hearing instrument (4), ambient sound containing speech is recorded from the user's surroundings. The speech contained in the recorded ambient sound is automatically converted into text data (T). This text data (T) is output to the user as a graphic representation (G) of text on a screen (40) of the hearing instrument (4) or of a peripheral device (22) connected thereto by data transmission technology and/or as synthesized speech (L) in the form of a sound signal. For the speech contained in the ambient sound, a direction of origin (R) and/or at least one speaker characteristic (S) is determined automatically and in a time-resolved manner. The graphic representation (G) of the text data (T) orThe synthesized speech (L) is varied in a time-resolved manner depending on the determined direction of origin (R) and/or the at least one speaker characteristic (S). In addition to or as an alternative to the immediate output to the user, the graphical representation (G) of the text data (T) or the synthesized speech (L) is recorded for later output.