Direction-Aware Speech-to-Text for Hearing Instrument Comprehension

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Hearing instrument users face challenges in understanding speech due to hearing impairments and disruptive background noise, unclear pronunciation, or unfamiliar accents, which existing technologies partially address but not effectively.

Innovation Solution

The method involves automatically detecting speech in ambient sound, converting it into text data, and outputting this data as graphical or synthesized speech, with direction and speaker traits being determined to enhance comprehension by varying the representation or recording based on the source and speaker characteristics.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If hearing instruments amplify ambient sound to compensate for hearing loss, then hearing ability is improved, but disruptive background noise and speech clarity deteriorate

Engineering Contradiction:
Improvehearing abilityVSAvoiddisruptive background noise
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent segments the audio spectrum into different frequency bands and applies selective processing to each band. The speech enhancement algorithm identifies and enhances speech-containing frequency regions while suppressing background noise in other regions, thereby improving speech clarity without excessive amplification of disruptive noises.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediate processing stage between sound capture and output, where speech enhancement algorithms analyze the ambient sound and generate enhanced speech signals. This intermediary processing layer selectively reinforces speech components while attenuating background noise, resolving the contradiction between overall amplification and speech clarity.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If speech is converted into text data and output graphically, then comprehension support is improved, but device complexity increases

Engineering Contradiction:
Improvespeech comprehension supportVSAvoidprocessing and output system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent introduces an external computing device as an intermediary to handle the complex speech-to-text conversion and graphical output functions. The hearing instrument communicates with this external device, which performs the computationally intensive speech recognition and text generation, thereby providing enhanced comprehension support without significantly increasing the complexity of the hearing instrument itself.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent leverages the multi-functionality of external computing devices (smartphones, tablets, computers) that can perform speech recognition, text generation, and graphical display functions. By utilizing these existing universal devices, the system avoids duplicating complex functionality within the hearing instrument, thus reducing device complexity while maintaining enhanced speech comprehension support.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If multiple speech sources are captured and processed, then speech source differentiation is improved, but processing time and computational load increase

Engineering Contradiction:
Improvespeech source differentiationVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent segments the speech processing task by assigning different processing responsibilities to the hearing instrument and external computing devices. The hearing instrument performs real-time spatial analysis and speech source identification using its microphones and signal processing capabilities, while less time-critical text conversion and display are handled externally, thereby achieving good speech source differentiation with minimized processing time.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary speech source identification and spatial analysis at the hearing instrument level before transmitting data to external devices. By pre-processing and identifying speech sources in real-time, the system reduces the computational load and processing time required for subsequent text generation and display, maintaining rapid response to multiple speech sources.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250308532A1Method for supporting the hearing comprehension of a hearing instrument user and hearing system with a hearing instrument
Publication Date: 2025.10.02 SIVANTOS PTE LTD
  • US20250308532A1 patent drawing
  • US20250308532A1 patent drawing
  • US20250308532A1 patent drawing

AI summary

A method for supporting hearing comprehension of a hearing instrument user includes using the hearing instrument to capture speech-containing ambient sound from surroundings. The speech is automatically converted into text data output to the user as a graphical representation of text on a screen of the hearing instrument or peripheral device connected thereto for data transmission and/or as synthesized speech as a sound signal. A direction of origin and/or at least one speaker trait for the speech are/is determined automatically and resolved relative to time. The graphical representation of the text data and synthesized speech vary based on the identified direction of origin and/or speaker trait in a manner resolved relative to time. Additionally or alternatively to immediate output to the user, the graphical representation of the text data and synthesized speech are recorded for later output. A hearing system is also provided.