Hearing Device Speech Recognition with Multimodal Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing hearing aids, particularly for individuals with profound hearing loss, struggle to provide clear and accurate speech recognition due to insufficient signal-to-noise ratios, and existing technologies rely solely on microphone input, which is inadequate for reliable word identification.

Innovation Solution

A hearing assistive device utilizing multiple data inputs, including microphones, lip reading, and speech recognition engines, with confidence levels and weighting factors to enhance word decipherment, and incorporating thermal sensors, context engines, and facial recognition to improve accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If multiple data inputs (microphone, lip reading, thermal sensors, context engines, facial recognition) are combined to improve speech recognition accuracy, then reliability and measurement precision improve, but device complexity increases

Engineering Contradiction:
Improvespeech recognition reliabilityVSAvoiddevice complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent combines multiple data input sources (microphone, camera for lip reading, thermal sensors, context engines, facial recognition systems) into a unified speech recognition system. These diverse inputs are merged through a processing system that integrates their outputs to determine the most likely spoken words, thereby improving reliability through redundancy and cross-validation of multiple sensing modalities.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The hearing device is designed with multi-functional capabilities, serving both as a traditional hearing aid and as a sophisticated speech recognition system. The device universally processes multiple types of data (acoustic, visual, thermal, contextual) to achieve accurate word identification, making it adaptable to various listening conditions and environments while maintaining a single integrated system architecture.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If multiple data inputs and processing engines are used to enhance speech recognition, then measurement precision improves, but the device complexity and processing requirements increase

Engineering Contradiction:
Improveword identification accuracyVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The speech recognition system is segmented into distinct functional modules: microphone input processing, camera-based lip reading processing, thermal sensor processing, context engine analysis, and facial recognition processing. Each module independently processes its specific data type and contributes to the overall word identification, allowing complex processing to be divided into manageable segments that can be optimized individually.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces intermediary processing layers that translate different data types into a common format for integration. The context engine acts as an intermediary that synthesizes information from multiple sources and applies linguistic knowledge to resolve ambiguities. Facial recognition and lip reading serve as intermediaries between visual cues and speech content, bridging different sensing modalities through a unified processing framework.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12424204B1Speech recognition hearing device with multiple supportive detection inputs
Publication Date: 2025.09.23 GN HEARING AS
  • US12424204B1 patent drawing
  • US12424204B1 patent drawing
  • US12424204B1 patent drawing

AI summary

A device and method for improving hearing devices by using computer speech recognition of words in the first instance, and then other elements of data acquisition to increase the confidence level that words recognized are accurately identified. If a required predetermined confidence level is reached or the engine(s) with the highest confidence is determined, the recognized/deciphered word from that engine(s) is determined to be the most likely deciphered word and is sent to the device users via a synthesized voice or actual voice. The system adds additional elements of data acquisition including: a) context analysis, b) facial recognition c) lip reading recognition, d) reflected sound wave recognition, e) thermal imaging of mouth expelled air movement to increase the confidence level of lip reading recognition and f) other means. Other modes include translation of foreign languages into a user's ear and using a heads up display to project the text version of words which the computer had deciphered or translated. The system may be triggered by eye moment, spoken command, hand movement or similar.