Silent Speech Recognition via Image Evaluation Means

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional automatic lip-reading methods rely on simultaneous audio and video recordings and require spoken language for training, limiting their applicability, especially for patients who cannot speak due to medical treatments.

Innovation Solution

A method that trains an image evaluation means to generate audio information from silent mouth movements, allowing for reliable lip reading without requiring acoustic input, using a multi-stage approach with image evaluation and speech evaluation means, which can be trained using machine learning and neural networks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional lip-reading methods use simultaneous audio and video recordings with spoken language for training, then speech recognition accuracy is improved, but applicability to patients who cannot speak deteriorates

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidapplicability to non-speaking patients
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system segments the speech recognition process into two independent stages: first generating audio information from visual mouth movements using a trained image evaluation device, then processing this generated audio through a speech evaluation device. This segmentation allows the system to function without requiring actual acoustic speech input, enabling use with non-speaking patients while maintaining recognition accuracy through the two-stage processing approach.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The invention introduces an intermediary functional component (the trained image evaluation device) that translates visual mouth movement data into artificial audio information. This intermediary bridges the gap between visual input and speech recognition processing, enabling the system to handle cases where direct acoustic speech is unavailable, thus expanding adaptability without sacrificing recognition accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Device complexity

If a single-stage neural network directly recognizes speech from video recordings, then system complexity is reduced, but reliability for silent speech recognition deteriorates

Engineering Contradiction:
Improvesystem complexityVSAvoidreliability for silent speech recognition
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The system divides the speech recognition task into two specialized stages: an image evaluation device that generates audio information from visual data, and a speech evaluation device that processes this audio information. This segmentation improves reliability for silent speech by dedicating each stage to its specific function, allowing the image evaluation device to specialize in extracting speech characteristics from mouth movements without the complexity burden of a single monolithic system.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The image evaluation device performs preliminary action by generating artificial audio information from visual mouth movement data before the speech recognition process begins. This preliminary generation of speech-like audio signals enables the subsequent speech evaluation device to process the input using conventional speech recognition algorithms, thereby improving reliability without requiring complete redesign of the entire system.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If the system generates audio information from silent mouth movements, then applicability to intensive care patients is improved, but training data requirements become more stringent

Engineering Contradiction:
Improveapplicability to intensive care patientsVSAvoidtraining data requirements
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

Instead of training the system with traditional audio-visual paired data and then applying it to silent speech, the invention inverts the approach by training the image evaluation device specifically to generate audio information from silent mouth movements. The training uses visual data from non-speaking patients as input and generates corresponding artificial audio as output, thereby expanding adaptability to intensive care patients while managing training data requirements through this inverted training paradigm.

Inventive Principle:
Principle #13The other way round (Inversion)

Data Source

PatentEP3940692B1Method for automatic lip reading using a functional component and providing the functional component
Publication Date: 2023.04.05 CLINOMIC MEDICAL GMBH
  • EP3940692B1 patent drawingFigure 1
  • EP3940692B1 patent drawingFigure 2
  • EP3940692B1 patent drawingFigure 3

AI summary

The invention relates to a method for providing at least one functional component (200) for automatic lip reading, wherein the following steps are carried out: - providing at least one recording (265) comprising audio information (270) about speech of a speaker (1) and image information (280) about a mouth movement of the speaker (1), - performing a training (255) of an image evaluation means (210), wherein the image information (280) is used for an input (201) of the image evaluation means (210) and the audio information (270) is used as a learning specification for an output (202) of the image evaluation means (210) in order to train the image evaluation means (210) to artificially generate the speech during a silent mouth movement.