Conditional Acoustic Model Sound Embedding for Speech Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speech recognition systems struggle with accuracy in uncommon accents, voice types, and environmental conditions such as noise, background voices, or music, limiting their usefulness in daily life.

Innovation Solution

The use of conditional acoustic models that condition on sound embeddings, which encode features from a key phrase to improve the accuracy of phoneme recognition in subsequent speech segments, even in challenging conditions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional speech recognition systems are used, then the system is simple and easy to operate, but the recognition accuracy deteriorates in uncommon accents, voice types, and environmental conditions

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidadaptability to uncommon accents and environmental conditions
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system performs preliminary action by encoding sound embeddings from a key phrase before processing the subsequent speech segment. This pre-computed embedding captures acoustic characteristics that are then used to condition the acoustic model, improving recognition accuracy for uncommon accents and environmental conditions without requiring complex real-time adaptations.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The invention changes parameters by introducing sound embeddings as conditional inputs to the acoustic model. Instead of using a fixed acoustic model, the system dynamically adjusts the model's behavior by conditioning it on the extracted sound embedding features, allowing adaptability to different speakers and environments while maintaining system simplicity.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If conditional acoustic models with sound embeddings are used, then speech recognition accuracy improves in diverse scenarios, but the device complexity increases

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidcomplexity of acoustic model processing
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system segments the speech processing into distinct components: key phrase detection, sound embedding extraction, and conditional acoustic model processing. This segmentation allows each component to be optimized independently and enables modular implementation, reducing overall system complexity while maintaining high recognition accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The sound embedding acts as an intermediary between the raw audio input and the acoustic model. This intermediate representation captures essential acoustic characteristics in a compact form, simplifying the processing required by the acoustic model while improving its ability to handle diverse speech conditions.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If conventional acoustic models are used, then the processing is fast and simple, but the recognition fails in noisy environments with background voices or music

Engineering Contradiction:
Improvereliability of speech recognition in noisy conditionsVSAvoidcomplexity of noise handling
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system performs preliminary noise characterization by extracting sound embeddings from the key phrase that contains information about the acoustic environment, including background noise, music, or other voices. This pre-computed embedding is then used to condition the acoustic model, enabling it to adapt to noisy conditions without requiring complex real-time noise processing.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentEP3905237B1Acoustic model conditioning on sound features
Publication Date: 2025.06.18 SOUNDHOUND INC
  • EP3905237B1 patent drawingFigure 1
  • EP3905237B1 patent drawingFigure 2
  • EP3905237B1 patent drawingFigure 3

AI summary

Systems and methods of speech recognition capture segments of speech audio having a key phrase shortly followed by an utterance. An encoder uses the key phrase segment to compute a sound embedding, which is stored. An acoustic model for speech recognition infers phonemes from the utterance audio signal using a model that is conditioned on the sound embedding as an input. The sound embedding may be held until another key phrase is captured or a session ends. The acoustic model and encoder may be jointly trained from speech data recordings that may be mixed with noise, the profile of mixed noise being the same for the key phrase segment and the utterance segment.