Context Acoustic Biasing Engine for Speech Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing automated speech recognition (ASR) systems face limitations in accuracy due to the reliance on small, less than optimal, training data sets, which result in misidentification of spoken words, especially those with multiple possible textual representations.

Innovation Solution

The implementation of a context acoustic biasing (CAB) engine that utilizes historical textual content and a domain-specific ontology to generate contextual term lists, which are then converted into acoustic representations and used to bias the ASR computer model's predictions, thereby improving accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of time

If ASR computer models are trained on smaller training data sets to reduce time and resource requirements, then training efficiency is improved, but recognition accuracy deteriorates

Engineering Contradiction:
Improvetraining timeVSAvoidrecognition accuracy
Core Design Contradiction:
Loss of timeVSMeasurement precision

Solution Approach 1:

The system performs preliminary actions by generating acoustic representations of concept terms from ontology data structures before runtime speech recognition. These pre-generated acoustic representations are stored and readily available to bias the ASR model predictions, eliminating the need for extensive training while improving accuracy at inference time.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary mechanism (context acoustic biasing engine) that bridges the gap between limited training data and high accuracy requirements. This engine uses ontology data structures to generate contextual acoustic representations that mediate the ASR prediction process, allowing accurate recognition without extensive training datasets.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of energy

If ASR systems use traditional training methods with limited data, then resource consumption is reduced, but misidentification of spoken words increases

Engineering Contradiction:
Improvecomputational resourcesVSAvoidword identification reliability
Core Design Contradiction:
Loss of energyVSReliability

Solution Approach 1:

The system employs self-service by using ontology data structures to automatically generate contextual acoustic representations without requiring external training data. The ASR model leverages these self-generated representations to improve its predictions, achieving high reliability without consuming additional computational resources for data collection and labeling.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent changes the parameters of the ASR system by introducing contextual acoustic representations derived from ontology data structures. This parameter change allows the system to achieve high word identification reliability without increasing computational resource consumption, as the biasing mechanism operates efficiently during inference.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If ASR models are trained extensively to improve accuracy, then recognition precision is improved, but training complexity increases

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidtraining process complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system extracts the essential contextual information from ontology data structures and uses it to generate acoustic representations that bias ASR predictions. This extraction approach achieves high recognition accuracy without the complexity of extensive training processes, as the contextual biasing is derived directly from structured knowledge representations.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent replaces the mechanical training process with a knowledge-based substitution approach. Instead of relying on iterative training with large datasets, the system uses ontology data structures to generate contextual acoustic representations that directly bias predictions, eliminating training complexity while maintaining high accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS12315496B2Ontology driven contextual automated speech recognition
Publication Date: 2025.05.27 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US12315496B2 patent drawing
  • US12315496B2 patent drawing
  • US12315496B2 patent drawing

AI summary

An automatic speech recognition (ASR) computing system and methodology are provided to predict a textual representation of received input speech data. A context acoustic biasing (CAB) engine of the ASR computing system receives historical textual content and an ontology data structure. The CAB engine matches key terms identified in the historical textual content with concepts present in the ontology data structure to generate a contextual term list data structure comprising the concept terms related to concepts matching the key terms. The CAB engine generates acoustic representations of the concept terms in the contextual term list data structure and inputs them to an ASR computer model of the ASR computing system which processes an input speech signal to generate a predicted textual representation of the input speech signal. The predicted textual representation is biased towards the acoustic representations of the concept terms in the contextual term list data structure.