Speech Recognition Training With Hybrid ASR Transcription Fusion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current audio transcription systems for hard-of-hearing individuals suffer from inaccuracies and delays due to reliance on human revoicing and speaker-dependent automatic speech recognition (ASR) systems, leading to increased costs and reduced effectiveness.

Innovation Solution

Implementing a hybrid system that combines speaker-dependent and speaker-independent ASR systems with real-time transcription fusion and human revoicing, utilizing revoiced and regular audio to generate accurate and synchronized transcriptions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If human revoicing is used for transcription, then transcription accuracy may be improved, but cost and time delay increase

Engineering Contradiction:
Improvetranscription accuracyVSAvoidtranscription delay
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent segments the transcription process into multiple parallel ASR systems (speaker-dependent and speaker-independent) that operate simultaneously on different audio inputs (regular audio and revoiced audio). This segmentation allows the system to process multiple transcription paths in parallel, reducing overall delay while maintaining accuracy through fusion of results.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent merges outputs from multiple ASR systems and audio sources through a fusion process. By combining transcriptions from speaker-dependent ASR, speaker-independent ASR, revoiced audio processing, and regular audio processing, the system achieves higher accuracy than any single method alone while reducing reliance on sequential human review.

Inventive Principle:
Principle #5Merging (Combining)

2Measurement precision

If speaker-dependent ASR systems are used, then transcription accuracy improves, but system complexity and cost increase

Engineering Contradiction:
Improvetranscription accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent implements a universal transcription system that handles multiple speaker types and communication scenarios through a single integrated architecture. The system processes both speaker-dependent and speaker-independent ASR inputs, along with revoiced and regular audio, making it adaptable to various users and situations without requiring separate specialized systems for each case.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent introduces a fusion module as an intermediary that receives and integrates outputs from multiple ASR systems. This mediator component simplifies the overall system architecture by providing a centralized point for combining results, managing complexity, and producing a unified transcription output without requiring direct integration between all individual ASR components.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If multiple ASR systems are implemented with real-time fusion, then transcription accuracy and speed improve, but computational resources and system complexity increase

Engineering Contradiction:
Improvetranscription speedVSAvoidcomputational resources
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The patent performs preliminary processing of audio signals through revoicing before they reach the ASR systems. By pre-processing audio to improve speech clarity and consistency, the system reduces the computational burden on subsequent ASR processing stages, allowing faster and more accurate transcription with reduced overall computational requirements.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12380877B2Training of speech recognition systems
Publication Date: 2025.08.05 SORENSON IP HOLDINGS LLC
  • US12380877B2 patent drawing
  • US12380877B2 patent drawing
  • US12380877B2 patent drawing

AI summary

A method may include obtaining first audio data of a first communication session between a first and second device and during the first communication session, obtaining a first text string that is a transcription of the first audio data and training a model of an automatic speech recognition system using the first text string and the first audio data. The method may further include in response to completion of the training, deleting the first audio data and the first text string and after deleting the first audio data and the first text string, obtaining second audio data of a second communication session between a third and fourth device and during the second communication session obtaining a second text string that is a transcription of the second audio data and further training the model of the automatic speech recognition system using the second text string and the second audio data.