Automatic Speech Recognition System Using Word Sequence Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current transcription systems for audio communications often produce inaccurate or delayed transcriptions, especially when relying on human assistance or revoicing, which increases costs and reduces accessibility for individuals who are hard of hearing.

Innovation Solution

The system incorporates a fully automatic speech recognition (ASR) system that selects between different transcription methods and systems during a communication session, fuses multiple transcriptions for improved accuracy, and trains language models using real-time audio data to enhance recognition of frequently occurring word sequences.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If human assistance or revoicing is used for transcription, then transcription accuracy may be improved, but costs increase and accessibility is reduced

Engineering Contradiction:
Improvetranscription accuracyVSAvoidautomation level
Core Design Contradiction:
Measurement precisionVSExtent of automation

Solution Approach 1:

The system enables automatic speech recognition to perform transcription without human intervention. The ASR system processes audio communications autonomously, generating transcriptions that are then processed through word sequence identification and language model training, eliminating the need for human transcribers or revoicing services.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical process of human transcription or revoicing with an automated electronic system. The ASR system, combined with word sequence analysis and language model training, substitutes human cognitive and vocal processes with computational algorithms that automatically generate and improve transcription accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If human assistance is used for transcription, then transcription quality may be improved, but transcription speed and real-time capability are reduced

Engineering Contradiction:
Improvetranscription qualityVSAvoidtranscription speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system performs real-time automatic transcription without human intervention. The ASR system continuously processes audio streams, identifies word sequences, and updates language models dynamically, enabling both high quality and real-time transcription capability simultaneously.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs preliminary processing by identifying and counting word sequences in real-time during communication sessions. This preliminary action accumulates training data continuously, allowing the language model to be refined without waiting for manual transcription completion, thus maintaining both quality and speed.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If multiple transcription methods are used, then transcription accuracy is improved, but system complexity increases

Engineering Contradiction:
Improvetranscription accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges multiple transcription approaches by combining ASR output with language model predictions. The system integrates word sequence identification from audio data with statistical language modeling, creating a unified transcription process that leverages both methods synergistically rather than operating them as separate systems.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The language model serves multiple functions within the system: it provides transcription predictions, identifies likely word sequences, and continuously learns from accumulated data. This multi-functionality reduces the need for separate dedicated components for each task, simplifying the overall system architecture while maintaining improved accuracy.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20200175962A1Training speech recognition systems using word sequences
Publication Date: 2020.06.04 SORENSON IP HOLDINGS LLC
  • US20200175962A1 patent drawing
  • US20200175962A1 patent drawing
  • US20200175962A1 patent drawing

AI summary

A method may include obtaining first audio data of a communication session between a first device and a second device, obtaining a text string that is a transcription of the first audio data, and selecting a contiguous sequence of words from the text string as a first word sequence. The method may further include comparing the first word sequence to multiple word sequences obtained before the communication session and in response to the first word sequence corresponding to one of the multiple word sequences, incrementing a counter of multiple counters associated with the one of the multiple word sequences. The method may also include deleting the text string and the first word sequence and training and after deleting the text string and the first word sequence, training a language model of an automatic transcription system using the multiple word sequences and the multiple counters.