Automatic Speech Recognition System Using Word Sequence Training
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current transcription systems for audio communications often produce inaccurate or delayed transcriptions, especially when relying on human assistance or revoicing, which increases costs and reduces accessibility for individuals who are hard of hearing.
Innovation Solution
The system incorporates a fully automatic speech recognition (ASR) system that selects between different transcription methods and systems during a communication session, fuses multiple transcriptions for improved accuracy, and trains language models using real-time audio data to enhance recognition of frequently occurring word sequences.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If human assistance or revoicing is used for transcription, then transcription accuracy may be improved, but costs increase and accessibility is reduced
Solution Approach 1:
The system enables automatic speech recognition to perform transcription without human intervention. The ASR system processes audio communications autonomously, generating transcriptions that are then processed through word sequence identification and language model training, eliminating the need for human transcribers or revoicing services.
Solution Approach 2:
The patent replaces the mechanical process of human transcription or revoicing with an automated electronic system. The ASR system, combined with word sequence analysis and language model training, substitutes human cognitive and vocal processes with computational algorithms that automatically generate and improve transcription accuracy.
2Measurement precision
If human assistance is used for transcription, then transcription quality may be improved, but transcription speed and real-time capability are reduced
Solution Approach 1:
The system performs real-time automatic transcription without human intervention. The ASR system continuously processes audio streams, identifies word sequences, and updates language models dynamically, enabling both high quality and real-time transcription capability simultaneously.
Solution Approach 2:
The system performs preliminary processing by identifying and counting word sequences in real-time during communication sessions. This preliminary action accumulates training data continuously, allowing the language model to be refined without waiting for manual transcription completion, thus maintaining both quality and speed.
3Measurement precision
If multiple transcription methods are used, then transcription accuracy is improved, but system complexity increases
Solution Approach 1:
The patent merges multiple transcription approaches by combining ASR output with language model predictions. The system integrates word sequence identification from audio data with statistical language modeling, creating a unified transcription process that leverages both methods synergistically rather than operating them as separate systems.
Solution Approach 2:
The language model serves multiple functions within the system: it provides transcription predictions, identifies likely word sequences, and continuously learns from accumulated data. This multi-functionality reduces the need for separate dedicated components for each task, simplifying the overall system architecture while maintaining improved accuracy.
Data Source
AI summary
A method may include obtaining first audio data of a communication session between a first device and a second device, obtaining a text string that is a transcription of the first audio data, and selecting a contiguous sequence of words from the text string as a first word sequence. The method may further include comparing the first word sequence to multiple word sequences obtained before the communication session and in response to the first word sequence corresponding to one of the multiple word sequences, incrementing a counter of multiple counters associated with the one of the multiple word sequences. The method may also include deleting the text string and the first word sequence and training and after deleting the text string and the first word sequence, training a language model of an automatic transcription system using the multiple word sequences and the multiple counters.


