Speech Recognition Output Augmentation for Untranscribable Utterances
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech recognition systems face high error rates due to untranscribable utterances such as slang, local references, and pseudowords, and lack integration with data platforms for further analysis.
Innovation Solution
An end-to-end computing system converts audio streams into intermediate representations, performs diarization, and transforms segments into object-based representations to decipher untranscribable utterances, integrating with data platforms for enhanced contextualization and analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional speech recognition systems are used, then processing speed is maintained, but accuracy deteriorates due to untranscribable utterances such as slang, local references, and pseudowords
Solution Approach 1:
The patent introduces an intermediary data platform that sits between the audio stream and the speech recognition system. This platform performs preliminary processing including converting audio to intermediate representations (spectrograms), voice activity detection, diarization, and phoneme separation before transcription. This intermediary layer handles untranscribable utterances by creating object-based representations that capture contextual information, thereby improving overall recognition accuracy and reducing error rates.
Solution Approach 2:
The patent segments the audio stream into distinct components: voice activity detection segments, diarization segments (separating different speakers), phoneme separation segments, and transcription segments. Each segment is processed independently and transformed into object-based representations. This segmentation allows the system to handle different types of utterances appropriately and improves accuracy by processing complex audio content in manageable parts.
2Adaptability or versatility
If a single integrated system is implemented to acquire, process, and analyze audio streams, then functionality is improved, but system complexity increases
Solution Approach 1:
The patent creates a universal data platform that performs multiple functions: audio-to-text conversion, object-based representation generation, contextual information extraction, and integration with external systems. This multi-functional platform handles diverse input types (audio streams, video feeds) and produces versatile outputs (transcriptions, object representations, contextual data) that can be used for various downstream applications including medical systems, music analysis, and general speech processing.
3Measurement precision
If untranscribable utterances are deciphered using other instances with similar characteristics, then transcription accuracy is improved, but processing time increases
Solution Approach 1:
The patent performs preliminary processing of audio streams by converting them to intermediate representations, performing voice activity detection, diarization, and phoneme separation before the actual transcription process. This preliminary action creates object-based representations that capture essential characteristics of the audio content, making the subsequent transcription process more efficient and accurate when handling untranscribable utterances by enabling comparison with pre-processed reference instances.
Data Source
AI summary
Computing systems methods, and non-transitory storage media are provided for obtaining an audio stream, converting the audio stream to an intermediate representation, performing diarization on the audio stream, separating the audio stream into individual speech constructs, performing speech recognition on the individual speech constructs by mapping each of the individual speech constructs, or consecutive individual speech constructs, to entries within a dictionary, to generate a transcription of the audio stream, generating an output indicative of the transcription and a result of the diarization, transforming the output into an object-based representation, and performing one or more operations on the object-based representation


