ASR-Guided Speech Compression for Selective Audio Retention
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Storing large quantities of speech data for analytics requires heavy compression, which destroys useful information such as speaker identity, emotion, and sentiment, making it difficult to achieve accurate recognition and analysis.
Innovation Solution
Implementing ASR processing and confidence estimation to govern compression levels, where confidently recognized utterances are heavily compressed, and those with low confidence or specific alerts (e.g., angry emotion, fraudster, or child age) remain uncompressed or lightly compressed, allowing for effective text representation and metadata extraction.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If heavy compression is applied to speech data for storage, then storage efficiency is improved, but information quality and recognition accuracy deteriorate
Solution Approach 1:
The patent applies different compression qualities to different segments of speech data based on their importance. Confidently recognized utterances are heavily compressed while uncertain or important utterances are kept in higher quality formats, optimizing the balance between storage efficiency and information retention.
Solution Approach 2:
The compression level is dynamically adjusted based on ASR confidence scores. The system transitions from static uniform compression to dynamic adaptive compression, where each utterance receives appropriate compression based on real-time confidence assessment.
2Device complexity
If uniform compression is applied to all speech data, then processing simplicity is improved, but recognition accuracy for uncertain utterances deteriorates
Solution Approach 1:
The patent changes the compression parameter based on ASR confidence scores. By introducing a confidence-based parameter adjustment mechanism, the system optimizes recognition accuracy for uncertain utterances while maintaining processing efficiency through automated decision-making.
3Loss of information
If all speech data is kept in high quality format, then information retention is improved, but storage costs and processing time deteriorate
Solution Approach 1:
The patent applies high quality retention selectively only to uncertain or important utterances identified by low ASR confidence scores, while heavily compressing confidently recognized portions. This local quality approach optimizes the balance between information retention and storage efficiency.
Solution Approach 2:
Instead of applying high quality retention to all data (excessive action), the patent applies it partially only where necessary based on confidence assessment, reducing overall storage costs while maintaining adequate information quality.
Data Source
AI summary
A process for compressing an audio speech signal utilizes ASR processing to generate a corresponding text representation and, depending on confidence in the corresponding text representation, selectively applies more, less, or no compression to the audio signal. The result is a compressed audio signal, with corresponding text, that is compact and well suited for searching, analytics, or additional ASR processing.


