Keyword Detection in Audio Data Using Speech Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for detecting keywords in audio recordings, particularly in low-quality telephone conversations, are inefficient and require manual scanning of hours of recordings, leading to resource wastage and high costs in various industries such as energy trading and national security.
Innovation Solution
A computer-based processing entity with speech recognition software is used to identify potential keyword occurrences in audio data, generating location data for selected subsets of audio to be played to operators for verification, and storing labels indicating the presence or absence of keywords, thereby automating the keyword detection process.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual scanning of audio recordings is used to identify keyword occurrences, then operators can detect keywords with high accuracy, but the process requires extensive time and human resources
Solution Approach 1:
The patent introduces an automatic keyword detection system that acts as an intermediary between the audio recordings and human operators. The system processes audio files, identifies potential keyword occurrences, and presents results to operators for verification, thereby reducing manual scanning time while maintaining detection accuracy through operator review of automated results
Solution Approach 2:
The patent replaces the mechanical process of manual audio scanning with an automated computer-based detection system using speech recognition technology. This substitution dramatically reduces search time from hours of manual listening to rapid automated processing, while the system maintains accuracy through configurable parameters and operator verification capabilities
2Reliability
If manual scanning of audio recordings is used to identify keyword occurrences, then operators can verify keyword presence, but the process consumes excessive human resources and costs
Solution Approach 1:
The automatic detection system serves as an intermediary that handles the bulk of keyword detection work, filtering and identifying potential occurrences. Human operators then verify results rather than performing initial detection, improving resource efficiency while maintaining reliability through this division of labor between automated detection and human verification
Solution Approach 2:
The system enables self-service keyword detection by automatically processing audio recordings and generating detection results without requiring human operators to manually scan each file. This automation improves productivity significantly while maintaining reliability through the system's configurable parameters and optional operator verification
3Productivity
If speech recognition software is used to automatically detect keywords, then processing speed increases, but false positives may occur reducing detection accuracy
Solution Approach 1:
The system implements feedback mechanisms where detection results are reviewed and verified. Operators can review automated detection results, confirm true positives, and correct false positives. This feedback loop continuously improves detection accuracy while maintaining the speed benefits of automated processing
Solution Approach 2:
The system performs partial manual verification rather than complete manual scanning. By automatically processing all recordings and then having operators verify only the detected keyword occurrences rather than listening to entire files, the system achieves high processing speed while maintaining accuracy through targeted verification of potential matches
Data Source
AI summary
Occurrences of one or more keywords in audio data are identified using a speech recognizer employing a language model to derive a transcript of the keywords. The transcript is converted into a phoneme sequence. The phonemes of the phoneme sequence are mapped to the audio data to derive a time-aligned phoneme sequence that is searched for occurrences of keyword phoneme sequences corresponding to the phonemes of the keywords. Searching includes computing a confusion matrix. The language model used by the speech recognizer is adapted to keywords by increasing the likelihoods of the keywords in the language model. For each potential occurrences keywords detected, a corresponding subset of the audio data may be played back to an operator to confirm whether the potential occurrences correspond to actual occurrences of the keywords.


