Audio Processing Matching Sound Units to Stop Playback
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speech dialogue systems face challenges in accurately detecting speech interruptions due to incomplete echo cancellation, low accuracy in speech activity detection, false recognition, and interference from environmental noise and unrelated sounds.
Innovation Solution
An audio processing method that receives an external input sound message during a first audio message playback, matches it with a pre-set receiving message associated with the audio message content, and stops playing the audio when the matching result meets a threshold, improving recognition accuracy and efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If AEC algorithm is used for echo cancellation, then echo removal is achieved, but prompt tone cannot be completely eliminated and remains in the output signal
Solution Approach 1:
The patent introduces a prompt tone database as an intermediary component that stores pre-collected prompt tone samples. During speech interruption detection, the system retrieves and compares these stored prompt tone samples with the current audio signal to identify and remove residual prompt tones that the AEC algorithm failed to eliminate, thereby improving both echo cancellation effectiveness and prompt tone elimination accuracy
Solution Approach 2:
The system implements a feedback mechanism where the output signal from AEC is continuously monitored and compared against the prompt tone database. When residual prompt tones are detected through comparison, the system feeds this information back to enhance the echo cancellation process, creating a closed-loop system that iteratively improves prompt tone removal while preserving user speech
2Extent of automation
If VAD or ASR technology is used for speech interruption detection, then speech detection is achieved, but accuracy is not high enough and false detection occurs under environmental noise interference
Solution Approach 1:
The patent segments the speech detection process into multiple independent stages: first using VAD for voice activity detection, then using ASR for speech recognition, and finally using prompt tone database comparison for verification. Each stage operates independently and contributes to the overall detection accuracy, reducing false detections by distributing the detection task across multiple specialized components rather than relying on a single automated system
Solution Approach 2:
The prompt tone database serves as an intermediary verification layer between the automated detection systems and the final speech interruption determination. By comparing detected signals against stored prompt tone samples, the system mediates between automated detection results and actual speech interruption events, filtering out false positives caused by environmental noise
3Adaptability or versatility
If speech recognition module is used for speech interruption detection, then speech recognition is achieved, but false recognition causes false detection of speech interruption event
Solution Approach 1:
The prompt tone database acts as an intermediary verification mechanism that checks the output of the speech recognition module. By comparing recognized speech patterns against stored prompt tone samples, the system identifies and corrects false recognitions before they result in false speech interruption detections, thereby maintaining the adaptability of speech recognition while improving detection reliability
Solution Approach 2:
The system implements feedback from the prompt tone database comparison results back to the speech recognition module. When false recognition is detected through database comparison, the feedback mechanism adjusts or corrects the recognition output, preventing false speech interruption events while preserving the module's ability to recognize valid user speech
4Ease of operation
If conventional speech interruption module is used, then speech interruption detection is achieved, but unrelated sounds from user cause speech interruption event
Solution Approach 1:
The prompt tone database serves as an intermediary filter between the microphone input and the speech interruption determination. By comparing all detected sounds against stored prompt tone samples, the system mediates between ease of operation (maintaining speech interruption functionality) and reliability (filtering out unrelated sounds like coughing or talking to others), only triggering interruption when a match is found
Data Source
AI summary
Systems and methods are provided for improving audio processing by receiving an external input sound message during playing a first audio message; matching the external input sound message with a receiving message to obtain a matching result, wherein the receiving message is associated with the first audio message in content, wherein the matching is based on a proportion of sound units in the sound message that hit sound units in the receiving message; determining whether the matching result meets a threshold; and upon determining that the matching result meets the threshold, stop playing the first audio message.


