Audio Processing Matching Sound Units to Stop Playback

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speech dialogue systems face challenges in accurately detecting speech interruptions due to incomplete echo cancellation, low accuracy in speech activity detection, false recognition, and interference from environmental noise and unrelated sounds.

Innovation Solution

An audio processing method that receives an external input sound message during a first audio message playback, matches it with a pre-set receiving message associated with the audio message content, and stops playing the audio when the matching result meets a threshold, improving recognition accuracy and efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If AEC algorithm is used for echo cancellation, then echo removal is achieved, but prompt tone cannot be completely eliminated and remains in the output signal

Engineering Contradiction:
Improveecho cancellation effectivenessVSAvoidprompt tone elimination accuracy
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The patent introduces a prompt tone database as an intermediary component that stores pre-collected prompt tone samples. During speech interruption detection, the system retrieves and compares these stored prompt tone samples with the current audio signal to identify and remove residual prompt tones that the AEC algorithm failed to eliminate, thereby improving both echo cancellation effectiveness and prompt tone elimination accuracy

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system implements a feedback mechanism where the output signal from AEC is continuously monitored and compared against the prompt tone database. When residual prompt tones are detected through comparison, the system feeds this information back to enhance the echo cancellation process, creating a closed-loop system that iteratively improves prompt tone removal while preserving user speech

Inventive Principle:
Principle #23Feedback

2Extent of automation

If VAD or ASR technology is used for speech interruption detection, then speech detection is achieved, but accuracy is not high enough and false detection occurs under environmental noise interference

Engineering Contradiction:
Improvespeech interruption detection automationVSAvoidspeech detection accuracy
Core Design Contradiction:
Extent of automationVSMeasurement precision

Solution Approach 1:

The patent segments the speech detection process into multiple independent stages: first using VAD for voice activity detection, then using ASR for speech recognition, and finally using prompt tone database comparison for verification. Each stage operates independently and contributes to the overall detection accuracy, reducing false detections by distributing the detection task across multiple specialized components rather than relying on a single automated system

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The prompt tone database serves as an intermediary verification layer between the automated detection systems and the final speech interruption determination. By comparing detected signals against stored prompt tone samples, the system mediates between automated detection results and actual speech interruption events, filtering out false positives caused by environmental noise

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If speech recognition module is used for speech interruption detection, then speech recognition is achieved, but false recognition causes false detection of speech interruption event

Engineering Contradiction:
Improvespeech recognition capabilityVSAvoidspeech interruption detection reliability
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The prompt tone database acts as an intermediary verification mechanism that checks the output of the speech recognition module. By comparing recognized speech patterns against stored prompt tone samples, the system identifies and corrects false recognitions before they result in false speech interruption detections, thereby maintaining the adaptability of speech recognition while improving detection reliability

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system implements feedback from the prompt tone database comparison results back to the speech recognition module. When false recognition is detected through database comparison, the feedback mechanism adjusts or corrects the recognition output, preventing false speech interruption events while preserving the module's ability to recognize valid user speech

Inventive Principle:
Principle #23Feedback

4Ease of operation

If conventional speech interruption module is used, then speech interruption detection is achieved, but unrelated sounds from user cause speech interruption event

Engineering Contradiction:
Improvespeech interruption functionalityVSAvoidspeech interruption accuracy
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The prompt tone database serves as an intermediary filter between the microphone input and the speech interruption determination. By comparing all detected sounds against stored prompt tone samples, the system mediates between ease of operation (maintaining speech interruption functionality) and reliability (filtering out unrelated sounds like coughing or talking to others), only triggering interruption when a match is found

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11133009B2Method, apparatus, and terminal device for audio processing based on a matching of a proportion of sound units in an input message with corresponding sound units in a database
Publication Date: 2021.09.28 ALIBABA GROUP HOLDING LTD
  • US11133009B2 patent drawing
  • US11133009B2 patent drawing
  • US11133009B2 patent drawing

AI summary

Systems and methods are provided for improving audio processing by receiving an external input sound message during playing a first audio message; matching the external input sound message with a receiving message to obtain a matching result, wherein the receiving message is associated with the first audio message in content, wherein the matching is based on a proportion of sound units in the sound message that hit sound units in the receiving message; determining whether the matching result meets a threshold; and upon determining that the matching result meets the threshold, stop playing the first audio message.