Audio Watermarking For Acoustic Echo Cancellation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Speech-enabled devices face challenges in accurately recognizing user speech due to acoustic echoes from synthesized audio, which complicates the speech recognition process.
Innovation Solution
A computer-implemented method that uses an audio watermark encoded in an alignment output audio stream to determine a time alignment value between the output audio stream and the input audio stream captured by a microphone array, allowing for effective cancellation of acoustic echoes using an acoustic echo canceler.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If synthesized audio is played back from an acoustic speaker, then the digital assistant can respond to user queries, but acoustic echo is captured by the microphone array which interferes with speech recognition
Solution Approach 1:
The system performs preliminary actions by outputting an alignment audio stream containing a watermark signal before playing back the synthesized audio response. This watermark is detected in the captured audio to determine time alignment values and generate impulse response characteristics in advance, enabling the acoustic echo canceler to effectively cancel acoustic echo during the actual speech recognition process.
Solution Approach 2:
The patent introduces a watermark signal as an intermediary element. This watermark is embedded in the alignment audio stream, detected in the captured audio, and used to derive time alignment values and impulse response characteristics that mediate between the synthesized audio output and the microphone array input, enabling effective acoustic echo cancellation.
2Measurement precision
If acoustic echo cancellation is performed, then speech recognition accuracy is improved, but time alignment between output and input audio streams must be precisely determined
Solution Approach 1:
The system creates a copy of the alignment audio stream through the acoustic channel by detecting the watermark signal in the captured audio. This copied signal contains the time alignment information needed to synchronize the output and input audio streams, simplifying the time alignment determination process while improving speech recognition accuracy.
Solution Approach 2:
The patent changes the parameter representation by using a watermark signal with specific frequency characteristics (e.g., ultrasonic frequencies above human hearing range) that can be easily detected and processed. This parameter change enables precise time alignment determination through spectral analysis rather than complex temporal correlation methods.
3Reliability
If a watermark signal is used for time alignment, then echo cancellation effectiveness is improved, but the alignment audio stream may become perceptible to users
Solution Approach 1:
The patent applies local quality by using a watermark signal with frequency characteristics localized in a specific range (e.g., ultrasonic frequencies above 20 kHz) that is imperceptible to human ears. This allows the watermark to carry time alignment information effectively while remaining inaudible to users, avoiding the harmful effect of audible interference.
Data Source
AI summary
A method includes receiving an audible response to a query, and prior to playing back the audible response, providing, for output from an acoustic speaker, an alignment output audio stream that encodes an audio watermark. The method also includes receiving an alignment input audio stream captured by a microphone array and encoding an acoustic echo of the audio watermark, processing the alignment input audio stream to detect the acoustic echo, and determining a time alignment value between the alignment output audio stream and the alignment input audio stream. The method also includes playing back a response output audio stream that encodes the audible response and receiving an input audio stream. The input audio stream includes acoustic echo corresponding to the audible response played back. The method also includes processing the input audio stream to generate a respective target audio signal that cancels the acoustic echo.


