STFT-based echomutator

The STFT-based echo mutator system addresses the challenge of acoustic echoes in speech-enabled devices by using an acoustic echo canceller and double-talk detector to enhance speech recognition accuracy in double-talk scenarios.

JP7764589B2Active Publication Date: 2025-11-05GOOGLE LLC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024516942
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-09-16
Filing Date
2021-12-11
Publication Date
2025-11-05
Estimated Expiration
2041-12-11

AI Technical Summary

Technical Problem

Acoustic echoes generated by synthesized playback audio in speech-enabled devices interfere with speech recognition systems, making it difficult to accurately process user speech during double-talk events.

Method used

An STFT-based echo mutator system that includes an acoustic echo canceller and a double-talk detector to identify and mute echo-only frames while allowing double-talk frames to be processed, using cross-correlation indicators to determine frame types.

Benefits of technology

Enhances speech recognition accuracy by effectively distinguishing and removing acoustic echoes, improving the performance of speech recognition systems in environments with overlapping speech.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007764589000007
    Figure 0007764589000007
  • Figure 0007764589000008
    Figure 0007764589000008
  • Figure 0007764589000009
    Figure 0007764589000009
Patent Text Reader

Abstract

A method (300) for short-time Fourier transform based echo muting includes receiving a microphone signal (202) including an acoustic echo (156) captured by a microphone and corresponding to audio content (154) from an acoustic speaker (118) and receiving a reference signal (158) including a sequence of frames representing the audio content. For each frame, the method includes processing using an acoustic echo canceller (210) configured to receive the respective frame as input to generate a respective output signal frame (206) that removes the acoustic echo from the respective frame, and determining using a double-talk detector (220) based on the respective frame and the output signal frame whether the respective frame includes a double-talk frame or an echo-only frame. The method further includes muting the respective output signal frame for each respective frame that includes an echo-only frame, and performing audio processing on the respective output signal frame for each respective frame that includes a double-talk frame.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to an echo mutator based on the short-time Fourier transform. [Background technology]

[0002] A speech-enabled device can generate synthesized playback audio and communicate the synthesized playback audio to one or more users in a speech environment. While the speech-enabled device outputs the synthesized playback audio, a microphone of the speech-enabled device may capture the synthesized playback audio as an acoustic echo while actively capturing speech spoken by a user directed at the speech-enabled device. Unfortunately, the acoustic echo resulting from the synthesized playback audio can make it difficult for a speech recognizer to recognize speech spoken by a user that occurs during the echo from the synthesized playback audio. Summary of the Invention [Means for solving the problem]

[0003] One aspect of the present disclosure provides a computer-implemented method for performing speech recognition using an STFT-based echo mutator. When executed on data processing hardware, the computer-implemented method causes the data processing hardware to perform operations including receiving a microphone signal including an acoustic echo captured by a microphone. The acoustic echo corresponds to audio content played from an acoustic speaker. The operations also include receiving a reference signal including a sequence of frames representing audio content sent on a reference channel before the audio content is played by the acoustic speaker. For each frame in the sequence of frames of the microphone signal, the operations also include processing each frame of the microphone signal using an acoustic echo canceller configured to receive as input each frame in the sequence of frames of the reference signal to generate a respective output signal frame that removes the acoustic echo from the respective frame of the microphone signal. The operations also include using a double-talk detector (DTD) to determine whether each frame of the microphone signal includes a double-talk frame or an echo-only frame based on the respective frame of the reference signal and the respective output signal frame. For each respective frame in the sequence of frames of the microphone signal that includes an echo-only frame, the operations also include muting the respective output signal frame. After muting the respective output signal frame for each respective frame in the sequence of frames of the microphone signal that includes an echo-only frame, the operations also include performing audio processing on the respective output signal frame for each respective frame in the sequence of frames of the microphone signal that includes a double-talk frame.

[0004] Implementations of the present disclosure may include one or more of the following optional features. In some implementations, a portion of the microphone signal further includes an audio signal representing target speech captured by the microphone, and the operations also include determining whether each frame of the microphone signal includes a double-talk frame or an echo-only frame when the respective frame of the microphone signal includes the audio signal representing the target speech, where the target speech is spoken while audio content is played from an acoustic speaker. In some examples, performing speech processing includes performing speech recognition using an automatic speech recognition (ASR) model. In some implementations, the operations further include transforming each respective frame of the microphone signal, the reference signal, and the output signal into a short-time Fourier transform (STFT) domain before determining whether the respective frame of the microphone signal includes a double-talk frame or an echo-only frame using a DTD.

[0005] In some examples, determining whether each frame of the microphone signal includes a double-talk frame or an echo-only frame includes calculating a respective first frame-level double-talk indicator based on cross-correlation between the respective frame of the microphone signal and the respective frame of the reference signal using the DTD, and calculating a respective second frame-level double-talk indicator based on cross-correlation between the respective frame of the microphone signal and the respective frame of the output signal using the DTD. These examples further include determining whether at least one of the respective first frame-level double-talk indicator or the respective second frame-level double-talk indicator satisfies a double-talk threshold, and determining that each frame of the microphone signal includes a double-talk frame when at least one of the respective first frame-level double-talk indicator or the respective second frame-level double-talk indicator satisfies the double-talk threshold. In these examples, determining whether each frame of the microphone signal includes a double-talk frame or an echo-only frame may further include determining that each frame of the microphone signal includes an echo-only frame when both the respective first frame-level double-talk indicator and the respective second frame-level double-talk indicator do not satisfy the double-talk threshold. Both the respective first frame-level double-talk indicator and the respective second frame-level double-talk indicator may be calculated over a predetermined range of frequency subbands.Additionally or alternatively, determining whether at least one of the respective first frame-level double-talk indicators or the respective second frame-level double-talk indicators satisfies a double-talk threshold may include determining that at least one of the respective first frame-level double-talk indicators or the respective second frame-level double-talk indicators satisfies the double-talk threshold when the smaller of the respective first frame-level double-talk indicators and the respective second frame-level double-talk indicators is less than the double-talk threshold.

[0006] In some implementations, for each frame in the sequence of frames of the microphone signal, the operations further include calculating, using the DTD, a respective first frame-level double-talk indicator based on cross-correlation between the respective frame of the microphone signal and one of the respective frame of the reference signal or the respective frame of the output signal, wherein determining whether the respective frame of the microphone signal includes a double-talk frame or an echo-only frame is based on the respective first frame-level double-talk indicator. In these implementations, the operations may further include calculating, using the DTD, a respective second frame-level double-talk indicator based on cross-correlation between the respective frame of the microphone signal and the other of the respective frame of the reference signal or the respective frame of the output signal, wherein determining whether the respective frame of the microphone signal includes a double-talk frame or an echo-only frame is further based on the respective second frame-level double-talk indicator.

[0007] In some examples, the acoustic echo canceller includes a linear acoustic echo canceller. In some implementations, the data processing hardware, the microphone, and the acoustic speaker are present on a user computing device. In some examples, performing speech processing on the respective output signal frames for each respective frame in the sequence of microphone signals including the double-talk frame includes performing speech processing on the respective output signal frames without performing acoustic echo suppression on the respective output signal frames.

[0008] Another aspect of the present disclosure provides a system for performing speech recognition using an STFT-based echo mutator. The system includes data processing hardware and memory hardware in communication with the data processing hardware. The memory hardware stores instructions that, when executed on the data processing hardware, cause the data processing hardware to perform operations including receiving a microphone signal including an acoustic echo captured by a microphone. The acoustic echo corresponds to audio content played from an acoustic speaker. The operations also include receiving a reference signal including a sequence of frames representing audio content sent on a reference channel before the audio speaker plays the audio content. For each frame in the sequence of frames of the microphone signal, the operations include processing the respective frame of the microphone signal using an acoustic echo canceller configured to receive as input the respective frame in the sequence of frames of the reference signal to generate a respective output signal frame that removes the acoustic echo from the respective frame of the microphone signal. The operations also include using a double-talk detector (DTD) to determine whether the respective frame of the microphone signal includes a double-talk frame or an echo-only frame based on the respective frame of the reference signal and the respective output signal frame. For each respective frame in the sequence of frames of the microphone signal that includes an echo-only frame, the operations include muting the respective output signal frame. After muting the respective output signal frame for each respective frame in the sequence of frames of the microphone signal that includes an echo-only frame, the operations include performing speech processing on the respective output signal frame for each respective frame in the sequence of frames of the microphone signal that includes a double-talk frame.

[0009] This aspect may include one or more of the following features. In some implementations, the portion of the microphone signal further includes an audio signal representing target speech captured by the microphone, and the operations also include determining whether each frame of the microphone signal includes a double-talk frame or an echo-only frame when the respective frame of the microphone signal includes the audio signal representing the target speech, where the target speech is spoken while audio content is played from an acoustic speaker. In some examples, performing speech processing includes performing speech recognition using an automatic speech recognition (ASR) model. In some implementations, the operations further include transforming each respective frame of the microphone signal, the reference signal, and the output signal into a short-time Fourier transform (STFT) domain before determining whether the respective frame of the microphone signal includes a double-talk frame or an echo-only frame using a DTD.

[0010] In some examples, determining whether each frame of the microphone signal includes a double-talk frame or an echo-only frame includes calculating a respective first frame-level double-talk indicator based on cross-correlation between the respective frame of the microphone signal and the respective frame of the reference signal using the DTD, and calculating a respective second frame-level double-talk indicator based on cross-correlation between the respective frame of the microphone signal and the respective frame of the output signal using the DTD. These examples further include determining whether at least one of the respective first frame-level double-talk indicator or the respective second frame-level double-talk indicator satisfies a double-talk threshold, and determining that each frame of the microphone signal includes a double-talk frame when at least one of the respective first frame-level double-talk indicator or the respective second frame-level double-talk indicator satisfies the double-talk threshold. In these examples, determining whether each frame of the microphone signal includes a double-talk frame or an echo-only frame may further include determining that each frame of the microphone signal includes an echo-only frame when both the respective first frame-level double-talk indicator and the respective second frame-level double-talk indicator do not satisfy the double-talk threshold. Both the respective first frame-level double-talk indicator and the respective second frame-level double-talk indicator may be calculated over a predetermined range of frequency subbands.Additionally or alternatively, determining whether at least one of the respective first frame-level double-talk indicators or the respective second frame-level double-talk indicators satisfies a double-talk threshold may include determining that at least one of the respective first frame-level double-talk indicators or the respective second frame-level double-talk indicators satisfies the double-talk threshold when the smaller of the respective first frame-level double-talk indicators and the respective second frame-level double-talk indicators is less than the double-talk threshold.

[0011] In some implementations, for each frame in the sequence of frames of the microphone signal, the operations further include calculating, using the DTD, a respective first frame-level double-talk indicator based on cross-correlation between the respective frame of the microphone signal and one of the respective frame of the reference signal or the respective frame of the output signal, wherein determining whether the respective frame of the microphone signal includes a double-talk frame or an echo-only frame is based on the respective first frame-level double-talk indicator. In these implementations, the operations may further include calculating, using the DTD, a respective second frame-level double-talk indicator based on cross-correlation between the respective frame of the microphone signal and the other of the respective frame of the reference signal or the respective frame of the output signal, wherein determining whether the respective frame of the microphone signal includes a double-talk frame or an echo-only frame is further based on the respective second frame-level double-talk indicator.

[0012] In some examples, the acoustic echo canceller includes a linear acoustic echo canceller. In some implementations, the data processing hardware, the microphone, and the acoustic speaker are present on a user computing device. In some examples, performing speech processing on the respective output signal frames for each respective frame in the sequence of microphone signals including the double-talk frame includes performing speech processing on the respective output signal frames without performing acoustic echo suppression on the respective output signal frames.

[0013] The details of one or more implementations of the disclosure are set forth in the accompanying drawings and the description below. Other aspects, features, and advantages will be apparent from the description and drawings, and from the claims. [Brief explanation of the drawings]

[0014] [Figure 1] 1 is a schematic diagram of an exemplary audio environment using an Acoustic Echo Cancellation (AEC) system. [Figure 2] FIG. 1 is a schematic diagram of an AEC system. [Figure 3] 3 is a flow diagram of an exemplary sequence of operations for a method of implementing an acoustic echo cancellation system. [Figure 4] FIG. 1 is a schematic diagram of an example computing device that may be used to implement the systems and methods described herein. DETAILED DESCRIPTION OF THE INVENTION

[0015] Like reference symbols in the various drawings indicate like elements.

[0016] A voice-enabled device can generate synthesized playback audio and communicate the synthesized playback audio to one or more users in a voice environment. Here, synthesized playback audio refers to audio generated by a voice-enabled device that originates from the voice-enabled device itself or a machine processing system associated with the voice-enabled device, rather than from a person or other source of audible sound external to the voice-enabled device. Generally, voice-enabled devices generate synthesized playback audio using a text-to-speech (TTS) system. A TTS system converts text into an audio representation of the text, which is modeled to resemble an audio representation of spoken speech using a human language.

[0017] While an audio output component (e.g., an acoustic speaker) of a voice-enabled device outputs synthesized playback audio, an audio capture component (e.g., a microphone) of the voice-enabled device may still be actively capturing (i.e., listening to) audio signals in the audio environment. This means that a portion of the synthesized playback audio output from the speaker is received at the audio capture component as acoustic echo. While actively capturing audio signals in the audio environment, the voice-enabled device may also be playing other types of audio content, such as media content (e.g., music), which may also be captured as acoustic echo by the audio capture component. Unfortunately, the acoustic echo resulting from the playback audio content (e.g., synthesized audio content) may make it difficult for a speech recognizer implemented in the voice-enabled device or implemented in a remote system communicating with the voice-enabled device to understand verbal utterances that occur amid the echo from the synthesized playback audio. In other words, voice-enabled devices often generate synthesized playback audio in response to queries or commands from a user of the voice-enabled device. For example, a user may ask the voice-enabled device, "What's the weather like today?" When the voice-enabled device receives this query or question from the user, the voice-enabled device, or a remote system in communication with the voice-enabled device, must first determine or process the verbal utterance from the user. By processing the verbal utterance, the voice-enabled device can recognize that the verbal utterance corresponds to a query from the user (e.g., regarding the weather) and that, as a query, the user expects a response from the voice-enabled device.

[0018] Typically, a voice-enabled device uses a voice recognition system (e.g., an automatic speech recognition (ASR) system) to determine the content of a verbal utterance. The voice recognition system receives an audio signal or audio data and generates a text transcript representing the letters, words, and / or sentences spoken in the audio signal. However, voice recognition can become more complicated when the voice capture component of the voice-enabled device receives echo and / or distortion simultaneously with all or part of one or more utterances spoken by a user into the voice-enabled device. For example, one or more microphones of the voice-enabled device are provided with a portion of the synthesized playback audio signal as echo or acoustic feedback. The echo from the synthesized playback audio combined with the one or more verbal utterances results in the voice-enabled device receiving an audio signal with overlapping speech. Here, overlapping speech refers to a double-talk event in which the acoustic echo from the synthesized playback audio occurs at the same time (i.e., simultaneously or concurrently) as one or more verbal utterances. During a double-talk event, the voice recognition system may have difficulty processing the audio signal received at the voice-enabled device. That is, overlapping speech may impair the ability of a speech recognition system to generate an accurate transcription of one or more verbal utterances. Without an accurate transcription from a speech recognition system, a speech-enabled device may be unable to respond accurately, or at all, to queries or commands from verbal utterances by a user. Alternatively, a speech-enabled device may want to avoid using its processing resources attempting to interpret audible sounds that are actually synthesized playback audio signals and / or echoes from the surroundings.

[0019] One approach to combating distortion or echo captured by the audio capture component of a voice-enabled device is to use an acoustic echo cancellation (AEC) system. In an AEC system, the AEC system uses an audio signal to remove echo associated with audio content played from an acoustic speaker. However, the audio signal employed by the AEC system to remove the echo inevitably produces residual echo, which can further degrade the performance of a voice recognition system. In wake word applications, in which a predetermined word or phrase is spoken to invoke voice recognition by a voice-enabled device, the residual echo generated by the AEC system improves the ability of the voice recognition system to find the wake word. However, in wake word-less applications, this residual echo adversely affects the performance of the voice recognition system. One way to reduce this residual echo is by processing the audio signal employed by the AEC system with a post-filter (e.g., an echo suppressor). However, voice recognition systems are generally sensitive to this post-filtering of the audio signal, making the use of a post-filter a suboptimal solution.

[0020] 1 , in some implementations, an audio environment 100 includes a user 10 communicating a verbal utterance 12 to a voice-enabled device 110 (also referred to as a device 110 or a user device 110). The user 10 (i.e., the speaker of the utterance 12) may speak the utterance 12 as a query or a command to solicit a response from the device 110. The device 110 is configured to capture sounds from one or more users 10 in the audio environment 100. Here, audio sounds may refer to verbal utterances 12 by the user 10 that function as audible queries, commands for the device 110, or audible communications captured by the device 110. A voice-enabled system of or associated with the device 110 may process the query for the command by answering the query and / or causing the command to be executed.

[0021] Here, the device 110 receives an acoustic echo 156 captured by an audio capture device 116 (also referred to as a microphone) and / or a microphone signal 202 including oral utterances 12 by the user 10. The acoustic echo 156 corresponds to audio content 154 played from an audio output device 118 (also referred to as an audio speaker). The device 110 may correspond to any computing device associated with the user 10 and capable of receiving the microphone signal 202. Some examples of the user device 110 include, but are not limited to, mobile devices (e.g., mobile phones, tablets, laptops, etc.), computers, wearable devices (e.g., smart watches), smart appliances, and Internet of Things (IoT) devices, smart speakers, etc. The device 110 includes data processing hardware 112 and memory hardware 114 that communicates with the data processing hardware 112 and stores instructions that, when executed by the data processing hardware 112, cause the data processing hardware 112 to perform one or more operations. In some examples, device 110 includes one or more applications (i.e., software applications), each of which may utilize one or more audio processing systems 140, 150, 200 associated with device 110 to perform various functions within the application. For example, device 110 includes an assistant application configured to communicate synthesized playback audio content 154 to user 10 to assist user 10 with various tasks. In other examples, the assistant application or media application is configured to play audio content 154 including media content (e.g., music, talk radio, podcast content, television / movie content). In further examples, the application corresponds to an assistant application configured to communicate / converse with user 10 or communicate audio content 154 as synthesized speech for playback from acoustic speaker 118 to assist user 10 in performing various tasks.For example, the assistant application may audibly output synthesized speech in response to queries / commands sent to the assistant application by user 10. In a further example, audio content 154 played from sound speaker 118 corresponds to notifications / alerts such as, without limitation, a timer expiring, an incoming phone call alert, a doorbell chime, an audio message, etc.

[0022] Device 110 includes a microphone 116 for capturing and converting audio data 14 in audio environment 100 into electrical microphone signals 202, and an acoustic speaker 118 for communicating / outputting playback audio content 154 (e.g., synthesized playback audio content). While device 110 implements a microphone 116 in the illustrated example, it may implement an array of microphones 116 without departing from the scope of this disclosure, whereby one or more microphones 116 in the array may not be physically present on device 110 but may communicate with interfaces / peripherals of device 110. For example, device 110 may correspond to a vehicle infotainment system that utilizes an array of microphones located throughout the vehicle. Similarly, the acoustic speakers 118 may include either one or more speakers present on the device 110, one or more speakers in communication with the device 110, or a combination of one or more speakers present on the device 110 and one or more other speakers that are physically removed from the device 110 but in communication with the device 110.

[0023] Additionally, device 110 may be configured to communicate with remote system 130 via network 120. Remote system 130 may include remote resources 132, such as remote data processing hardware 134 (e.g., a remote server or CPU) and / or remote memory hardware 136 (e.g., a remote database or other storage hardware). Device 110 may utilize remote resources 132 to perform various functions related to speech processing and / or delivery of synthesized playback. For example, device 110 may be configured to perform speech recognition using speech recognition system 140 (e.g., using speech recognition model 145). Additionally, device 110 may be configured to perform text-to-speech conversion using TTS system 150 and acoustic echo cancellation using AEC system 200. These systems 140, 150, 200 may reside on device 110 (referred to as on-device systems) or may reside remotely (e.g., on remote system 130) but communicate with device 110. In some examples, some of the systems 140, 150, 200 reside locally or on-device, while other portions reside remotely. In other words, any of these systems 140, 150, 200 may be local or remote in any combination. For example, when the size or processing requirements of a system 140, 150, 200 are significant, the system 140, 150, 200 may reside on a remote system 130. However, when the device 110 can support the size or processing requirements of one or more of the systems 140, 150, 200, one or more of the systems 140, 150, 200 may reside on the device 110 using the data processing hardware 112 and / or memory hardware 114. Optionally, one or more of the systems 140, 150, 200 may reside both locally / on-device and remotely.For example, one or more of systems 140, 150, 200 may run on remote system 130 by default when a connection to network 120 between device 110 and remote system 130 is available, but when the connection is lost or network 120 is unavailable, systems 140, 150, 200 instead run locally on device 110.

[0024] Speech recognition system 140 receives audio signal 202 as input and transcribes the audio signal into transcription 142 as output. Generally, by converting audio signal 202 into transcription 142, speech recognition system 140 enables device 110 to recognize when spoken utterance 12 from user 10 corresponds to a query, a command, or some other form of audio communication. Transcription 142 refers to a sequence of text that device 110 may then use to generate a response to the query or command. For example, if user 10 asks device 110 the question, "What's the weather like today?", device 110 passes an audio signal corresponding to the question, "What's the weather like today?" to speech recognition system 140. Speech recognition system 140 converts the audio signal into a transcription that includes the text, "What's the weather like today?". Device 110 may then use the text, or a portion of the text, to determine a response to the query. For example, to determine the weather for the current day (i.e., today), device 110 passes text (e.g., "What's the weather like today?") or a specific portion of text (e.g., "weather" and "today") to a search engine. The search engine may then return one or more search results, and device 110 interprets the search results to generate a response for user 10.

[0025] In some implementations, device 110 or a system associated with device 110 identifies text 152 that device 110 communicates to user 10 in response to the query of spoken utterance 12. Device 110 may then use TTS system 150 to convert text 152 into corresponding synthesized playback audio 154 for device 110 to communicate to user 10 (e.g., audibly communicate to user 10) in response to the query of spoken utterance 12. In other words, TTS system 150 receives text 152 as input and converts the text 152 into output synthesized playback audio 154, which is an audio signal defining an audible representation of text 152. In some examples, TTS system 150 includes a text encoder that processes text 152 into an encoded format (e.g., text embedding). Here, TTS system 150 may use a trained text-to-speech model to generate synthesized playback audio 154 from the encoded format of text 152. Once generated, TTS system 150 communicates synthesized playback audio 154 to device 110 to enable device 110 to output synthesized playback audio 154. For example, device 110 outputs synthesized playback audio 154 of "It's sunny today" at speaker 118 of device 110.

[0026] 1, when device 110 outputs synthesized playback audio 154 (e.g., synthetic speech), the synthesized playback audio 154 produces an echo 156 that is captured by microphone 116. Unfortunately, in addition to the echo 156, microphone 116 may also be simultaneously capturing another verbal utterance 12 from user 10 that corresponds to the target speech directed at device 110. For example, FIG. 1 depicts user 10 further inquiring about the weather in verbal utterance 12 to device 110 by stating, "How's tomorrow?" as device 110 outputs synthesized playback audio content 154. In particular, the user 10 speaks an utterance 12 as part of a continued conversation scenario in which the device 110 keeps the microphone 116 open and the speech recognition system 140 active to allow the user 10 to provide a follow-up query for recognition by the speech recognition system 140 without requiring the user 10 to say a hot word (e.g., a predetermined word or phrase that, when detected, triggers the device 110 to invoke speech recognition). When the utterance 12 spoken by the user 10 includes a hot word, an echo 156 helps the speech recognition system 140 convert the audio signal 14 into a transcription. However, when the device 110 does not require the user 10 to say a hot word, the user device 110 may process the microphone signal 202 including the echo 156, causing the user device 110 to process its own playback audio 154 output from the speaker 118.

[0027] Here, the spoken utterance 12 and the echo 156 are both simultaneously captured by the microphone 116 to form the microphone signal 202. In other words, the microphone signal 202 includes a portion including only the echo 156 corresponding to the played audio content 154 output from the speaker 118 before the user 10 speaks the utterance 12, an overlapping portion (e.g., overlapping region 204) where the utterance 12 spoken by the user 10 overlaps with a portion of the played audio content 154 output from the speaker 118, and a portion including only the utterance 12 spoken by the user 10 after the acoustic speaker 118 stops outputting the played audio content 154.

[0028] 1 , the overlapping region 204 in the captured microphone signal 202 corresponds to a double-talk event, which indicates when a portion of the utterance 12 and a portion of the synthesized playback audio 154 overlap each other in the captured microphone signal 202. During a double-talk event, the speech recognition system 140 may have difficulty recognizing a subsequent query 12 corresponding to the weather question "How's tomorrow?" in the audio signal 202 captured by the microphone 116 because the utterance 12 is mixed with an echo 156 of the synthesized playback audio content 154. As discussed above, one way to reduce the echo 156 in the microphone signal 202 is by processing the microphone signal 202 with an echo suppressor filter. However, the speech recognition system 140 has difficulty processing this type of filtered microphone signal 202.

[0029] To address this, device 110 includes an AEC system 200 that processes microphone signal 202 and provides output to speech recognition system 140. AEC system 200 (FIG. 2) includes an acoustic echo canceller 210, a double-talk detector 220, and an echo mutator 230, and is configured to mute audio output signal frames identified as echo-only, while allowing audio frames containing double-talk to pass for processing by speech recognition system 140. In other words, AEC system 200 receives microphone signal 202 that includes subsequent query 12 mixed with echo 156. For each audio frame of the microphone signal, AEC system 200 processes the audio frame to determine whether the audio frame is echo-only or contains double-talk.

[0030] 2, the acoustic echo canceller 210 of the AEC system 200 is configured to receive a microphone signal 202 that may simultaneously include respective frames of target speech 12 directed at the device and a reference signal 158 corresponding to each frame representing the playback audio 154 captured by the microphone 116. For each frame in the sequence of frames, the acoustic echo canceller 210 processes the microphone signal 202 using the reference signal 158 to remove the echo 156 in the microphone signal 202, thereby generating an output signal frame 206. In some examples, the AEC system 200 further transforms each respective frame of the microphone signal 202, the reference signal 158, and the output signal 206 into the short-time Fourier transform (STFT) domain using the following equation: Y(l, k) =Σ n y[n]w[n - lL] exp(-j2πkn / N), k = 0,...,N - 1 (1) where L represents the frame-hop, w[n] represents the possible N-point analysis window, l represents the frame index, and k represents the subband index.

[0031] The double-talk detector 220 then receives each frame of the STFT domain of the microphone signal m[n] 202 (denoted as M(l,k)), each frame of the STFT domain of the reference signal x[n] 158 (denoted as X(l,k)), and each frame of the STFT domain of the residual echo output signal r[n] 206 (denoted as R(l,k)), and determines whether each frame contains a double-talk frame or an echo-only frame. To accomplish this, the double-talk detector 220 uses two double-talk indicators 208 for each subband of each respective frame. The two double-talk indicators 208a, 208b may be expressed as follows:

[0032]

number

[0033] In the formula, c PQ (k) represents the complex cross-correlation of each subband k at lag 0 in the STFT domain of the reproduced audio frame 154 P(l, k) and Q(l, k), where P(l, k) and Q(l, k) are associated with the time-domain signals p[n] and q[n], respectively. The complex cross-correlation c PQ (k) is expressed as follows: c PQ (k) = E{P(l, k)Q * (l, k)} (3) where E{·} represents the mathematical expectation and ·* represents the complex conjugate. Furthermore, the energy of subband k may be expressed as:

[0034]

number

[0035] The cross-correlation of each subband determines the energy of the subband by using an exponentially weighted average over the STFT domain using the following formula: c PQ (l, k) =ρc PQ (l - 1, k) + (1 -ρ)P(l, k)Q * (l, k) σ 2 (l, k) = ρσ 2 (l - 1, k) + (1 -ρ)P(l, k)P * (l, k) (5) where ρ represents an exponential weighting factor (i.e., forgetting factor). For example, an exponential weighting factor ρ of 0.9 may strike a balance between estimation accuracy and response time. Using this, equation (2) may be reformulated and expressed as follows:

[0036]

number

[0037] In some examples, the first double talk indicator 208a is calculated by cross-correlation between the microphone signal 202 and the reference signal 158, while the second double talk indicator 208b is calculated by cross-correlation between the microphone signal 202 and the output signal 206.

[0038] The double-talk detector 220 calculates double-talk indicators 208a, 208b for each respective frame by calculating a weighted average of the subband indicators across a limited number of subbands using the indicators and subband energies. For example, subbands between 700 Hz and 2400 Hz may be used to determine the double-talk indicators 208a, 208b. Each double-talk indicator 208a, 208b is weighted by the energy of the residual within a given subband by cross-correlating each frame in a sequence of frames to determine the energy within the band. A higher residual energy within a subband indicates a greater likelihood of the presence of double talk, causing the double-talk indicator 208a, 208b to indicate that the frame contains double talk. The first double-talk indicator 208a may be calculated as follows:

[0039]

number

[0040] Similarly, the second double talk indicator 208b may be calculated as follows:

[0041]

number

[0042] Once the double-talk detector 220 calculates the double-talk indicators 208a, 208b as an output, the output is provided to the echo mutator 230 as an input for determining whether at least one of the double-talk indicators 208a, 208b meets a double-talk threshold. In some examples, the double-talk detector 220 calculates only one of the first or second double-talk indicators 208a, 208b as an input to the echo mutator 230 for determining whether the double-talk threshold is met. Given a threshold τ, the double-talk detector 220 determines whether each frame contains only echo, which should be muted, or double-talk, which should be passed. This is determined as follows:

[0043]

number

[0044] When at least one of the double-talk indicators 208a, 208b meets the double-talk threshold, the respective frame is passed for processing by the speech recognition system 140. When both double-talk indicators 208a, 208b do not meet the double-talk threshold, the respective frame is identified as echo only and muted (i.e., not passed for processing by the speech recognition system). Requiring the smaller of the double-talk indicators 208a, 208b to meet the double-talk threshold favors passing the frame over muting it.

[0045] The double-talk threshold τ may be calculated using a data-driven approach. Metadata is collected from the utterance to segment the utterance into echo-only and double-talk periods. Frame-level double-talk indicators 208a, 208b are then calculated for every frame in the training dataset. For a given threshold τ, the percentage of misclassifications of echo-only and double-talk frames is calculated. This may be repeated for a range of thresholds τ, with the goal of identifying a threshold τ that results in misclassification of 10% or less double-talk frames.

[0046] 3 includes a flow diagram of an example sequence of operations of a method 300 for performing speech recognition using the STFT-based echo mutator 230. At operation 302, the method 300 includes receiving a microphone signal 202 including an acoustic echo 156 captured by the microphone 116. The acoustic echo 156 corresponds to audio content 154 played from the sound speaker 118. At operation 304, the method 300 includes receiving a reference signal 158 including a sequence of frames representing the audio content 154 sent on a reference channel before the sound speaker 118 plays the audio content 154.

[0047] For each frame in the sequence of frames of the microphone signal 202, the method 300 includes, at operation 306, processing the respective frame of the microphone signal 202 using an acoustic echo canceller 210 configured to receive as input the respective frame in the sequence of frames of the reference signal 158 to generate a respective output signal frame 206 that removes the acoustic echo 156 from the respective frame of the microphone signal 202. At operation 308, the method 300 includes using a double-talk detector 220 to determine whether the respective frame of the microphone signal 202 includes a double-talk frame or an echo-only frame based on the respective frame of the reference signal 158 and the respective output signal frame 206. For each respective frame in the sequence of frames of the microphone signal 202 that includes an echo-only frame, the method 300 at operation 310 includes muting the respective output signal frame 206. In operation 312, after muting the respective output signal frame 206 for each respective frame in the sequence of frames of the microphone signal 202 that includes echo-only frames, the method 300 includes performing audio processing on the respective output signal frame 206 for each respective frame in the sequence of frames of the microphone signal 202 that includes double-talk frames.

[0048] 4 is a schematic diagram of an exemplary computing device 400 that may be used to implement the systems and methods described herein. Computing device 400 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The components shown herein, their connections and relationships, and their functionality are intended to be merely exemplary and are not intended to limit the implementation of the present disclosure as described and / or claimed herein.

[0049] Computing device 400 includes a processor 410, memory 420, a storage device 430, a high-speed interface / controller 440 that connects to memory 420 and a high-speed expansion port 450, and a low-speed interface / controller 460 that connects to a low-speed bus 470 and storage device 430. Each of components 410, 420, 430, 440, 450, and 460 are interconnected using various buses and may be mounted on a common motherboard or otherwise suitable. Thus, processor 410 (also referred to as “data processing hardware 410,” which may include data processing hardware 112 of user computing device 110 or data processing hardware 134 of remote system 130) can process instructions for execution within computing device 400, including instructions stored in memory 420 or on storage device 430, and display graphical information for a graphical user interface (GUI) on an external input / output device, such as a display 480, coupled to high-speed interface 440. In other implementations, multiple processors and / or multiple buses may be used, along with multiple memories and types of memory, as appropriate, and multiple computing devices 400 may be connected (e.g., as a server bank, a group of blade servers, or a multiprocessor system) with each device providing a portion of the required operations.

[0050] Memory 420 (also referred to as “memory hardware 420,” which may include memory hardware 114 of user computing device 110 or memory hardware 136 of remote system 130) non-transitoryly stores information within computing device 400. Memory 420 may be a computer-readable medium, a volatile memory unit, or a non-volatile memory unit. Non-transitory memory 420 may be a physical device used to temporarily or permanently store programs (e.g., sequences of instructions) or data (e.g., program state information) for use by computing device 400. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electronically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and disk or tape.

[0051] The storage device 430 can provide mass storage for the computing device 400. In some implementations, the storage device 430 is a computer-readable medium. In various different implementations, the storage device 430 can be a floppy disk device, a hard disk device, an optical disk device, or an array of devices including a tape device, a flash memory or other similar solid-state memory device, or a device in a storage area network or other configuration. In additional implementations, a computer program product is tangibly embodied in an information carrier. The computer program product includes instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer-readable or machine-readable medium, such as the memory 420, the storage device 430, or memory on the processor 410.

[0052] The high-speed controller 440 manages bandwidth-intensive operations for the computing device 400, while the low-speed controller 460 manages less bandwidth-intensive operations. Such assignment of roles is merely exemplary. In some implementations, the high-speed controller 440 is coupled to the memory 420, to the display 480 (e.g., through a graphics processor or accelerator), and to a high-speed expansion port 450, which may accept various expansion cards (not shown). In some implementations, the low-speed controller 460 is coupled to the storage device 430 and to a low-speed expansion port 490. The low-speed expansion port 490, which may include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), may be coupled to one or more input / output devices, such as a keyboard, pointing device, scanner, etc., or may be coupled to a network device, such as a switch or router, for example, via a network adapter.

[0053] Computing device 400 may be implemented in many different forms, as shown in the figure. For example, computing device 400 may be implemented as a standard server 400a, or multiple times within a group of such servers 400a, as a laptop computer 400b, or as part of a rack server system 400c.

[0054] Various implementations of the systems and techniques described herein may be realized in digital electronic and / or optical circuitry, integrated circuits, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementation in one or more computer programs executable and / or interpretable on a programmable system that includes at least one programmable processor, which may be special-purpose or general-purpose, coupled to receive data and instructions from and transmit data and instructions to a storage system, at least one input device, and at least one output device.

[0055] A software application (i.e., a software resource) may refer to computer software that causes a computing device to perform tasks. In some examples, a software application may be referred to as an "application," "app," or "program." Exemplary applications include, but are not limited to, system diagnostic applications, system management applications, system maintenance applications, word processing applications, spreadsheet applications, messaging applications, media streaming applications, social networking applications, and gaming applications.

[0056] Non-transitory memory may be a physical device used to temporarily or permanently store programs (e.g., sequences of instructions) or data (e.g., program state information) for use by a computing device. Non-transitory memory may be volatile and / or non-volatile addressable semiconductor memory. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electronically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and disk or tape.

[0057] These computer programs (also known as programs, software, software applications, or code) contain machine instructions for a programmable processor and may be implemented in a high-level procedural and / or object-oriented programming language, and / or assembly / machine language. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, non-transitory computer-readable medium, apparatus, and / or device (e.g., magnetic disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0058] The processes and logic flows described herein may be implemented by one or more programmable processors, also referred to as data processing hardware, executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows may also be implemented by special-purpose logic circuitry, such as a field-programmable gate array (FPGA) or an application-specific integrated circuit (ASIC). Processors suitable for executing computer programs include, by way of example, both general-purpose and special-purpose microprocessors, as well as any one or more processors of any type of digital computer. A processor typically receives instructions and data from a read-only memory or a random-access memory, or both. Essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. A computer typically also includes one or more mass storage devices, such as magnetic, magneto-optical, or optical disks, for storing data, or is operatively coupled to receive data from or transfer data to such mass storage devices, or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include, by way of example, all forms of non-volatile memory, media, and memory devices, including semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices, magnetic disks, e.g., internal hard disks or removable disks, magneto-optical disks, and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.

[0059] To provide for interaction with a user, one or more aspects of the present disclosure may be implemented on a computer having a display device, e.g., a CRT (cathode ray tube), LCD (liquid crystal display) monitor, or touch screen, for displaying information to the user, and optionally a keyboard and pointing device, e.g., a mouse or trackball, by which the user can provide input to the computer. Other types of devices may also be used to provide for interaction with the user; for example, feedback provided to the user may be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback, and input from the user may be received in any form, including acoustic, speech, or tactile input. Additionally, the computer may interact with the user by sending documents to and receiving documents from a device used by the user, e.g., by sending a web page to a web browser on the user's client device in response to a request received from the web browser.

[0060] Although several implementations have been described, it will be understood that various modifications may be made without departing from the spirit and scope of the present disclosure. Accordingly, other implementations are within the scope of the following claims. [Explanation of symbols]

[0061] 10 users 12 Oral utterances, subsequent queries, and target speech 14 Audio Data 100 Audio Environment 110 Voice-enabled devices, devices, user devices, user computing devices 112 Data Processing Hardware 114 Memory Hardware 116 Audio capture device, microphone 118 Audio output device, acoustic speaker 120 Network 130 Remote Systems 132 Remote Resources 134 Remote Data Processing Hardware 136 Remote Memory Hardware 140 Speech processing systems, speech recognition systems 142 Transcription 145 Speech Recognition Model 150 Speech processing systems, TTS systems 152 Text 154 Audio Content, Playback Audio Content, Synthesized Playback Audio, Playback Audio Frames 156 Acoustic Echo 158 Reference signal 200 Audio processing system, AEC system 202 Microphone Signal 204 overlapping areas 206 Output signal, output signal frame, residual echo output signal 208, 208a, 208b Double Talk Indicator 210 Acoustic Echo Canceller 220 Double Talk Detector 230 Echomuta 300 ways 400 computing devices 400a Standard Server 400b laptop computer 400c Rack Server System 410 Processors, Components, and Data Processing Hardware 420 Memory, Components, and Memory Hardware 430 Storage Devices, Components 440 High-Speed ​​Interface / Controller, Components 450 High-Speed ​​Expansion Port, Component 460 Low-speed interface / controller, component 470 Slow Bus 480 display 490 Low-Speed ​​Expansion Port

Claims

1. A computer-implemented method (300) that, when executed on data processing hardware (410), causes the data processing hardware (410) to perform operations, the operations comprising: receiving a microphone signal (202) including an acoustic echo (156) captured by a microphone (116), the acoustic echo (156) corresponding to audio content (154) played from an audio speaker (118); receiving a reference signal (158) including a sequence of frames representing the audio content (154) sent on a reference channel before the sound speaker (118) plays the audio content (154); For each frame in the sequence of frames of the microphone signal (202), processing each frame of the microphone signal (202) using an acoustic echo canceller (210) configured to receive as input a respective frame in a sequence of frames of the reference signal (158) to generate a respective output signal frame (206) that removes the acoustic echo (156) from the respective frame of the microphone signal (202); determining, using a double-talk detector (DTD) (220), based on the respective frames of the reference signal (158) and the respective output signal frames (206), whether the respective frames of the microphone signal (202) comprise double-talk frames or echo-only frames, calculating, using the DTD (220), each first frame-level double-talk indicator (208a) based on a cross-correlation between the respective frame of the microphone signal (202) and the respective frame of the reference signal (158); calculating, using the DTD (220), a respective second frame-level double-talk indicator (208b) based on a cross-correlation between the respective frame of the microphone signal (202) and the respective output signal frame (206); determining whether at least one of the respective first frame-level double-talk indicator (208a) or the respective second frame-level double-talk indicator (208b) satisfies a double-talk threshold; and for each respective frame in the sequence of frames of the microphone signal (202) including the echo-only frame, muting the respective output signal frame (206); and muting the respective output signal frame (206) for each respective frame in the sequence of frames of the microphone signal (202) including the echo-only frame, and then performing speech processing on the respective output signal frame (206) for each respective frame in the sequence of frames of the microphone signal (202) including the double-talk frame.

2. a portion of the microphone signal (202) further comprising an audio signal representing target speech (12) captured by the microphone (116), the target speech (12) being spoken while the audio content (154) is played from the acoustic speaker (118); 2. The computer-implemented method of claim 1, wherein the operations further comprise determining whether the respective frame of the microphone signal includes the double-talk frame or the echo-only frame when the respective frame of the microphone signal includes the audio signal representing the target speech.

3. 3. The computer-implemented method (300) of claim 1 or 2, wherein performing speech processing includes performing speech recognition using an automatic speech recognition (ASR) model (145).

4. 4. The computer-implemented method of claim 1, wherein the operations further comprise transforming each respective frame of the microphone signal, the reference signal, and the output signal into a short-time Fourier transform (STFT) domain before using the DTD to determine whether the respective frame of the microphone signal comprises the double-talk frame or the echo-only frame.

5. Determining whether the respective frames of the microphone signal (202) include the double-talk frame or the echo-only frame, 5. The computer-implemented method of claim 1, further comprising: determining that the respective frame of the microphone signal comprises the double-talk frame when at least one of the respective first frame-level double-talk indicator (208a) or the respective second frame-level double-talk indicator (208b) satisfies the double-talk threshold.

6. 6. The computer-implemented method of claim 5, wherein determining whether the respective frames of the microphone signal include the double-talk frame or the echo-only frame further comprises determining that the respective frames of the microphone signal include the echo-only frame when both the respective first frame-level double-talk indicator (208a) and the respective second frame-level double-talk indicator (208b) do not satisfy the double-talk threshold.

7. 7. The computer-implemented method of claim 5, wherein the respective first frame-level double-talk indicators and the respective second frame-level double-talk indicators are both calculated over a predetermined range of frequency subbands.

8. 8. The computer-implemented method of claim 5, wherein determining whether at least one of the respective first frame-level double-talk indicators (208 a) or the respective second frame-level double-talk indicators (208 b) satisfies the double-talk threshold comprises determining that at least one of the respective first frame-level double-talk indicators (208 a) or the respective second frame-level double-talk indicators (208 b) satisfies the double-talk threshold when a smaller one of the respective first frame-level double-talk indicators (208 a) and the respective second frame-level double-talk indicators (208 b) is less than the double-talk threshold.

9. The operation comprises, for each frame in the sequence of frames of the microphone signal (202): using the DTD (220) to calculate a respective first frame-level double-talk indicator (208 a) based on a cross-correlation between the respective frame of the microphone signal (202) and one of the respective frame of the reference signal (158) or the respective output signal frame (206); 9. The computer-implemented method (300) of any one of claims 1 to 8, wherein determining whether the respective frame of the microphone signal (202) comprises the double-talk frame or the echo-only frame is based on the respective first frame-level double-talk indicator (208a).

10. The operation comprises, for each frame in the sequence of frames of the microphone signal (202): using the DTD (220) to calculate a respective second frame-level double-talk indicator (208b) based on a cross-correlation between the respective frame of the microphone signal (202) and the other of the respective frame of the reference signal (158) or the respective output signal frame (206); 10. The computer-implemented method of claim 9, wherein determining whether the respective frame of the microphone signal comprises the double-talk frame or the echo-only frame is further based on the respective second frame-level double-talk indicator.

11. 11. The computer-implemented method (300) of any one of claims 1 to 10, wherein the acoustic echo canceller (210) comprises a linear acoustic echo canceller.

12. 12. The computer-implemented method (300) of any one of claims 1 to 11, wherein the data processing hardware (410), the microphone (116), and the acoustic speaker (118) reside on a user computing device (110).

13. 2. The computer-implemented method of claim 1, wherein performing speech processing on the respective output signal frames for each respective frame in the sequence of the microphone signals including the double-talk frame comprises performing speech processing on the respective output signal frames without performing acoustic echo suppression on the respective output signal frames.

14. data processing hardware (410); Memory hardware (114, 136) in communication with the data processing hardware (410), storing instructions that, when executed on the data processing hardware (410), cause the data processing hardware (410) to perform operations, the operations including: receiving a microphone signal (202) including an acoustic echo (156) captured by a microphone (116), the acoustic echo (156) corresponding to audio content (154) played from an audio speaker (118); receiving a reference signal (158) including a sequence of frames representing the audio content (154) sent on a reference channel before the sound speaker (118) plays the audio content (154); For each frame in the sequence of frames of the microphone signal (202), processing each frame of the microphone signal (202) using an acoustic echo canceller (210) configured to receive as input a respective frame in a sequence of frames of the reference signal (158) to generate a respective output signal frame (206) that removes the acoustic echo (156) from the respective frame of the microphone signal (202); determining, using a double-talk detector (DTD) (220), based on the respective frames of the reference signal (158) and the respective output signal frames (206), whether the respective frames of the microphone signal (202) comprise double-talk frames or echo-only frames, calculating, using the DTD (220), each first frame-level double-talk indicator (208a) based on a cross-correlation between the respective frame of the microphone signal (202) and the respective frame of the reference signal (158); calculating, using the DTD (220), a respective second frame-level double-talk indicator (208b) based on a cross-correlation between the respective frame of the microphone signal (202) and the respective output signal frame (206); determining whether at least one of the respective first frame-level double-talk indicator (208a) or the respective second frame-level double-talk indicator (208b) satisfies a double-talk threshold; and for each respective frame in the sequence of frames of the microphone signal (202) including the echo-only frame, muting the respective output signal frame (206); and and memory hardware (114, 136), including muting the respective output signal frames (206) for each respective frame in the sequence of frames of the microphone signal (202) including the echo-only frames, and then performing audio processing on the respective output signal frames (206) for each respective frame in the sequence of frames of the microphone signal (202) including the double-talk frames.

15. a portion of the microphone signal (202) further comprising an audio signal representing target speech (12) captured by the microphone (116), the target speech (12) being spoken while the audio content (154) is played from the acoustic speaker (118); 15. The system of claim 14, wherein the operations further include determining whether the respective frame of the microphone signal (202) includes the double-talk frame or the echo-only frame when the respective frame of the microphone signal (202) includes the audio signal representing the target speech (12).

16. 16. The system of claim 14 or 15, wherein performing speech processing includes performing speech recognition using an automatic speech recognition (ASR) model (145).

17. 17. The system of claim 14, wherein the operations further include transforming each respective frame of the microphone signal, the reference signal, and the output signal into a short-time Fourier transform (STFT) domain before using the DTD to determine whether the respective frame of the microphone signal includes the double-talk frame or the echo-only frame.

18. Determining whether the respective frames of the microphone signal (202) include the double-talk frame or the echo-only frame, 18. The system of claim 14, further comprising: determining that the respective frame of the microphone signal comprises the double-talk frame when at least one of the respective first frame-level double-talk indicator (208a) or the respective second frame-level double-talk indicator (208b) satisfies the double-talk threshold.

19. 20. The system of claim 18, wherein determining whether the respective frames of the microphone signal (202) include the double-talk frame or the echo-only frame further comprises determining that the respective frames of the microphone signal (202) include the echo-only frame when both the respective first frame-level double-talk indicator (208a) and the respective second frame-level double-talk indicator (208b) do not satisfy the double-talk threshold.

20. 20. The system of claim 18 or 19, wherein the respective first frame-level double-talk indicator (208a) and the respective second frame-level double-talk indicator (208b) are both calculated over a predetermined range of frequency subbands.

21. 21. The system of claim 18, wherein determining whether at least one of the respective first frame-level double-talk indicators (208 a) or the respective second frame-level double-talk indicators (208 b) satisfies the double-talk threshold comprises determining that at least one of the respective first frame-level double-talk indicators (208 a) or the respective second frame-level double-talk indicators (208 b) satisfies the double-talk threshold when a smaller one of the respective first frame-level double-talk indicators (208 a) and the respective second frame-level double-talk indicators (208 b) is less than the double-talk threshold.

22. The operation comprises, for each frame in the sequence of frames of the microphone signal (202): using the DTD (220) to calculate a respective first frame-level double-talk indicator (208 a) based on a cross-correlation between the respective frame of the microphone signal (202) and one of the respective frame of the reference signal (158) or the respective output signal frame (206); 22. The system of claim 14, wherein determining whether the respective frame of the microphone signal (202) comprises the double-talk frame or the echo-only frame is based on the respective first frame-level double-talk indicator (208a).

23. The operation comprises, for each frame in the sequence of frames of the microphone signal (202): using the DTD (220) to calculate a respective second frame-level double-talk indicator (208b) based on a cross-correlation between the respective frame of the microphone signal (202) and the respective frame of the reference signal (158) or the respective frame of the output signal; 23. The system of claim 22, wherein determining whether the respective frame of the microphone signal (202) comprises the double-talk frame or the echo-only frame is further based on the respective second frame-level double-talk indicator (208b).

24. 24. The system of any one of claims 14 to 23, wherein the acoustic echo canceller (210) comprises a linear acoustic echo canceller.

25. 25. The system of any one of claims 14 to 24, wherein the data processing hardware (410), the microphone (116), and the acoustic speaker (118) reside on a user computing device (110).

26. 26. The system of claim 14, wherein performing speech processing on the respective output signal frame (206) for each respective frame in the sequence of the microphone signal (202) including the double talk frame comprises performing speech processing on the respective output signal frame (206) without performing acoustic echo suppression on the respective output signal frame (206).

Citation Information

Patent Citations

  • Echo canceller and voice communication equipment provided with this echo canceller

    JP2000252885A

  • Audio processing apparatus

    JP2007235770A

  • Telecommunication apparatus

    JP2009094802A

  • Minimizing echo due to speaker-to-microphone coupling changes in an acoustic echo canceler

    WO2020154300A1