Single microphone acoustic echo and noise suppression

By using a delay estimator and a masking technique generated by a deep neural network, speech, echo, and noise components in the microphone signal are separated, solving the problem of poor nonlinear echo processing in existing technologies. This achieves effective echo and noise suppression in untrained environments and improves the quality of speech signals.

CN121889852APending Publication Date: 2026-04-17SYNAPTICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SYNAPTICS INC
Filing Date
2024-08-28
Publication Date
2026-04-17

Smart Images

  • Figure CN121889852A_ABST
    Figure CN121889852A_ABST
Patent Text Reader

Abstract

The invention provides a method, a device and a system for audio signal processing. The present implementations more particularly relate to speech enhancement techniques for separating microphone signals into speech, echo, and noise signals. In some aspects, a speech enhancement system may include a delay estimator and an acoustic echo and noise (AEN) decoupling filter. The delay estimator receives a microphone signal via the microphone and receives a far-end audio signal for output via the speaker, and estimates a reference audio signal based on a delay between the microphone signal and the far-end audio signal. In some aspects, an AEN decoupling filter may determine a speech mask, an echo mask, and a noise mask based on a microphone signal and a reference audio signal, and may suppress echo and noise components of the microphone signal based on the determined set of masks.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications This patent application claims priority to U.S. Nonprovisional Patent Application No. 18 / 460,442 (filed September 1, 2023, entitled “SINGLE-MICROPHONE ACOUSTIC ECHO AND NOISE SUPPRESSION”), which has been assigned to the assignee of this application and is incorporated herein by reference in its entirety. Technical Field

[0002] This implementation generally involves audio signal processing, and specifically involves single-microphone acoustic echo and noise suppression techniques. Background Technology

[0003] Many hands-free communication devices (such as Voice over Internet Protocol (VoIP) phones, speakerphones, and mobile phones configured to operate in hands-free mode) include a microphone and a speaker positioned relatively close to each other. The microphone is configured to convert sound waves from the surrounding environment into an audio signal (also known as a "microphone signal"), which can be transmitted over a communication channel to a remote device. The speaker is configured to convert the audio signal received from the remote device into sound waves that can be heard by a nearby user. Because the speaker is close to the microphone, the microphone signal can include a speech component (representing audio originating from the nearby user), an echo component (representing audio emitted by the speaker), and a noise component (representing ambient audio from the background environment).

[0004] Acoustic echo cancellation (AEC) refers to various techniques that attempt to eliminate or suppress the echo component of a microphone signal. Many existing AEC techniques rely on a linear transfer function that approximates the impulse response between the speaker and the microphone. For example, this linear transfer function can be determined using adaptive filters (such as the Normalized Least Mean Square (NLMS) algorithm) that model the acoustic coupling (or channel) between the speaker and microphone. However, the convergence rate of the NLMS algorithm can depend on the double-talk state (such as when a near-end user and a far-end user are speaking simultaneously) and changes to the echo path. Furthermore, such linear transfer functions cannot account for nonlinearities introduced along the echo path by the various mechanical components of the amplifier and speaker. Therefore, there is a need for further improvements to the quality of speech in the microphone signal. Summary of the Invention

[0005] This overview is provided to present, in a simplified form, the selection of concepts further described in the detailed description below. This overview is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter.

[0006] One innovative aspect of the subject matter of this disclosure can be implemented using a speech enhancement method. The method includes the steps of: receiving a first audio signal via a microphone; receiving a second audio signal for output via a speaker; estimating a reference audio signal based on a delay between the first and second audio signals; determining a plurality of masks based on the first and reference audio signals, wherein the plurality of masks includes a speech mask associated with speech components of the first audio signal, an echo mask associated with echo components of the first audio signal, and a noise mask associated with noise components of the first audio signal; and suppressing the echo and noise components of the first audio signal at least in part based on the plurality of masks.

[0007] Another innovative aspect of the subject matter of this disclosure can be implemented as a speech enhancement system, the system comprising a processing system and a memory. The memory stores instructions that, when executed by the processing system, cause the speech enhancement system to: receive a first audio signal via a microphone; receive a second audio signal for output via a speaker; estimate a reference audio signal based on a delay between the first and second audio signals; determine a plurality of masks based on the first and reference audio signals, wherein the plurality of masks includes a speech mask associated with speech components of the first audio signal, an echo mask associated with echo components of the first audio signal, and a noise mask associated with noise components of the first audio signal; and suppress the echo and noise components of the first audio signal at least in part based on the plurality of masks. Attached Figure Description

[0008] This implementation is illustrated by way of example and is not intended to be limited to the figures in the accompanying drawings.

[0009] Figure 1 An example hands-free communication system is shown.

[0010] Figure 2 A block diagram of an example speech enhancement system based on some implementation methods is shown.

[0011] Figure 3 A block diagram of an example acoustic echo and noise (AEN) decoupling system is shown, based on some implementation methods.

[0012] Figure 4A A block diagram of an example audio mask generation system based on some implementation methods is shown.

[0013] Figure 4B Another block diagram is shown for an example audio mask generation system based on some implementation methods.

[0014] Figure 5 Another block diagram of an example speech enhancement system based on some implementations is shown.

[0015] Figure 6An illustrative flowchart depicting example operations for speech enhancement based on some implementations is shown. Detailed Implementation

[0016] In the following description, numerous specific details, such as examples of particular components, circuits, and processes, are set forth to provide a thorough understanding of this disclosure. As used herein, the term “coupled” means a direct connection to or a connection via one or more intermediary components or circuits. The terms “electronic system” and “electronic device” may be used interchangeably to refer to any system capable of electronically processing information. Similarly, specific terminology is set forth in the following description and for purposes of explanation to provide a thorough understanding of various aspects of this disclosure. However, it will be apparent to those skilled in the art that these specific details are not required to practice the exemplary embodiments. In other instances, well-known circuits and devices are shown in block diagram form to avoid obscuring this disclosure. Some portions of the following detailed description are presented as procedures, logic blocks, processes, and other symbolic representations of the operation of data bits within computer memory.

[0017] These descriptions and representations are the means by which those skilled in the art of data processing most effectively communicate the substance of their work to others skilled in the art. In this disclosure, procedures, logic blocks, processes, or the like are conceived as self-consistent sequences of steps or instructions that lead to desired results. These steps are those that require physical manipulation of physical quantities. These quantities typically (but not necessarily) take the form of electrical or magnetic signals that can be stored, transmitted, combined, compared, and otherwise manipulated in a computer system. However, it should be noted that all these and similar terms are to be associated with appropriate physical quantities and are merely convenient labels applied to those quantities.

[0018] Unless specifically stated in a manner distinct from the following discussion, it will be understood that throughout this application, discussions using terms such as “access,” “receive,” “send,” “use,” “select,” “determine,” “normalize,” “multiply,” “average,” “monitor,” “compare,” “apply,” “update,” “measure,” “derive,” or the like refer to the actions and processes of a computer system (or similar electronic computing device) that manipulates data represented as physical (electronic) quantities within the registers and memories of the computer system and converts them into other data similarly represented as physical quantities within the computer system’s memory or registers or other such information storage, transmission, or display devices.

[0019] In the figures, a single box may be described as performing a function or multiple functions; however, in practice, the function or multiple functions performed by that box may be performed in a single component or through multiple components and / or may be performed using hardware, using software, or using a combination of hardware and software. To clearly illustrate this interchangeability of hardware and software, various illustrative components, boxes, modules, circuits, and steps are generally described below in terms of their functional aspects. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the system as a whole. Those skilled in the art may implement the described functionality in different ways for each specific application, but such implementation decisions should not be construed as causing a departure from the scope of this disclosure. Similarly, example input devices may include components other than those shown, including well-known components such as processors, memory, and the like.

[0020] Unless specifically described as implemented in a particular manner, the techniques described herein can be implemented in hardware, software, firmware, or any combination thereof. Any feature described as a module or component can also be implemented together as an integrated logic device, or as discrete but interoperable logic devices. If implemented in software, the techniques can be implemented at least in part by a non-transitory processor-readable storage medium comprising instructions that, when executed, perform one or more of the methods described above. The non-transitory processor-readable data storage medium can form part of a computer program product, which may include packaging materials.

[0021] Non-transitory processor-readable storage media may include random access memory (RAM) (such as synchronous dynamic random access memory (SDRAM)), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), flash memory, other known storage media, and the like. The technology may also be implemented, at least in part, by a processor-readable communication medium that carries or transmits code in the form of instructions or data structures, and that the code may be accessed, read, and / or executed by a computer or other processor.

[0022] The various illustrative logic blocks, modules, circuits, and instructions described in conjunction with the embodiments disclosed herein can be executed by one or more processors (or processing systems). As used herein, the term "processor" can refer to any general-purpose processor, special-purpose processor, conventional processor, controller, microcontroller, and / or state machine capable of executing instructions or scripts of one or more software programs stored in memory.

[0023] As described above, many hands-free communication devices include a microphone and a speaker positioned relatively close to each other. Thus, the microphone signal captured by the microphone can include a speech component (representing audio originating from a nearby user), an echo component (representing audio emitted by the speaker), and a noise component (representing ambient audio from the background environment). Acoustic echo cancellation (AEC) refers to various techniques that attempt to eliminate or suppress the echo component of the microphone signal. Many existing AEC techniques rely on a linear transfer function that approximates the impulse response between the speaker and the microphone. However, such a linear transfer function cannot account for the nonlinearities introduced along the echo path by the various mechanical components of the amplifier and speaker.

[0024] Some modern AEC techniques rely on machine learning to separate the speech components of a microphone signal from echo and noise components. Machine learning (which generally includes a training phase and an inference phase) is a technique used to improve the ability of a computer system or application to perform a task. During the training phase, the machine learning system is given one or more "answers" along with a large amount of raw training data associated with those answers. The machine learning system analyzes the training data to learn a set of rules that can be used to describe each of the one or more answers. During the inference phase, the machine learning system uses the learned set of rules to infer answers from new data. Unlike linear AEC filters, machine learning models can be trained to account for nonlinear distortion along the echo path.

[0025] Many existing machine learning models for AEC are trained to enhance speech by jointly suppressing the echo and noise components of the microphone signal. However, because the echo and noise components originate from different audio sources, such machine learning models may perform poorly in untrained environments, such as environments with speakers or background audio sources different from those used for training. Aspects of this disclosure recognize that machine learning models can also be used to decouple the echo component of the microphone signal from the noise component of the microphone signal. Therefore, the microphone signal can be decomposed into separate speech, echo, and noise signals (representing the speech, echo, and noise components of the microphone signal, respectively).

[0026] Various aspects generally involve audio signal processing, and more specifically, speech enhancement techniques for separating microphone signals into speech, echo, and noise signals. In some aspects, a speech enhancement system may include a delay estimator and an acoustic echo and noise (AEN) decoupling filter. The delay estimator receives the microphone signal via the microphone. and the far-end audio signal for output via the speaker And based on microphone signals and remote audio signal The delay between them is used to estimate the reference audio signal. In some implementations, the microphone signal... May include speech components echo components and noise components Among them, echo components With reference audio signal Related.

[0027] In some aspects, AEN decoupling filters can be based on microphone signals. and reference audio signal A set of masks is determined, and microphone signals can be suppressed based on the determined set of masks. echo components and noise components In some implementations, this set of masks may include a voice mask. echo mask and noise mask ,in , ,and In some implementations, the AEN decoupling filter may include one trained to be based on microphone signals. and reference audio signal This involves using a neural network to infer multiple outputs. In such implementations, the AEN decoupling filter can determine a mask based on the outputs inferred by the neural network. , and .

[0028] The specific implementations of the subject matter described in this disclosure can be implemented to achieve one or more of the following potential advantages. By using microphone signals... and reference audio signal To generate three masks , and Various aspects of this disclosure can improve the quality of speech in microphone signals. For example, the AEN decoupling filter can mask speech. Application to microphone signals This makes only the speech components The output is sent to a remote device. In some aspects, the AEN decoupling filter can be based on an echo mask. and noise mask The echo signals were respectively and noise signals From microphone signal Further separation. Because of the echo signal. With noise signal Decoupling allows the aspects of this disclosure to suppress acoustic echoes and noise in microphone signals more effectively than existing AEC techniques, even in untrained environments.

[0029] Figure 1 An example hands-free communication system 100 is shown. System 100 includes a set of communication devices 110 and 120 communicatively coupled via a wired or wireless communication channel (not shown for simplicity). More specifically, the first communication device 110 is located in a remote environment (also referred to as a "remote device") and the second communication device 120 is located in a near environment (also referred to as a "near-end device").

[0030] The remote device 110 includes a microphone 112 and a speaker 114. The microphone 112 is configured to detect acoustic waves propagating through the remote environment. Figure 1 In this example, such acoustic waves may include speech 102 from user 101 (also referred to as a "remote user") in a remote environment. Microphone 112 converts the detected acoustic waves into an electrical signal 103 (also referred to as a "remote audio signal") representing the acoustic waveform. Remote device 110 is configured to transmit the remote audio signal 103 to near-end device 120 and receive microphone signals 109 from near-end device 120. Speaker 114 is configured to convert the microphone signals 109 into acoustic waves that can be heard in a remote environment.

[0031] The near-end device 120 includes a speaker 122 and a microphone 124. The speaker 122 is configured to convert a far-end audio signal 103 into acoustic waves 104 that can be heard in a near-end environment. The microphone 124 is configured to detect the acoustic waves propagating through the near-end environment. Figure 1 In this example, such acoustic waves may include acoustic waves 104 (also referred to as “acoustic echoes”) output by speaker 122, speech 106 from user 105 (also referred to as “near-end user”) in the near-end environment, and ambient noise 108 generated by one or more background audio sources 107. Microphone 124 converts the detected acoustic waves into microphone signals 109, which are transmitted to remote device 110.

[0032] Acoustic echo 104 and background noise 108 may mix with and distort the user speech 106 detected by microphone 124. Therefore, microphone signal 109 may include speech components (representing user speech 106), echo components (representing acoustic echo 104), and noise components (representing background noise 108). In some aspects, near-end device 120 may improve the quality of speech in microphone signal 109 by suppressing echo and noise components of microphone signal 109 or otherwise increasing the signal-to-echo ratio (SER) and signal-to-noise ratio (SNR) of microphone signal 109 (also referred to as "speech enhancement"). Thus, microphone signal 109 may include a relatively unchanged copy of speech components, with only minor remnants of echo and noise components (if present).

[0033] Figure 2 A block diagram of an example speech enhancement system 200 according to some implementations is shown. The speech enhancement system 200 is configured to receive microphone signals. and remote audio signal And based on the received audio signal and An enhanced audio signal 201 is generated. More specifically, the voice enhancement system 200 can suppress microphone signals. The acoustic echoes and noise in the sound are used to generate an enhanced audio signal 201.

[0034] In some implementations, microphone signals and remote audio signal They can be respectively Figure 1 Examples of microphone signal 109 and far-end audio signal 103. For example, microphone signal... May include speech components echo components and noise components ,in l It is a frame index and k It is a frequency index associated with the time and frequency domain: For example, refer to Figure 1 Speech components It can represent user voice 106, echo component It can represent acoustic echo 104, and noise components It can represent background noise of 108.

[0035] The speech enhancement system 200 includes a delay estimator 210 and an acoustic echo and noise (AEN) decoupling filter 220. The delay estimator 210 is configured to estimate the microphone signal. and remote audio signal The delay (δ) between them is calculated, and a reference audio signal is generated based on the estimated delay. For reference Figure 1 The acoustic echo 104, detected by microphone 124, represents a delayed version of the far-end audio signal 103. More specifically, the microphone signal... echo components Can be described as a remote audio signal Functions: in This describes the speaker 122 in the microphone signal. The effect of the nonlinear function on, and It is the acoustic transfer function between speaker 122 and microphone 124.

[0036] In some implementations, the delay estimator 210 may estimate the microphone signal based on the generalized cross-correlation phase transform (GCC-PHAT) algorithm. and remote audio signal The delay δ between signals. For example, audio signals. and They can be expressed as time-domain signals respectively. and : in Represents audio signal and The far-end speech component in each audio signal; and They represent audio signals respectively. and Noise components in; α It is related to the second audio signal The associated decay factor; and D It is the first audio signal Second audio signal The delay between (in the time domain).

[0037] All parties involved in this disclosure recognize the time domain delay. D It is possible to calculate the audio signal and cross-correlation And determined: in It is the expected value, and makes The minimized value of τ provides information about the time-domain delay.D (And therefore, the estimation of the delay δ in the time-frequency domain).

[0038] In some implementations, microphone signals and reference audio signal It can be directly passed to the AEN decoupling filter 220. In some other implementations, the speech enhancement system 200 may also include an acoustic echo cancellation (AEC) filter 230, which is configured to be based on the reference audio signal. Reduce microphone signal Acoustic echo in the sound. In some implementations, the AEC filter 230 may rely on a linear transfer function to approximate the impulse response between the speaker 122 and the microphone 124. For example, an adaptive filter (such as the Normalized Least Mean Square (NLMS) algorithm) that models the acoustic coupling (or channel) between the speaker 122 and the microphone 124 may be used to determine the linear transfer function.

[0039] More specifically, the AEC filter 230 can estimate the signal to be emitted from the microphone signal based on the direct path attenuation factor (γ). The subtracted echo. In other words, the AEC filter 230 can adjust the reference audio signal via the direct path attenuation factor γ. To generate signals that can be obtained from the microphone The adjusted reference audio signal that was subtracted : in This represents the filtered microphone signal. As described above, the linear AEC filter cannot account for the nonlinearities introduced along the echo path by the various mechanical components of the amplifier and speaker 124. Therefore, the filtered microphone signal... Some residual echoes can be retained. ,in : .

[0040] The AEN decoupling filter 220 is configured to filter microphone signals. and reference audio signal An enhanced audio signal 201 is generated. For example, in some implementations, the enhanced audio signal 201 may include a microphone signal. or only speech components In some respects, the AEN decoupling filter 220 can make the echo components... or With microphone signal or noise components Decoupling. More specifically, the AEN decoupling filter 220 can decouple microphone signals. or Decomposed into speech components, including microphone signals. The first audio signal, including only the echo component of the microphone signal. or The second audio signal and the noise-only component including the microphone signal. The third audio signal.

[0041] In some implementations, the AEN decoupling filter 220 can use a machine learning (ML) model 222 to decouple the microphone signal. or The audio signal is decomposed into its components. As described above, machine learning generally includes a training phase and an inference phase. During the training phase, the machine learning system is provided with one or more "answers" and a large amount of raw training data associated with those answers. The machine learning system analyzes the training data to learn a set of rules (also known as a "machine learning model") that can be used to describe each of the one or more answers. During the inference phase, the machine learning system can use the learned set of rules to infer answers from new data. Unlike linear AEC filters, machine learning models can be trained to account for nonlinearities along the echo path. In some aspects, the AEN decoupling filter 220 can use the component audio signal to further suppress the filtered microphone signal. Acoustic noise and echoes in the environment.

[0042] Figure 3 A block diagram of an example acoustic echo and noise (AEN) decoupling system 300 according to some implementations is shown. The AEN decoupling system 300 is configured to receive a microphone signal 302 and a reference audio signal 304, and generate a set of component audio signals 322-326 based on the received audio signals 302 and 304. In some implementations, the AEN decoupling system 300 may be... Figure 2 An example of the AEN decoupling filter 220. For example, refer to... Figure 2 Microphone signal 302 can be a microphone signal. or An example of any signal, and the reference audio signal 304 can be the reference audio signal. or An example of any reference audio signal.

[0043] The AEN decoupling system 300 includes a deep neural network (DNN) 310 and a mask generator 320. The DNN 310 is configured to infer multiple (N) outputs 312(1)–312(N) from a microphone signal 302 and a reference audio signal 304 based on the neural network model. Deep learning is a specific form of machine learning in which the inference (and training) phases are performed on multiple layers, resulting in a more abstract dataset in each successive layer. Due to the way it processes information (similar to a biological nervous system), deep learning architectures are often called “artificial neural networks.” For example, each layer of an artificial neural network may consist of one or more “neurons.” Neurons can be interconnected across various layers, allowing input data to be processed and passed from one layer to the next. More specifically, neurons in each layer can perform different transformations on the output data from the previous layer, so that the final output of the neural network leads to the desired inference. The collection of transformations associated with the various layers of the network is called a “neural network model.”

[0044] The mask generator 320 is configured to generate a set of audio masks based on the outputs 312(1)–312(N) of the DNN 310. In some aspects, the set of audio masks may include a speech mask associated with the speech components of the microphone signal 302. Echo mask associated with the echo component of microphone signal 302 Noise mask associated with the noise components of microphone signal 302 Audio mask , and It can be used to decompose the microphone signal 302 into component audio signals 322-326. In some implementations, the AEN decoupling system 300 can mask the voice signal. The microphone signal 302 is applied to obtain the voice signal 322. In some other implementations, the AEN decoupling system 300 can mask the echo. The microphone signal 302 is applied to obtain the echo signal 324. Furthermore, in some implementations, the AEN decoupling system 300 can mask the noise. The microphone signal 302 is applied to obtain the noise signal 326.

[0045] In some implementations, the microphone signal 302 can be Figure 2 microphone signal One example. Referring, for instance, to Equation 1, the voice signal 322 may include a microphone signal. only speech components The echo signal 324 may include a microphone signal. echo components only Furthermore, the noise signal 326 may include a microphone signal. Noise component only ,in: .

[0046] In some other implementations, the microphone signal 302 can be Figure 2 Filtered microphone signal One example. Referring, for instance, to Equation 3, the speech signal 322 may include a filtered microphone signal. only speech components The echo signal 324 may include a filtered microphone signal. Only the residual echo Furthermore, the noise signal 326 may include a filtered microphone signal. Noise component only ,in: .

[0047] In some implementations, the DNN 310 can be trained to infer multiple (M) outputs for each component of the microphone signal 302 (where N = 3 * M). In such implementations, the mask generator 320 can generate an audio mask using the M DNN outputs per group. , and The corresponding audio mask in. In some other implementations, the DNN 310 can be trained to infer multiple (P) outputs for only two components of the microphone signal 302 (where N = 2 * P). In such an implementation, the mask generator 320 can use the DNN outputs 312(1)-312(N) to generate two of the audio masks and can generate a third audio mask (based on the other two audio masks). For example, referring to Equations 3-6, the audio mask , and The sum must be 1: Therefore, audio mask , or Any audio mask can be determined based on the sum of the other two audio masks.

[0048] In some implementations, the mask generator 320 can determine the speech mask based on the DNN outputs 312(1)-312(N). and echo mask Furthermore, it can be further based on voice masks. and echo mask Determine the noise mask : In this type of implementation, DNN 310 can be implemented using a smaller or more compact neural network model (compared to a neural network model trained to infer the output associated with all three audio masks). In other words, for a given neural network size, Equation 8 allows DNN 310 to produce more accurate inference results compared to a neural network model trained to infer the output associated with all three audio masks.

[0049] Figure 4A A block diagram of an example audio mask generation system 400 according to some implementations is shown. The audio mask generation system 400 is configured to output a DNN-based solution. and To generate a set of audio masks , or In some implementations, the audio mask generation system 400 can be... Figure 3 An example of a mask generator 320. See, for example, reference... Figure 3 DNN output and It could be an example of DNN output 312(1)-312(N).

[0050] The audio masking generation system 400 includes a speech masking generation component 402, an echo masking generation component 404, and a noise masking generation component 406. In some implementations, the speech masking generation component 402 may be based on DNN output. and complementary speech mask Generate speech mask ,in and More specifically, the speech mask generation component 402 can be based on the DNN output. , and Determine the amplitude of the speech mask : in It is a smooth approximation (also known as a "soft-addition function") of the rectified linear unit (ReLU) activation function, and It is a function of the sum of positive divisors.

[0051] The speech mask generation component 402 can also be based on DNN output. , and Determine the amplitude of the complementary speech mask : .

[0052] The speech mask generation component 402 can further base its output on the amplitude of the speech mask. Complementary speech mask amplitude and DNN output and Determine the phase of the speech mask ,in: Inverse transform sampling can be used to extract... And calculate To sample and .

[0053] In some implementations, the echo mask generation component 404 can be based on the DNN output. and complementary echo mask Generate echo mask ,in and More specifically, the echo mask generation component 404 can be based on the DNN output. , and Determine the amplitude of the echo mask : .

[0054] Echo mask generation component 404 can also be based on DNN output , and Determine the amplitude of the complementary echo mask : .

[0055] The echo mask generation component 404 can further base its model on the amplitude of the echo mask. Amplitude of complementary echo masks and DNN output and Determine the phase of the echo mask ,in: .

[0056] In some implementations, the noise mask generation component 406 may be based on a speech mask. and echo mask Generate noise mask More specifically, the noise mask generation component 406 can generate a noise mask based on Equation 8. (such as references) Figure 3 (Described).

[0057] For reference Figure 2 The microphone signal described echo components It can be expressed as a far-end audio signal that produces an echo. Functions (such as those shown in Equation 2). Thus, it is also possible to obtain a reference audio signal. or Use masks to derive microphone signals separately or echo components or For example, residual echoes This can be expressed as a reference audio signal that is adjusted. Associated reference mask Functions: .

[0058] By combining Equations 5 and 9, the echo mask... It can be expressed as a reference mask. Functions: in It is the reference audio signal for each frequency point. With microphone signal The ratio. Therefore, the ratio It must be defined between 0 and 1. In some implementations, the ratio... It can be limited (e.g., for values ​​less than 0 and for values ​​greater than 1). In some other implementations, the ratio... It can be normalized, where: .

[0059] All aspects of this disclosure further recognize that: complementary reference masks can be used as a basis. Estimating complementary echo masks ,in: in It is the reference audio signal for each frequency point. With microphone signal The complementary ratio. Therefore, the ratio It must also be defined between 0 and 1. In some implementations, the ratio... It can be limited (e.g., for values ​​less than 0 and for values ​​greater than 1). In some other implementations, the ratio... It can be normalized, where: .

[0060] The example in Equations 9-13 assumes that the speech enhancement system 300 receives filtered microphone signals respectively. and adjusted reference audio signal As Figure 3 The microphone signal 302 and the reference audio signal 304 are used. However, in some other implementations, the voice enhancement system 300 may receive the initial microphone signal separately. and the initial reference audio signal As the microphone signal 302 and the reference audio signal 304. In this implementation, the reference audio signal... Microphone signal and its echo components The filtered microphone signals can be replaced in equations 9-13 respectively. Reference audio signal and residual echo .

[0061] Figure 4B Another block diagram of an example audio mask generation system 410 according to some implementations is shown. The audio mask generation system 410 is configured to use DNN-based outputs. and Generate a set of audio masks , or In some implementations, the audio mask generation system 410 can be... Figure 3 An example of a mask generator 320. See, for example, reference... Figure 3 DNN output and It could be an example of DNN output 312(1)-312(N).

[0062] Apart from Figure 4A In addition to the speech mask generation component 402 and the noise mask generation component 406, the audio mask generation system 400 also includes a reference mask generation component 412 and an echo mask generation component 414. In some implementations, the reference mask generation component 412 may be based on the DNN output. , and Determine the magnitude of the reference mask : in It is a reference audio signal With microphone signal The ratio (as described in reference to Equation 11), and It is a reference audio signal With microphone signal The complementary ratio (as described in Equation 13).

[0063] In some implementations, the reference mask generation component 412 can be further based on the DNN output. , and Determine the magnitude of the complementary reference mask : .

[0064] In some implementations, the echo mask generation component 414 may be based on the amplitude of the reference mask. Magnitude of complementary reference mask and DNN output and Generate echo mask More specifically, the echo mask generation component 414 can be based on the amplitude of the reference mask. and the magnitude of the complementary reference mask Determine the amplitude of the echo mask respectively The amplitude of the complementary echo mask : .

[0065] The echo mask generation component 414 can further base its model on the amplitude of the echo mask. Amplitude of complementary echo masks and DNN output and Determine the phase of the echo mask ,in: .

[0066] Figure 5 Another block diagram of an example speech enhancement system 500 according to some implementations is shown. The speech enhancement system 500 can be configured to generate an enhanced audio signal based on a microphone signal and a reference audio signal. In some implementations, the speech enhancement system 500 may be... Figure 2 An example of a speech enhancement system 200.

[0067] The voice enhancement system 500 includes a device interface 510, a processing system 520, and a memory 530. The device interface 510 is configured to communicate with audio communication devices (such as...) Figure 1The device interface 510 communicates with one or more components of the near-end device 120. In some implementations, the device interface 510 may include a microphone interface (I / F) 512 and a speaker interface (I / F) 514. The microphone interface 512 is configured to receive microphone signals via a microphone (such as microphone 124). The speaker interface 514 is configured to receive far-end audio signals for output via a speaker (such as speaker 122). For example, the speaker interface 514 may receive far-end audio signals from a far-end device (such as far-end device 110).

[0068] Memory 530 may include audio data storage 532, configured to store frames of microphone signals and reference audio signals, as well as any intermediate signals that may be generated by the speech enhancement system 500 as a result of generating enhanced audio signals. Memory 530 may also include non-transitory computer-readable media (including one or more non-volatile memory elements, such as EPROM, EEPROM, flash memory, or hard drives, and other examples), which may store at least the following software (SW) modules: The delay estimation SW module 534 is used to estimate a reference audio signal based on the delay between the microphone signal and the far-end audio signal; A mask generation module SW 536 is used to determine multiple masks based on a microphone signal and a reference audio signal, wherein the multiple masks include a speech mask associated with the speech component of the microphone signal, an echo mask associated with the echo component of the microphone signal, and a noise mask associated with the noise component of the microphone signal; and The speech enhancement module 538 is used to suppress echo and noise components of the microphone signal, at least in part, based on multiple masks. Each software module includes instructions that, when executed by the processing system 520, cause the speech enhancement system 500 to execute a corresponding function.

[0069] Processing system 520 may include any suitable one or more processors capable of executing scripts or instructions of one or more software programs stored in speech enhancement system 500 (such as in memory 530). For example, processing system 520 may execute delay estimation module 534 to estimate a reference audio signal based on the delay between a microphone signal and a far-end audio signal. Processing system 520 may also execute mask generation module 536 to determine multiple masks based on the microphone signal and the reference audio signal, wherein the multiple masks include a speech mask associated with the speech components of the microphone signal, an echo mask associated with the echo components of the microphone signal, and a noise mask associated with the noise components of the microphone signal. Furthermore, processing system 520 may execute speech enhancement module 538 to suppress the echo and noise components of the microphone signal, at least in part, based on the multiple masks.

[0070] Figure 6 An illustrative flowchart depicting an example operation 600 for speech enhancement according to some implementations is shown. In some implementations, example operation 600 may be generated by a speech enhancement system (such as...) Figure 2 and Figure 5 (Executes any of the 200 or 500 voice enhancement systems).

[0071] The speech enhancement system receives a first audio signal (610) via a microphone. The speech enhancement system also receives a second audio signal (620) for output via a speaker. The speech enhancement system estimates a reference audio signal (630) based on the delay between the first and second audio signals. In some aspects, the speech enhancement system may perform an AEC operation on the first audio signal based on the reference audio signal. In some implementations, the AEC operation may be associated with a linear filter.

[0072] The speech enhancement system determines multiple masks based on a first audio signal and a reference audio signal, wherein the multiple masks include a speech mask associated with the speech components of the first audio signal. Echo mask associated with the echo component of the first audio signal and a noise mask associated with the noise components of the first audio signal. (640). The speech component may include audio originating from a near-end user associated with the microphone, the echo component may include audio output by a speaker based on the second audio signal, and the noise component may include audio not originating from a near-end user and not output by a speaker. The speech enhancement system also suppresses the echo and noise components of the first audio signal at least in part based on a plurality of masks (650).

[0073] In some aspects, determining multiple masks may include inferring multiple outputs from a first audio signal and a reference audio signal based on a neural network model. In some implementations, determining multiple masks may further include determining a first subset of the multiple outputs and a complementary speech mask. Estimated speech mask ,in ; based on a second subset of multiple outputs and complementary echo masks Estimated echo mask ,in ; and based on voice masking and echo mask Determine the noise mask .

[0074] In some implementations, voice masking The estimation may include determining the speech mask based on one or more first outputs from a first subset of multiple outputs. The amplitude; determine the complementary speech mask based on one or more first outputs. The amplitude; and based on the speech mask Amplitude, complementary speech mask The amplitude and one or more second outputs from a first subset of multiple outputs determine the speech mask. The phase.

[0075] In some implementations, echo masking The estimation may include determining the echo mask based on one or more first outputs from a second subset of multiple outputs. The amplitude; determine the complementary echo mask based on one or more first outputs. The amplitude; and based on echo masking Amplitude, complementary echo mask The amplitude and one or more second outputs in a second subset of multiple outputs determine the echo mask. The phase.

[0076] In some other implementations, echo masking The estimation may include a second subset based on multiple outputs and a complementary reference mask. Estimate the reference mask associated with the reference audio signal ,in In this implementation, the speech enhancement system can further determine a reference mask based on one or more first outputs from a second subset of multiple outputs. The amplitude; determine the complementary reference mask based on one or more first outputs. The amplitude; based on the reference mask The amplitude determines the echo mask. The amplitude; based on complementary reference masks Amplitude determination of complementary echo mask The amplitude; and based on echo masking Amplitude, complementary echo mask The amplitude and one or more second outputs in a second subset of multiple outputs determine the echo mask. The phase.

[0077] Those skilled in the art will appreciate that any of a wide variety of different technologies and techniques can be used to represent information and signals. For example, data, instructions, commands, information, signals, bits, symbols, and chips that can be referenced throughout the description above can be represented by voltage, current, electromagnetic waves, magnetic fields or magnetic particles, light fields or optical particles, or any combination thereof.

[0078] Furthermore, those skilled in the art will appreciate that the various illustrative logic blocks, modules, circuits, and algorithm steps described in conjunction with the aspects disclosed herein can be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability between hardware and software, various illustrative components, blocks, modules, circuits, and steps have been generally described above in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the system as a whole. Those skilled in the art may implement the described functionality in different ways for each specific application, but such implementation decisions should not be construed as causing a departure from the scope of this disclosure.

[0079] The methods, sequences, or algorithms described in conjunction with the aspects disclosed herein can be implemented directly in hardware, as a software module executed by a processor, or a combination of both. The software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor such that the processor can read information from and write information to the storage medium. Alternatively, the storage medium may be integrated into the processor.

[0080] In the foregoing description, embodiments have been described with reference to specific examples thereof. However, it will be apparent that various modifications and changes may be made therein without departing from the broader scope of this disclosure as set forth in the appended claims. Therefore, this description and the accompanying drawings should be viewed in an illustrative rather than restrictive sense.

Claims

1. A method for speech enhancement, comprising: The first audio signal is received via the microphone; Receives a second audio signal for output via a speaker; The reference audio signal is estimated based on the delay between the first audio signal and the second audio signal; Multiple masks are determined based on the first audio signal and the reference audio signal, the multiple masks including a speech mask associated with the speech components of the first audio signal. Echo mask associated with the echo component of the first audio signal and a noise mask associated with the noise components of the first audio signal. ;as well as The echo component and the noise component of the first audio signal are suppressed at least in part based on the plurality of masks.

2. The method as described in claim 1, wherein, The speech component includes audio originating from a near-end user associated with the microphone, the echo component includes audio output by the speaker based on the second audio signal, and the noise component includes audio not originating from the near-end user and not output by the speaker.

3. The method of claim 1, further comprising: Before determining the plurality of masks, acoustic echo cancellation (AEC) is performed on the first audio signal based on the reference audio signal.

4. The method of claim 3, wherein, The AEC operation is associated with a linear filter.

5. The method of claim 1, wherein, The determination of the plurality of masks includes: Multiple outputs are inferred from the first audio signal and the reference audio signal based on a neural network model.

6. The method of claim 5, wherein, The determination of the plurality of masks further includes: Based on the first subset of the multiple outputs and the complementary speech mask Estimate the speech mask ,in ; Based on the second subset of the multiple outputs and the complementary echo mask Estimate the echo mask ,in ;as well as Based on the voice mask and the echo mask Determine the noise mask .

7. The method of claim 6, wherein, For the voice mask The estimates include: The speech mask is determined based on one or more first outputs from the first subset of the plurality of outputs. The amplitude; The complementary speech mask is determined based on the one or more first outputs. The amplitude; and Based on the voice mask The amplitude and the complementary speech mask The amplitude and one or more second outputs in the first subset of the plurality of outputs determine the speech mask. The phase.

8. The method of claim 6, wherein, For the echo mask The estimates include: The echo mask is determined based on one or more first outputs from the second subset of the plurality of outputs. The amplitude; The complementary echo mask is determined based on the one or more first outputs. The amplitude; and Based on the echo mask The amplitude and the complementary echo mask The amplitude and one or more second outputs in the second subset of the plurality of outputs determine the echo mask. The phase.

9. The method of claim 6, wherein, For the echo mask The estimates include: The second subset based on the multiple outputs and the complementary reference mask Estimate the reference mask associated with the reference audio signal. ,in .

10. The method of claim 9, wherein, For the echo mask The estimate also includes: The reference mask is determined based on one or more first outputs from the second subset of the plurality of outputs. The amplitude; The complementary reference mask is determined based on the one or more first outputs. The amplitude; Based on the reference mask The amplitude determines the echo mask. The amplitude; Based on the complementary reference mask The amplitude determines the complementary echo mask. The amplitude; and Based on the echo mask The amplitude and the complementary echo mask The amplitude and one or more second outputs in the second subset of the plurality of outputs determine the echo mask. The phase.

11. A speech enhancement system, comprising: Processing system; as well as The memory stores instructions that, when executed by the processing system, cause the speech enhancement system to: The first audio signal is received via the microphone; Receives a second audio signal for output via a speaker; The reference audio signal is estimated based on the delay between the first audio signal and the second audio signal; Multiple masks are determined based on the first audio signal and the reference audio signal, the multiple masks including a speech mask associated with the speech components of the first audio signal. Echo mask associated with the echo component of the first audio signal and a noise mask associated with the noise components of the first audio signal. ; as well as The echo component and the noise component of the first audio signal are suppressed at least in part based on the plurality of masks.

12. The speech enhancement system of claim 11, wherein, The speech component includes audio originating from a near-end user associated with the microphone, the echo component includes audio output by the speaker based on the second audio signal, and the noise component includes audio not originating from the near-end user and not output by the speaker.

13. The speech enhancement system of claim 11, wherein, The execution of the instruction also enables the speech enhancement system to: Before determining the plurality of masks, acoustic echo cancellation (AEC) is performed on the first audio signal based on the reference audio signal.

14. The speech enhancement system of claim 13, wherein, The AEC operation is associated with a linear filter.

15. The speech enhancement system of claim 11, wherein, The determination of the plurality of masks includes: Multiple outputs are inferred from the first audio signal and the reference audio signal based on a neural network model.

16. The speech enhancement system of claim 15, wherein, The determination of the plurality of masks further includes: Based on the first subset of the multiple outputs and the complementary speech mask Estimate the speech mask ,in ; Based on the second subset of the multiple outputs and the complementary echo mask Estimate the echo mask ,in ;as well as Based on the voice mask and the echo mask Determine the noise mask .

17. The speech enhancement system of claim 16, wherein, For the voice mask The estimates include: The speech mask is determined based on one or more first outputs from the first subset of the plurality of outputs. The amplitude; The complementary speech mask is determined based on the one or more first outputs. The amplitude; and Based on the voice mask The amplitude and the complementary speech mask The amplitude and one or more second outputs from the first subset of the plurality of outputs determine the speech mask. The phase.

18. The speech enhancement system of claim 16, wherein, For the echo mask The estimates include: The echo mask is determined based on one or more first outputs from the second subset of the plurality of outputs. The amplitude; The complementary echo mask is determined based on the one or more first outputs. The amplitude; and Based on the echo mask The amplitude and the complementary echo mask The amplitude and one or more second outputs in the second subset of the plurality of outputs determine the echo mask. The phase.

19. The speech enhancement system of claim 16, wherein, For the echo mask The estimates include: The second subset based on the multiple outputs and the complementary reference mask Estimate the reference mask associated with the reference audio signal. ,in .

20. The speech enhancement system of claim 19, wherein, For the echo mask The estimate also includes: The reference mask is determined based on one or more first outputs from the second subset of the plurality of outputs. The amplitude; The complementary reference mask is determined based on the one or more first outputs. The amplitude; Based on the reference mask The amplitude determines the echo mask. The amplitude; Based on the complementary reference mask The amplitude determines the complementary echo mask. The amplitude; and Based on the echo mask The amplitude and the complementary echo mask The amplitude and one or more second outputs in the second subset of the plurality of outputs determine the echo mask. The phase.