Unified post-filter for audio filter system of vehicle

By processing audio signals in the short-time Fourier transform domain and utilizing parameterized variants of the unified post-filter and Wiener filter, the problem of residual echo is solved, thereby improving speech quality and noise suppression.

CN121905205APending Publication Date: 2026-04-21GM GLOBAL TECHNOLOGY OPERATIONS LLC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GM GLOBAL TECHNOLOGY OPERATIONS LLC
Filing Date
2024-12-18
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing technologies often result in audible residual echoes during telephone calls or microphone exchanges. Traditional linear acoustic echo cancellation methods are unable to effectively suppress filter misalignment and nonlinear echo components, leading to a decline in voice quality.

Method used

A computer-implemented method is used to process audio signals in the short-time Fourier transform domain through a unified post-filter. Residual echo spectrum estimation is generated using noise smoothing factor, steering vector, and directional and coherent masks. Spectral shaping is performed using a parameterized variant of the Wiener filter to reduce distortion and suppress echoes.

Benefits of technology

It effectively reduces residual echo, improves voice quality, enhances noise suppression, and improves the audio communication experience within the vehicle.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121905205A_ABST
    Figure CN121905205A_ABST
Patent Text Reader

Abstract

A computer-implemented method, when executed by data processing hardware, causes the data processing hardware to perform various operations. These operations include receiving an audio signal from a sensor array at a unified post-filter, converting the audio signal to a short time Fourier transform (STFT) domain via a conversion function, determining a speech presence probability based on the converted audio signal, and determining a noise smoothing factor based on the speech presence probability. The operations further include estimating, via the unified post-filter, a noise power spectral density based on the noise smoothing factor, estimating, via the unified post-filter, a steering vector of the desired source, and generating, via the unified post-filter, a directivity-based mask and a coherence-based mask. Operations further include generating a residual echo spectrum estimate based on the directivity-based mask and the coherence-based mask, and setting one or more spectrum shaping factors based on the residual echo spectrum estimate.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] introduce

[0002] The information provided in this section is intended to generally introduce the background of this disclosure. The work of the inventors listed herein (within the scope disclosed in this section) and aspects of the specification that may otherwise not be considered prior art at the time of filing are neither expressly nor implied to be prior art to this disclosure.

[0003] This disclosure generally relates to a unified post-filter, and more specifically, to a unified post-filter for suppressing in-vehicle noise and residual echo.

[0004] During telephone calls or other microphone exchanges, audible residual echoes may exist. For example, a speaker may hear their own voice after speaking. To suppress echo components, typical linear acoustic echo cancellation first generates an estimate of the echo signal and then subtracts it from the microphone signal. However, residual echoes remain due to filter offset, reverberation, and nonlinear echo components. Therefore, an improved filter is needed to improve speech quality and enhance overall noise suppression by reducing distortion. Summary of the Invention

[0005] In some aspects, the computer-implemented method causes the data processing hardware to perform various operations when executed by the data processing hardware. These operations include: receiving an audio signal from a sensor array at a unified post-filter; converting the audio signal to the short-time Fourier transform (STFT) domain via a transformation function; determining the probability of speech presence based on the converted audio signal; and determining a noise smoothing factor based on the probability of speech presence. The operations also include: estimating the noise power spectral density based on the noise smoothing factor via the unified post-filter; estimating the steering vector of the desired source via the unified post-filter; and generating a directional mask and a coherence mask via the unified post-filter. The operations further include: generating a residual echo spectrum estimate based on the directional mask and the coherence mask; and setting one or more spectral shaping factors based on the residual echo spectrum estimate.

[0006] In some examples, the audio signal may include a desired source, residual ambient noise, and residual echo. Optionally, determining the noise power spectral density may include determining the probability of an active speaker. In some cases, generating a directionality-based mask may include utilizing spatial information and distinguishing the desired source of the audio signal from the residual echo. Additionally or alternatively, generating a coherence-based mask may include masking the estimated echo of the audio signal. In other cases, generating a residual echo spectral estimate may include extracting the residual echo of the audio signal from the original echo of the audio signal. The operation may also include implementing a parameterized variant of the Wiener filter via a unified post-filter.

[0007] On the other hand, an audio filter system for a vehicle includes data processing hardware and memory hardware communicating with the data processing hardware. The memory hardware stores instructions that, when executed on the data processing hardware, cause the data processing hardware to perform operations. These operations include: receiving an audio signal from a sensor array at a unified post-filter; converting the audio signal to the short-time Fourier transform (STFT) domain via a transformation function; determining a speech presence probability based on the converted audio signal; and determining a noise smoothing factor based on the speech presence probability. The operations also include: estimating a noise power spectral density based on the noise smoothing factor via the unified post-filter; estimating a steering vector of a desired source via the unified post-filter; and generating a directional mask and a coherence-based mask via the unified post-filter. The operations further include: generating a residual echo spectrum estimate based on the directional mask and the coherence-based mask; and setting one or more spectral shaping factors based on the residual echo spectrum estimate.

[0008] In some examples, the audio signal may include a desired source, residual ambient noise, and residual echo. Optionally, determining the noise power spectral density may include determining the probability of an active speaker. In some cases, generating a directionality-based mask may include utilizing spatial information and distinguishing the desired source of the audio signal from the residual echo. Additionally or alternatively, generating a coherence-based mask may include masking the estimated echo of the audio signal. In other cases, generating a residual echo spectral estimate may include extracting the residual echo of the audio signal from the original echo of the audio signal. The operation may also include implementing a parameterized variant of the Wiener filter via a unified post-filter.

[0009] In another aspect, an audio filter system for a vehicle includes data processing hardware and memory hardware communicating with the data processing hardware. The memory hardware stores instructions that, when executed on the data processing hardware, cause the data processing hardware to perform operations. These operations include: receiving an audio signal from a sensor array at a unified post-filter; converting the audio signal to the short-time Fourier transform (STFT) domain via a transformation function; determining a speech presence probability based on the converted audio signal; determining a noise smoothing factor based on the speech presence probability; and estimating a noise power spectral density based on the noise smoothing factor via the unified post-filter. The operations also include: estimating a steering vector of a desired source via the unified post-filter; generating a directional mask and a coherence-based mask via the unified post-filter; generating a residual echo spectrum estimate based on the directional mask and the coherence-based mask; setting one or more spectral shaping factors based on the residual echo spectrum estimate; and implementing a parameterized variant of a Wiener filter via the unified post-filter.

[0010] In some examples, the audio signal may include a desired source, residual ambient noise, and residual echo. Optionally, determining the noise power spectral density may include determining the probability of an active speaker. In some cases, generating a directionality-based mask may include utilizing spatial information and distinguishing the desired source of the audio signal from the residual echo. Additionally or alternatively, generating a coherence-based mask may include masking the estimated echo of the audio signal. In other cases, generating a residual echo spectral estimate may include extracting the residual echo of the audio signal from the original echo of the audio signal.

[0011] This disclosure provides the following examples:

[0012] Example 1. A computer-implemented method, when executed by data processing hardware, causes the data processing hardware to perform operations including the following:

[0013] Audio signals are received from the sensor array at a unified post-filter;

[0014] The audio signal is converted to the Short Time Fourier Transform (STFT) domain via a conversion function;

[0015] The probability of speech presence is determined based on the converted audio signal;

[0016] The noise smoothing factor is determined based on the probability of speech presence.

[0017] The noise power spectral density is estimated based on the noise smoothing factor via the unified post-filter;

[0018] The steering vector of the desired source is estimated via the unified post-filter;

[0019] Direction-based masks and coherence-based masks are generated via the unified post-filter;

[0020] Residual echo spectrum estimation is generated based on the directional mask and the coherence mask; and

[0021] One or more spectral shaping factors are set based on the residual echo spectrum estimation.

[0022] Example 2. The method according to Example 1, wherein the audio signal includes the desired source, residual ambient noise, and residual echo.

[0023] Example 3. According to the method of Example 1, wherein determining the noise power spectral density includes determining the probability of the active speaker.

[0024] Example 4. The method according to Example 1, wherein generating a directionality-based mask includes utilizing spatial information and distinguishing the desired source from the residual echo of the audio signal.

[0025] Example 5. The method according to Example 1, wherein generating a coherence-based mask includes masking an estimated echo of the audio signal.

[0026] Example 6. The method according to Example 1, wherein generating the residual echo spectrum estimate includes extracting the residual echo of the audio signal from the original echo of the audio signal.

[0027] Example 7. The method according to Example 1 further includes a parameterized variant of the Wiener filter implemented via the unified post-filter.

[0028] Example 8. An audio filter system for a vehicle, the audio filter system comprising:

[0029] Data processing hardware; and

[0030] Memory hardware that communicates with the data processing hardware, the memory hardware storing instructions that, when executed on the data processing hardware, cause the data processing hardware to perform operations including the following:

[0031] Audio signals are received from the sensor array at a unified post-filter;

[0032] The audio signal is converted to the Short Time Fourier Transform (STFT) domain via a conversion function;

[0033] The probability of speech presence is determined based on the converted audio signal;

[0034] The noise smoothing factor is determined based on the probability of the speech presence.

[0035] The noise power spectral density is estimated based on the noise smoothing factor via the unified post-filter;

[0036] The steering vector of the desired source is estimated via the unified post-filter;

[0037] Direction-based masks and coherence-based masks are generated via the unified post-filter;

[0038] Residual echo spectrum estimation is generated based on the directional mask and the coherence mask; and

[0039] One or more spectral shaping factors are set based on the residual echo spectrum estimation.

[0040] Example 9. The system according to Example 8, wherein the audio signal includes the desired source, residual ambient noise, and residual echo.

[0041] Example 10. The system according to Example 8, wherein determining the noise power spectral density includes determining the probability of the active speaker.

[0042] Example 11. The system according to Example 8, wherein generating a directionality-based mask includes utilizing spatial information and distinguishing the desired source from the residual echo of the audio signal.

[0043] Example 12. The system according to Example 8, wherein generating a coherence-based mask includes masking an estimated echo of the audio signal.

[0044] Example 13. The system according to Example 8, wherein generating residual echo spectrum estimation includes extracting the residual echo of the audio signal from the original echo of the audio signal.

[0045] Example 14. The system according to Example 8 further includes a parameterized variant of the Wiener filter implemented via the unified post-filter.

[0046] Example 15. An audio filter system for a vehicle, the audio filter system comprising:

[0047] Data processing hardware; and

[0048] Memory hardware that communicates with the data processing hardware, the memory hardware storing instructions that, when executed on the data processing hardware, cause the data processing hardware to perform operations including the following:

[0049] Audio signals are received from the sensor array at a unified post-filter;

[0050] The audio signal is converted to the Short Time Fourier Transform (STFT) domain via a conversion function;

[0051] The probability of speech presence is determined based on the converted audio signal;

[0052] The noise smoothing factor is determined based on the probability of the speech presence.

[0053] The noise power spectral density is estimated based on the noise smoothing factor via the unified post-filter;

[0054] The steering vector of the desired source is estimated via the unified post-filter;

[0055] Direction-based masks and coherence-based masks are generated via the unified post-filter;

[0056] Residual echo spectrum estimation is generated based on the directional mask and the coherence mask;

[0057] Based on the residual echo spectrum estimation, one or more spectrum shaping factors are set; and

[0058] A parameterized variant of the Wiener filter is achieved via the unified post-filter.

[0059] Example 16. The system according to Example 15, wherein the audio signal includes the desired source, residual ambient noise, and residual echo.

[0060] Example 17. The system according to Example 15, wherein determining the noise power spectral density includes determining the probability of the active speaker.

[0061] Example 18. The system according to Example 15, wherein generating a directionality-based mask includes utilizing spatial information and distinguishing the desired source from the residual echo of the audio signal.

[0062] Example 19. The system according to Example 15, wherein generating a coherence-based mask includes masking an estimated echo of the audio signal.

[0063] Example 20. The system according to Example 15, wherein generating a residual echo spectrum estimate includes extracting the residual echo of the audio signal from the original echo of the audio signal. Attached Figure Description

[0064] The accompanying drawings described herein are for illustrative purposes only and are not intended to limit the scope of this disclosure.

[0065] Figure 1 This is an exemplary schematic diagram of a vehicle equipped with an audio filter system according to the present disclosure;

[0066] Figure 2 This is an exemplary block diagram of an audio filter system according to the present disclosure;

[0067] Figure 3 This is an example schematic diagram of an audio filter system according to this disclosure; and

[0068] Figure 4 This is an example flowchart of an audio filter system based on this disclosure.

[0069] Throughout the accompanying drawings, the corresponding reference numerals indicate the corresponding parts. Detailed Implementation

[0070] The example configuration will now be described more fully with reference to the accompanying drawings. The example configuration is provided so that this disclosure will be comprehensive and will fully convey the scope of this disclosure to those skilled in the art. Specific details, such as examples of specific components, devices, and methods, are set forth to provide a comprehensive understanding of the configuration of this disclosure. It will be apparent to those skilled in the art that the specific details are not required, the example configuration may be embodied in many different forms, and the specific details and example configuration should not be construed as limiting the scope of this disclosure.

[0071] The terminology used herein is for describing specific exemplary configurations only and is not intended to be limiting. As used herein, the singular articles “a,” “an,” and “the” may also be intended to include plural forms unless the context explicitly indicates otherwise. The terms “comprising,” “including,” and “having” are inclusive and therefore specify the presence of features, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, elements, components, and / or groups thereof. The method steps, processes, and operations described herein should not be construed as necessarily requiring them to be performed in the specific order discussed or illustrated, unless explicitly identified as such. Additional or alternative steps may be employed.

[0072] When an element or layer is referred to as “on another element or layer,” “joined to,” “connected to,” “attached to,” or “coupled to” another element or layer, it may be located directly on, joined to, connected to, attached to, or coupled to that other element or layer, or there may be intermediate elements or layers present. Conversely, when an element is referred to as “directly on another element or layer,” “directly joined to,” “directly connected to,” “directly attached to,” or “directly coupled to” another element or layer, there may be no intermediate elements or layers present. Other terms used to describe relationships between elements should be interpreted in a similar manner (e.g., “between” vs. “directly between,” “adjacent” vs. “directly adjacent,” etc.). As used herein, the term “and / or” includes any and all combinations of one or more of the related listed items.

[0073] The terms “first,” “second,” “third,” etc., are used herein to describe various elements, components, regions, layers, and / or sections. These elements, components, regions, layers, and / or sections should not be limited by these terms. These terms are used only to distinguish one element, component, region, layer, or section from another. Unless the context explicitly indicates otherwise, terms such as “first,” “second,” and other numerical terms do not imply order or sequence. Therefore, the first element, component, region, layer, or section discussed below may be referred to as the second element, component, region, layer, or section without departing from the teachings of the example configuration.

[0074] In this application, the term "module" is replaced by the term "circuit" as defined below. The term "module" may refer to, be part of, or include the following: application-specific integrated circuit (ASIC); digital, analog, or mixed-signal analog / digital discrete circuit; digital, analog, or mixed-signal analog / digital integrated circuit; combinational logic circuit; field-programmable gate array (FPGA); processor (shared, dedicated, or grouped) that executes code; memory (shared, dedicated, or grouped) that stores code executed by the processor; other suitable hardware components that provide the aforementioned functionality; or combinations of some or all of the foregoing, such as in a system-on-a-chip.

[0075] The term "code" as used above can include software, firmware, and / or microcode, and can refer to programs, routines, functions, classes, and / or objects. The term "shared processor" covers a single processor that executes some or all of the code from multiple modules. The term "group processor" covers a processor that, in combination with additional processors, executes some or all of the code from one or more modules. The term "shared memory" covers a single memory that stores some or all of the code from multiple modules. The term "group memory" covers memory that, in combination with additional memory, stores some or all of the code from one or more modules. The term "memory" can be a subset of the term "computer-readable medium." The term "computer-readable medium" does not include transient electrical and electromagnetic signals propagating through a medium, and therefore can be considered tangible and non-transient memory. Non-limiting examples of non-transient memory include tangible computer-readable media, including non-volatile memory, magnetic storage devices, and optical storage devices.

[0076] The apparatus and methods described in this application may be implemented, in part or in whole, by one or more computer programs executed by one or more processors. The computer programs include processor-executable instructions stored on at least one non-transitory tangible computer-readable medium. The computer programs may also include and / or depend on stored data.

[0077] A software application (i.e., a software resource) can refer to computer software that instructs a computing device to perform a task. In some examples, a software application may be referred to as an "application," "app," or "program." Example applications include, but are not limited to, system diagnostic applications, system management applications, system maintenance applications, word processing applications, spreadsheet applications, messaging applications, media streaming applications, social networking applications, and game applications.

[0078] Non-transient memory can be a physical device used for temporary or permanent storage of programs (e.g., instruction sequences) or data (e.g., program state information) for use by a computing device. Non-transient memory can be volatile and / or non-volatile addressable semiconductor memory. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electronically erasable programmable read-only memory (EEPROM) (e.g., commonly used in firmware, such as bootloaders). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase-change memory (PCM), and magnetic disks or magnetic tapes.

[0079] These computer programs (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor and can be implemented using high-level procedural and / or object-oriented programming languages ​​and / or assembly / machine languages. As used herein, the terms “machine-readable medium” and “computer-readable medium” refer to any computer program product, non-transitory computer-readable medium, apparatus, and / or device (e.g., disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0080] Various implementations of the systems and techniques described herein can be implemented in digital electronic and / or optical circuits, integrated circuits, specially designed ASICs (Application-Specific Integrated Circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be dedicated or general-purpose, coupled to receive data and instructions from a storage system, at least one input device, and at least one output device, and to transmit data and instructions to the storage system, at least one input device, and at least one output device.

[0081] The processes and logic described in this specification can be executed by one or more programmable processors (also known as data processing hardware) that execute one or more computer programs to perform functions by manipulating input data and generating output. The processes and logic can also be executed by special-purpose logic circuitry, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits). Processors suitable for executing computer programs include, for example, both general-purpose and special-purpose microprocessors, and any one or more processors of any kind of digital computer. Typically, the processor receives instructions and data from read-only memory or random access memory, or both. The basic elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, or operably coupled to receive data from or transfer data to, or both, such as magnetic disks, magneto-optical disks, or optical disks. However, a computer does not necessarily need to have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROMs and DVD-ROMs. Processors and memory may be supplemented by or incorporated into dedicated logic circuitry.

[0082] To provide interaction with a user, one or more aspects of this disclosure can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube), LCD (liquid crystal display) monitor, or touchscreen) for displaying information to the user and optionally having a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback, such as visual, auditory, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. Furthermore, the computer can interact with the user by sending documents to and receiving documents from the device used by the user; for example, by sending web pages to a web browser on the user's client device in response to a request received from a web browser.

[0083] refer to Figure 1-3The audio filter system 10 for a vehicle 100 according to this disclosure includes an electronic control unit (ECU) 12 configured with a uniform post-filter 14. It is also contemplated that, in other examples, the audio filter system 10 may be used in computer systems other than those relating to a vehicle. These examples include, but are not limited to, mobile devices, headsets, headphones, speaker systems, and any other feasible devices utilizing voice-to-voice and / or voice-to-machine (ASR) communication via an audio system. For illustrative purposes, the audio filter system 10 is described herein with respect to a vehicle 100.

[0084] Audio filter system 10 may be electrically coupled to sensor array 102 of vehicle 100 to receive audio signal 16. In some examples, sensor array 102 may include a microphone array and / or speakers within vehicle 100, configured to at least partially capture audio signal 16. Digital signal 17 may also be processed by audio filter system 10 and is generally associated with audio signal 16. Audio signal 16 includes desired source 16a, perceived ambient noise 16b, and perceived echo 16c. After audio signal 16 passes through acoustic echo canceller (AEC) and beamformer, the remaining signal includes residual ambient noise 18a and residual echo 18b, as described below. For example, input 46 of unified post-filter 14 can be represented by the following equation:

[0085] (1)x(k,n)=s(k,n)+r(k,n)+v(k,n)

[0086] Where x(k,n) is the audio signal 16 after linear processing 46; s(k,n) is the desired source 16a; r(k,n) is the residual echo 18b; and v(k,n) is the residual ambient noise 18a.

[0087] As described herein, the audio filter system 10 coordinates the ECU 12, which has a unified post-filter 14, to monitor the audio signal 16 and adjust the audio output from the audio filter system 10. The ECU 12 includes data processing hardware 20 and memory hardware 22, which stores operations and instructions of the audio filter system 10 that can be executed by the data processing hardware 20. In some examples, the unified post-filter 14 may be configured as part of the ECU 12. In other examples, the unified post-filter 14 may be separate from the ECU 12 but may communicate with it.

[0088] As described herein, the unified post-filter 14 is configured to isolate the desired source 16a from residual ambient noise 18b and residual echo 18b. For example, the unified post-filter 14 is configured to identify a linear operator 24 that estimates the desired source 16a using a mean square error (MSE) 26. The unified post-filter 14 is configured to determine the MSE 26 of the difference between the desired source 16a and the audio signal 16 received as a whole. For example, the linear operator 24 can be determined using the following equation:

[0089] (2)G opt (k, n)argmin G E[(s(k,n)-G(k,n)x(k,n)) 2 ]

[0090] Among them, G opt (k,n) is a linear operator 24; argmin G Find the optimal post-filter to apply on x; E is the expectation operator; s(k,n) is the expectation source 16a; G(k,n) is the post-filter that approximates the expectation source (s) applied on (x); and x(k,n) is the single-channel input signal of the post-filter.

[0091] In one example, the audio filter system 10 utilizes linear operator 24 to define a modified Wiener filter 28. For instance, the audio filter system 10 may introduce a parameterized variant 28a of the Wiener filter 28, which can be used to generate and refine the unified post-filter 14. The parameterized variant 28a of the Wiener filter 28 utilizes parameters 30, including a spectral shaping factor 30a, an audio emphasis factor 30b, and an overestimation factor 30c for residual echo 18b. These parameters 30 are pre-tuned to optimize speech quality and suppress interference. For example, the parameterized variant 28a of the Wiener filter 28 can be expressed by the following equation:

[0092]

[0093] in It is a parameterized variant 28a; α(k) is a spectral shaping factor 30a; β(k) is an audio emphasis factor 30b; and γ(k) is an overestimation factor 30c of residual echo 18b.

[0094] Parametric variant 28a also utilizes the unified parameter 32 of the final unified post-filter 14. The unified parameter 32 is the autocorrelation of the signal and includes the near speaker 32a with respect to the desired value (in equation (3) Φ). ss (k,n) represents the residual environmental noise 16b (in equation (3) denoted by Φ). vv(k,n) represents) and residual echo 32c (in equation (3) denoted by Φ rr (k,n) represents the information. The audio filter system 10 estimates each parameter 30 and can manipulate the parameters 30 to improve the overall performance of the unified post-filter 14. Specifically, the audio filter system 10 can manipulate or otherwise modify the unified parameter 32 to refine the parameterized variant 28a and improve the unified post-filter 14. For example, the unified post-filter 14 is configured to take into account all unified parameters 32, which ultimately improves the filtering of the received audio signal 16.

[0095] Before estimating the uniformity parameter 32, the audio filter system 10 uses a transformation function 34 to transform the audio signal 16 to the short-time Fourier transform (STFT) domain 36. The audio filter system 10 can use the STFT domain 36 to compute or otherwise generate a noise spectrum estimate 38. The noise spectrum estimate 38 is related to the ambient noise 16b. The audio filter system 10 can also use the transformed audio signal 16 from the STFT domain 36 to determine the speech presence probability 40. The audio filter system 10 uses a scale from zero (0) to one (1) when performing the speech presence probability 40, such that the speech presence probability 40 is a value between zero (0) and one (1). For example, the audio filter system 10 uses the speech presence probability 40 to detect whether a speaker is active. If the probability of speaker activity is between zero (0) and one (1), the audio filter system 10 will not estimate the uniformity parameter 32 because the audio signal 16 will include both the ambient noise 16b and the speech.

[0096] In addition to the speech presence probability 40, the noise spectrum estimation 38 is also based on a predefined noise smoothing factor 42, which can be stored in the memory hardware 22. The audio filter system 10 uses the predefined noise smoothing factor 42 to calculate the estimated noise smoothing factor 42a. For example, the estimated noise smoothing factor 42a can be calculated using the following equation:

[0097]

[0098] in It is the estimated noise smoothing factor 42a; λ v (k) is a predefined noise smoothing factor 42; η(k,n) is the speech presence probability 40, and obtains a value between zero (0) and one (1). The audio filter system 10 uses the estimated noise smoothing factor 42a to smooth the ambient noise 16b and the audio signal 16 itself. The audio filter system 10 can be configured to remove any desired source 16a from the audio signal 16 when estimating the noise smoothing factor 42a and determining the speech presence probability 40, in order to focus on the ambient noise 16b and / or residual ambient noise 18a in the audio signal 16.

[0099] Once the audio filter system 10 has determined or otherwise estimated the estimated noise smoothing factor 42a, it can estimate the noise power spectral density 44. The noise power spectral density 44 can be used to determine the active speaker probability 40a of the speech presence probability 40. The active speaker probability 40a can be estimated, for example, using the following equation:

[0100]

[0101] in The residual environmental noise power spectral density is 44. λ is the estimated noise smoothing factor 42a; x(k,n) is the single-channel input signal of the post-filter 46; v (k) is a predefined noise smoothing factor of 42; Φ v (k, n) is the ambient noise 16b; and η(k, n) is the speech presence probability 40. Smoothing is performed between the audio signal 16 in the frame and the previous estimate. The audio filter system 10 can identify the active speaker based on the speech presence probability 40. As mentioned above, the speech presence probability 40 can vary between zero (0) and one (1), such that if the speech presence probability 40 is closer to one (1), the noise is not updated and the previous estimate is used. If the speech presence probability 40 is close to zero (0), the audio filter system 10 can update the noise estimate. Therefore, noise is estimated in the absence of speech, which means that when the speech presence probability 40 is close to one (1), the audio filter system 10 relies on the previous noise estimate.

[0102] Ambient noise 16b and / or residual ambient noise 18a are estimated in the complexity domain 48. The complexity domain 48 includes the microphone's spectral power and phase power, and is typically related to the square of the channel input signal 46, as described in more detail below. To obtain the residual echo spectrum estimate 52, the audio filter system 10 generates a directionality-based mask 54 and a coherence-based mask 56, as described in more detail below. The residual echo spectrum estimate 52 can be determined using the following equation:

[0103]

[0104] in It is a residual echo spectrum estimate 52; M d (k, n) is a directional mask 54; M C (k, n) is a coherence-based mask 56; and It is the estimated echo, which is estimated in the acoustic echo canceller module as part of linear acoustic echo cancellation.

[0105] The directional mask 54 uses spatial information to distinguish the desired source 16a from the residual echo 18b. For example, the directional mask 54 is associated with the spectral domain 48, which informs the audio filter system 10 of the environmental domain of the vehicle 100. The spectral domain 48 provides environmental information to the audio filter system 10. In some cases, the directional mask 54 can be determined at least partially by the audio filter system 10 using a transient beam 58 from the spectral domain 48. The transient beam 58 can be determined using the following equation:

[0106]

[0107] Where ψ(k, n) is the instantaneous beam 58; i is the index for searching the maximum number of available speakers; x(k, n) is the input signal of the unified post-filter 14 (i.e., the output of the beamformer module); I is the identity matrix; h i It is the guiding vector 50; and e(k,n) is the component vector 60.

[0108] The unified post-filter 14 receives an estimated steering vector 50 from the beamforming module. The steering vector 50 is used to generate a directional mask 54. The unified post-filter 14 is used to estimate the components of the desired source 16a in the vehicle 100. This component can be a relatively conservative function 62 or an acoustic function 64. The steering vector 50 represents knowledge of the desired source 16a within the vehicle 100. For example, the audio filter system 10 can utilize the spatial information 104 of the vehicle 100 in distinguishing the desired source 16a from the residual echo 18b of the audio signal 16. If there are more than one (1) speaker, the audio filter system 10 can use a blocking matrix to block the speakers and identify the desired source 16a. For example, the audio filter system 10 can use the component vector 60 to compare with the steering vector 50 to obtain a value.

[0109] Component vector 60 is a sensor array 102 without linear echo. Therefore, the audio filter system 10 can examine the relationship between component vector 60 and steering vector 50 using the received audio signal 16. If only a near-end signal is detected, the audio signal 16 will become zero (0). If residual echo 18b is detected, the audio filter system 10 can determine the relationship between residual echo 18b and the value. Although the value may change over time, component vector 60 is related to information concerning the desired source 16a and also contains some information related to residual echo 18b.

[0110] Referring again to equation (7), the audio filter system 10 can utilize or extract the signal-to-noise ratio (SNR) level to estimate the possible direction of the desired source 16a. For example, if a near-end signal is present, the SNR level may be high, while if no near-end signal is present, the SNR level will be low. The audio filter system 10 can use the SNR level to determine the presence of a near-end signal, which can help the audio filter system 10 attenuate the audio signal 16. For example, if the audio filter system 10 can identify the presence of a near-end signal, the audio filter system 10 can attenuate the audio signal 16 to estimate only the residual echo 18b. As a result, the audio filter system 10 generates a directionality-based mask 54. For example, a directionality-based mask can be derived from the following equation:

[0111]

[0112] Where M d (k, n) is a directional mask 54; <ψ(k, n)> r ψ(k,n) is the average value of the instantaneous beam 58; and ψ(k,n) is the instantaneous beam 58. If the SNR level is high, the audio filter system 10 infers that the directional mask 54 will be close to zero (0). In determining the directional mask 54, the audio filter system 10 checks the instantaneous beam 58 against the average value of the instantaneous beam 58. The average value of the instantaneous beam 58 is determined using the active echo 18b. If only the residual echo 18b is present, the directional mask 54 will be close to one (1). If only the near-end signal is present, the directional mask 54 will be close to zero (0) because the instantaneous beam 58 will be high and the average value is preserved. Preserving the average value means that the average value is calculated only based on the active reference beam, excluding the near-end beam.

[0113] The audio filter system 10 also generates a coherence-based mask 56, which is determined by comparing the coherence 70 between the estimated echo 72 and the input at the sensor array 102. The coherence-based mask 56 is a complementary mask to the directional mask 54, allowing the audio filter system 10 to utilize both the directional mask 54 and the coherence-based mask 56. The coherence-based mask 56 provides an indication of a frequency beam with a high probability of echo presence. The audio filter system 10 measures the correlation 70 between the audio signal 16 and the estimated echo 72, estimated using linear echo cancellation 74. If the correlation is high, the audio filter system 10 will have an indication of a frequency box with echoes, and may also be able to indicate the presence of residual echo 18b. Therefore, the audio filter system 10 can use linear echo cancellation 74 to attenuate the echoes to eliminate the echoes and / or residual echo 18b. The coherence-based mask can be derived from the following exemplary equation:

[0114]

[0115]

[0116]

[0117]

[0118] In equation (9), μ(k,n) is the coherence 70; E is the expectation operator; and d(k,n) is the audio signal. It is the estimated echo 72; It is the estimated variance of audio signal 16; and This is the estimated variance of the echo 72. In equation (10), ρ(k,n) is the result of subtracting the estimated noise floor from the spectrum of the input spectrum of the unified post-filter; x(k,n) is the input signal of the unified post-filter; and The noise power spectral density is 44. In equation (11), The naive estimation of the residual echo 18b based on the input signal x(k,n)58; μ(k,n) represents the coherence 70. For the naive residual echo estimation from the previous time frame; in equation (11), M C (k, n) is a coherence-based mask 56; It is a naive estimate of residual echoes.58; It is the instantaneous beam 58, which is the estimated echo from the acoustic echo canceller; and ∈ is a small number.

[0119] The audio filter system 10 utilizes a directional mask 54, a coherence mask 56, and a uniform post-filter 14 to generate a residual echo spectrum estimate 52. The residual echo spectrum estimate 52 can be determined by multiplying the masks 54 and 56 by the power of the residual echo 18b to mask the estimated echo 72. Therefore, the audio filter system 10 can mask the residual echo 18b by masking the estimated echo 72. As a result, the audio filter system 10 can set one or more spectrum shaping factors 30a based on the residual echo spectrum estimate 52 to extract the residual echo 18b of the audio signal 16 from the original echo 16d of the audio signal 16. The directional mask 54 further enhances the residual echo estimate 18b by cleaning up the estimate based on the presence of near-end signals. For example, the audio filter system 10 attenuates the noise power spectral density 44 and extracts the residual echo 18b from the estimated echo 72. When the coherence-based mask 56 is used to equalize the power of the estimated echo 72 from or from the residual echo 18b, the directivity-based mask 54 further enhances and cleans up the near-end presence of the estimated echo 72.

[0120] For details, please refer to the following: Figure 4 The diagram illustrates an example method 400 for an audio filter system 10. At 402, a unified post-filter 14 receives an audio signal 16 from a sensor array 102. At 404, the audio signal 16 is converted to the STFT domain 36 via a conversion function 34. At 406, the audio filter system 10 determines a speech presence probability 40 based on the converted audio signal 16, and at 408 determines a noise smoothing factor 42 based on the speech presence probability 40. The unified post-filter 14 estimates a noise power spectral density 44 based on the noise smoothing factor 42 at 410 and estimates a steering vector 50 of the desired source 16a at 412. The unified post-filter 14 generates a directionality-based mask 54 and a coherence-based mask 56 at 414, and generates a residual echo spectrum estimate 52 based on the directionality-based mask 54 and the coherence-based mask 56 at 416. The audio filter system 10 sets one or more spectral shaping factors 30a at 416 based on the residual echo spectrum estimate 52, and implements a parameterized variant 28a of the Wiener filter 28 at 418 via a unified post-filter 14.

[0121] Many implementations have been described. However, it should be understood that various modifications can be made without departing from the spirit and scope of this disclosure. Therefore, other implementations are within the scope of the following claims.

[0122] The foregoing description has been provided for illustrative and descriptive purposes. It is not intended to be exhaustive or limiting of this disclosure. Individual elements or features of a particular configuration are generally not limited to that particular configuration, but are interchangeable where applicable and can be used in the selected configuration, even if not specifically shown or described. They can also be varied in many ways. These variations should not be considered as departing from this disclosure, and all such modifications are intended to be included within the scope of this disclosure.

Claims

1. A computer-implemented method, when executed by data processing hardware, causing the data processing hardware to perform operations including the following: Audio signals are received from the sensor array at a unified post-filter; The audio signal is converted to the Short Time Fourier Transform (STFT) domain via a conversion function; The probability of speech presence is determined based on the converted audio signal; The noise smoothing factor is determined based on the probability of speech presence. The noise power spectral density is estimated based on the noise smoothing factor via the unified post-filter; The steering vector of the desired source is estimated via the unified post-filter; Direction-based masks and coherence-based masks are generated via the unified post-filter; Residual echo spectrum estimation is generated based on the directional mask and the coherence mask; as well as One or more spectral shaping factors are set based on the residual echo spectrum estimation.

2. The method of claim 1, wherein the audio signal comprises the desired source, residual ambient noise, and residual echo.

3. The method of claim 1, wherein determining the noise power spectral density includes determining the probability of the active speaker.

4. The method of claim 1, wherein generating a directionality-based mask includes utilizing spatial information and distinguishing the desired source from the residual echo of the audio signal.

5. The method of claim 1, wherein generating a coherence-based mask includes masking an estimated echo of the audio signal.

6. The method of claim 1, wherein generating the residual echo spectrum estimate comprises extracting the residual echo of the audio signal from the original echo of the audio signal.

7. The method of claim 1, further comprising implementing a parameterized variant of the Wiener filter via the unified post-filter.

8. An audio filter system for a vehicle, the audio filter system comprising: Data processing hardware; and Memory hardware that communicates with the data processing hardware, the memory hardware storing instructions that, when executed on the data processing hardware, cause the data processing hardware to perform operations including the following: Audio signals are received from the sensor array at a unified post-filter; The audio signal is converted to the Short Time Fourier Transform (STFT) domain via a conversion function; The probability of speech presence is determined based on the converted audio signal; The noise smoothing factor is determined based on the probability of the speech presence. The noise power spectral density is estimated based on the noise smoothing factor via the unified post-filter; The steering vector of the desired source is estimated via the unified post-filter; Direction-based masks and coherence-based masks are generated via the unified post-filter; Residual echo spectrum estimation is generated based on the directional mask and the coherence mask; as well as One or more spectral shaping factors are set based on the residual echo spectrum estimation.

9. The system of claim 8, wherein the audio signal includes the desired source, residual ambient noise, and residual echo.

10. The system of claim 8, wherein determining the noise power spectral density includes determining the probability of the active speaker.