Echo cancellation method and device, electronic equipment, storage medium and vehicle
By adaptively selecting the target reference channel and using a combination of foreground and background filters, the problems of large computational complexity and poor echo cancellation effect in the existing technology are solved, and efficient echo cancellation is achieved in complex sound fields.
Patent Information
- Application Number
- CN202410323262.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-20
- Publication Date
- 2025-09-23
AI Technical Summary
In scenarios with strong reverberation, large echoes, and complex echo paths, the existing fixed downmixing method cannot effectively leverage the importance of different reference channels, resulting in large computational complexity and poor echo cancellation effect.
By determining the instantaneous energy and smoothed energy of multiple reference channels corresponding to the current frame reference signal, the appropriate target reference channel is selected for adaptive downmixing, and the estimated echo signal is generated using foreground and background filters for flexible echo cancellation.
Flexible and adaptive echo cancellation is achieved under different sound fields and sound effects, which reduces the overall calculation amount and improves the accuracy and efficiency of echo cancellation.
Smart Images

Figure CN120690218A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of automobile technology, and in particular to an echo cancellation method, device, electronic device, storage medium, and vehicle. Background Art
[0002] Acoustic Echo Cancellation (AEC) is based on the principle of echo cancellation. It uses an adaptive filter to generate a noise signal with equal amplitude and opposite phase to the echo, achieving echo cancellation. Due to the large number of reference channels, this requires significant computation. Downmixing the reference channels can effectively reduce this computational complexity. A common approach is to set a fixed downmix mode based on experience, reducing the number of reference signal channels. The filter then adaptively adjusts based on the reference signal input, generating a canceling echo signal. This signal cancels the echo signal of the near-end reference signal, eliminating the far-end echo signal and delivering a pure vocal output.
[0003] However, for scenes with strong reverberation, large echoes, and complex echo paths, using a fixed downmixing method for all sound fields and sound effects is not conducive to leveraging the importance of different reference channels. A downmixing method that is more flexible and adaptable to sound fields and sound effects is needed. Summary of the Invention
[0004] The present application provides an echo cancellation method, device, electronic device, storage medium and vehicle, which can flexibly adapt to the characteristics of different sound fields and sound effects to adaptively downmix the reference signal, which is conducive to giving full play to the importance of different reference channels and can effectively reduce the overall computational complexity.
[0005] In a first aspect, an embodiment of the present application provides an echo cancellation method, comprising: determining the instantaneous energy of each reference channel in a plurality of reference channels corresponding to a current frame reference signal; determining the smoothed energy of the corresponding reference channel based on the instantaneous energy of each reference channel; determining a plurality of target reference channels from the plurality of reference channels based on the smoothed energy of each reference channel, wherein the smoothed energy corresponding to each target reference channel is greater than or equal to an energy threshold; downmixing the current frame reference signal based on the plurality of target reference channels to obtain downmixed reference signals corresponding to the plurality of target reference channels; inputting the downmixed reference signal into a filter module to generate an estimated echo signal corresponding to the current frame reference signal; and performing echo cancellation on a near-end signal received by a microphone based on the estimated echo signal to obtain a target audio signal in the near-end signal. In some embodiments of the present application, based on the smoothed energy of each reference channel, multiple target reference channels are determined from the multiple reference channels, including: among the multiple reference channels, the reference channels whose corresponding smoothed energy is greater than or equal to the energy threshold are respectively determined as candidate reference channels to obtain multiple candidate reference channels; when the number of channels of the multiple candidate reference channels is less than or equal to the number of target channels, the multiple candidate reference channels are determined as the multiple target reference channels; when the number of channels of the multiple candidate reference channels is greater than the number of target channels, the multiple target reference channels are determined from the multiple candidate reference channels based on a priori rules, and the number of channels of the multiple target reference channels is equal to the number of target channels.
[0006] In some embodiments of the present application, the prior rule includes at least one of the following: a candidate reference channel with a high priority is the target reference channel; a candidate reference channel in a reference channel merging combination is the target reference channel, and different reference channel merging combinations include different reference channels among multiple reference channels.
[0007] In some embodiments of the present application, the filter module includes a foreground filter and a background filter, and the background filter is an adaptive filter; the downmixed reference signal is input into the filter module to generate an estimated echo signal corresponding to the current frame reference signal, including: inputting the downmixed reference signal into the foreground filter to obtain a first estimated echo signal; inputting the downmixed reference signal into the background filter to obtain a second estimated echo signal; based on the estimated echo signal, performing echo cancellation on the near-end signal received by the microphone to obtain a target audio signal in the near-end signal, including: performing echo cancellation on the near-end signal based on the first estimated echo signal to obtain a first error signal; performing echo cancellation on the near-end signal based on the second estimated echo signal to obtain a second error signal; and determining the smaller value of the first error signal and the second error signal as the target audio signal.
[0008] In some embodiments of the present application, after performing echo cancellation on the near-end signal based on the second estimated echo signal to obtain a second error signal, the method further includes: when the first error signal is smaller than the second error signal, updating the filter coefficient of the background filter to the filter coefficient of the foreground filter; when the first error signal is greater than the second error signal, updating the filter coefficient of the foreground filter to the filter coefficient of the background filter.
[0009] In some embodiments of the present application, after downmixing the current frame reference signal based on the multiple target reference channels to obtain the downmixed reference signals corresponding to the multiple target reference channels, and before inputting the downmixed reference signal into the background filter to obtain the second estimated echo signal, the method further includes: recording a downmixing mode label of the current frame reference signal, where the downmixing mode label is used to indicate the downmixing mode of the multiple reference channels downmixed into the multiple target reference channels; and adaptively updating the filter coefficients of the background filter based on the downmixing mode label of the current frame reference signal when the downmixing mode label of the current frame reference signal is different from the downmixing mode label of the previous frame reference signal.
[0010] In some embodiments of the present application, inputting the downmixed reference signal into the background filter to obtain a second estimated echo signal includes: when the downmixing mode label of the current frame reference signal is different from the downmixing mode label of the previous frame reference signal, in the process of adaptively filtering the downmixed reference signal through the background filter, adaptively adjusting the iterative step size of the background filter based on the downmixing mode label of the current frame reference signal.
[0011] In a second aspect, an embodiment of the present application provides an echo cancellation device, comprising: a determination module for determining the instantaneous energy of each reference channel in a plurality of reference channels corresponding to a current frame reference signal; based on the instantaneous energy of each reference channel, determining the smoothed energy of the corresponding reference channel; based on the smoothed energy of each reference channel, determining a plurality of target reference channels from the plurality of reference channels, the smoothed energy corresponding to each target reference channel being greater than or equal to an energy threshold; a downmixing module for downmixing the current frame reference signal based on the plurality of target reference channels to obtain downmixed reference signals corresponding to the plurality of target reference channels; a generation module for inputting the downmixed reference signal into a filter module to generate an estimated echo signal corresponding to the current frame reference signal; and an echo cancellation module for performing echo cancellation on a near-end signal received by a microphone based on the estimated echo signal to obtain a target audio signal in the near-end signal.
[0012] In some embodiments of the present application, the determination module is specifically used to determine, among the multiple reference channels, the reference channels whose corresponding smoothed energy is greater than or equal to the energy threshold as candidate reference channels, to obtain multiple candidate reference channels; when the number of channels of the multiple candidate reference channels is less than or equal to the number of target channels, determine the multiple candidate reference channels as the multiple target reference channels; when the number of channels of the multiple candidate reference channels is greater than the number of target channels, determine the multiple target reference channels from the multiple candidate reference channels based on a priori rules, and the number of channels of the multiple target reference channels is equal to the number of target channels.
[0013] In some embodiments of the present application, the prior rule includes at least one of the following: a candidate reference channel with a high priority is the target reference channel; a candidate reference channel in a reference channel merging combination is the target reference channel, and different reference channel merging combinations include different reference channels among multiple reference channels.
[0014] In some embodiments of the present application, the filter module includes a foreground filter and a background filter, and the background filter is an adaptive filter; the generation module is specifically configured to input the downmixed reference signal into the foreground filter to obtain a first estimated echo signal; and input the downmixed reference signal into the background filter to obtain a second estimated echo signal;
[0015] The echo cancellation module is specifically configured to perform echo cancellation on the near-end signal based on the first estimated echo signal to obtain a first error signal; perform echo cancellation on the near-end signal based on the second estimated echo signal to obtain a second error signal; and determine the smaller value of the first error signal and the second error signal as the target audio signal.
[0016] In some embodiments of the present application, the device also includes: an updating module, which is used to update the filter coefficients of the background filter to the filter coefficients of the foreground filter when the first error signal is less than the second error signal after performing echo cancellation on the near-end signal based on the second estimated echo signal to obtain a second error signal; and when the first error signal is greater than the second error signal, update the filter coefficients of the foreground filter to the filter coefficients of the background filter.
[0017] In some embodiments of the present application, the device also includes: a recording module for recording a downmixing mode label of the current frame reference signal after downmixing the current frame reference signal based on the multiple target reference channels to obtain downmixed reference signals corresponding to the multiple target reference channels and before inputting the downmixed reference signal into the background filter to obtain a second estimated echo signal, wherein the downmixing mode label is used to indicate the downmixing mode of the multiple reference channels being downmixed into the multiple target reference channels; and an updating module for adaptively updating the filter coefficients of the background filter based on the downmixing mode label of the current frame reference signal when the downmixing mode label of the current frame reference signal is different from the downmixing mode label of the previous frame reference signal.
[0018] In some embodiments of the present application, the echo cancellation module is specifically configured to adaptively adjust the iterative step size of the background filter based on the downmixing mode label of the current frame reference signal when the downmixing mode label of the current frame reference signal is different from the downmixing mode label of the previous frame reference signal.
[0019] In a third aspect, an embodiment of the present application provides an electronic device, comprising: a processor, wherein the processor is configured to execute a computer program stored in a memory, wherein when the computer program is executed by the processor, the steps of any one of the echo cancellation methods provided in the first aspect are implemented.
[0020] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any one of the echo cancellation methods provided in the first aspect.
[0021] In a fifth aspect, an embodiment of the present application provides a vehicle, comprising: the echo cancellation device as described in the second aspect, or the electronic device as described in the third aspect, or the computer-readable storage medium as described in the fourth aspect.
[0022] In a sixth aspect, an embodiment of the present application provides a computer program product, wherein the computer program product includes a computer program or instructions. When the computer program product runs on a processor, the processor executes the computer program or instructions to implement the steps of the echo cancellation method as described in the first aspect.
[0023] In the seventh aspect, an embodiment of the present application provides a chip, which includes a processor, a memory and a communication interface, the communication interface is coupled to the processor, the memory is used to store programs or instructions that can be run on the processor, and the processor is used to execute the program or instructions to implement the steps of the echo cancellation method described in the first aspect.
[0024] The technical solution provided by the embodiments of the present application has the following advantages over the prior art: in the embodiments of the present application, the instantaneous energy of each reference channel among multiple reference channels corresponding to the current frame reference signal is determined; based on the instantaneous energy of each reference channel, the smoothed energy of the corresponding reference channel is determined; based on the smoothed energy of each reference channel, multiple target reference channels are determined from the multiple reference channels, and the smoothed energy corresponding to each target reference channel is greater than or equal to an energy threshold; based on the multiple target reference channels, the current frame reference signal is downmixed to obtain downmixed reference signals corresponding to the multiple target reference channels; the downmixed reference signal is input into a filter module to generate an estimated echo signal corresponding to the current frame reference signal; based on the estimated echo signal, the near-end signal received by the microphone is echo cancelled to obtain the target audio signal in the near-end signal. The reference channels that play an important role in the reference signals corresponding to different sound fields and sound effects are usually different, and the smoothing energy corresponding to the reference channels that play an important role is relatively high. Therefore, the present application determines multiple target reference channels from multiple reference channels based on the relationship between the smoothing energy and the energy threshold of multiple reference channels corresponding to the current frame reference signal, and then downmixes the current frame reference signal to obtain the downmixed reference signal corresponding to the multiple target reference channels, and then performs echo cancellation of the multi-reference channel reference signal based on the downmixed reference signal. In this way, in the process of multi-channel reference echo cancellation, it is possible to flexibly match the characteristics of different sound fields and sound effects to adaptively downmix the reference signal, which is conducive to giving full play to the importance of different reference channels and can effectively reduce the overall computational complexity. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0026] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0027] Figure 1 A flowchart of an echo cancellation method provided in this application;
[0028] Figure 2 A flowchart of another echo cancellation method provided by this application;
[0029] Figure 3 A flowchart of another echo cancellation method provided by this application;
[0030] Figure 4 A schematic structural diagram of an echo cancellation device provided in this application;
[0031] Figure 5 A schematic diagram of the hardware structure of an electronic device provided in this application. DETAILED DESCRIPTION
[0032] In order to more clearly understand the above-mentioned objectives, features and advantages of the present application, the scheme of the present application will be further described below. It should be noted that, in the absence of conflict, the embodiments of the present application and the features therein can be combined with each other.
[0033] In the following description, many specific details are set forth to facilitate a full understanding of the present application, but the present application can also be implemented in other ways different from those described herein; it is obvious that the embodiments in the specification are only part of the embodiments of the present application, not all of the embodiments.
[0034] The terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of this application can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "first," "second," and the like generally distinguish objects of a class and do not limit the number of objects. For example, the first object may be one or more.
[0035] First, some nouns or terms involved in the claims and description of the present invention are explained below.
[0036] Acoustic Echo Cancellation (AEC): Based on the principle of sound wave interference, it generates an echo with equal amplitude and opposite phase to the original echo. These two sound waves interfere in space, creating a "quiet zone" to eliminate the echo. Simply put, echo cancellation technology uses an adaptive filter to simulate and remove the audio signal generated by the echo.
[0037] Adaptive filter: A filter that can automatically adjust the filtering characteristics according to the input signal.
[0038] Reference loudspeakers: Loudspeakers placed at different locations in the vehicle to generate different sound effects and sound fields. The multi-channel loudspeaker signal serves as a reference signal.
[0039] Reference signal: This is a sound signal (audio signal) played through multiple reference channels set at different positions and directions, creating a more three-dimensional and immersive sound effect. A reference channel can correspond to one speaker or multiple speakers.
[0040] Channel separation: Multi-channel speaker signals distribute different sound signals to different speakers or sound boxes, with each speaker playing a specific channel. In this way, the sound can be positioned and distributed in space, enhancing the three-dimensional and spatial perception of the audio.
[0041] Surround sound system: Common surround sound systems such as 5.1-channel, 7.1-channel, 5.1.2-channel, 7.1.4-channel, etc., create sound effects surrounding the audience by arranging multiple speakers in different positions, including front, center, rear and subwoofers.
[0042] The 5.1.2 channels of the car system: 5 represents 5 channels, including: left front door channel, right front door channel, center channel speaker, left rear door channel, right rear door channel; 1 represents the subwoofer channel; 2 represents 2 channels, including: left front 3D channel, right front 3D channel.
[0043] The 7.1.4 channels of the car system: 7 represents 7 channels, including: left front door channel, right front door channel, center channel, left rear door channel, right rear door channel, left rear surround channel, and right rear surround channel; 1 represents the subwoofer channel; 4 represents 4 channels, including: left front 3D channel, right front 3D channel, left rear 3D channel, and right rear 3D channel.
[0044] Instantaneous energy: This refers to the energy value of an audio signal at a specific moment. It is often used to describe the intensity or loudness of an audio signal at that moment. Instantaneous energy can be calculated by performing time-domain analysis on the audio signal, such as squaring or integrating the signal.
[0045] Smoothed Energy: This is the energy value obtained by smoothing the instantaneous energy. The purpose of smoothing is to reduce high-frequency variations in the instantaneous energy, resulting in a more stable and continuous energy representation. Smoothed energy can be achieved by applying a filter or averaging the instantaneous energy.
[0046] Microphone: A microphone is placed at the corresponding seat position to collect the near-end signal, which includes the far-end echo signal;
[0047] Downmixing (mix): reducing the number of reference signal channels and combining multi-channel speaker signals into one channel, thereby reducing the amount of calculation;
[0048] Least Mean Square Filter (LMS Filter): Modifies the filter coefficients by minimizing the mean square value of the error signal, adaptively simulating the desired ideal filter, where the error signal is the difference between the ideal reference signal and the actual output signal.
[0049] Recursive Least Squares (RLS) is an adaptive filtering algorithm based on the least squares method. It uses observed data to estimate the required filter coefficients. The rule is to minimize the mean square error of the prediction error and adaptively optimize the filtering effect.
[0050] The present application is applied to echo cancellation scenarios with strong reverberation, large echo, and complex echo path, especially for echo cancellation scenarios within vehicles. The electronic device in the embodiments of the present application can be a vehicle-mounted terminal, or a mobile phone, notebook, computer, etc. that communicates with the vehicle. The specific method can be determined based on actual conditions and is not limited here.
[0051] The following describes an echo cancellation scenario: When a far-end speaker plays sound, an acoustic echo signal is generated at the near-end. As a result, the near-end microphone collects a signal containing both the echo signal and the target audio signal (such as a human voice). This signal is then transmitted to the subsequent speech recognition module. This echo signal can lead to speech misrecognition and other problems. Therefore, acoustic echo cancellation (AEC) is required to remove the echo signal from the near-end signal. For ease of description, this term will be referred to as echo cancellation.
[0052] The technical solution of this application is explained in detail below through several specific embodiments.
[0053] Figure 1 A flow chart of an echo cancellation method provided by this application is shown as follows: Figure 1 As shown, the echo cancellation method may include the following steps 101 to 106.
[0054] 101. Determine the instantaneous energy of each reference channel in a plurality of reference channels corresponding to a current frame reference signal.
[0055] The current frame reference signal is a speaker signal of multiple reference channels.
[0056] 102. Based on the instantaneous energy of each reference channel, determine the smoothed energy of the corresponding reference channel.
[0057] Generally speaking, the instantaneous energy of each reference channel corresponding to the reference signals of different sound fields and sound effects is different, and the smoothing energy of each reference channel corresponding to the reference signals of different sound fields and sound effects is also different. Reference channels with different smoothing energies have different effects on the generation of sound fields and sound effects. The higher the smoothing energy of the reference channel, the more important its role in the generation of sound fields and sound effects. Therefore, when downmixing, retaining the reference channels that play an important role in the generation of sound fields and sound effects can ensure that the impact on the sound field and sound effects is minimized after downmixing, and thus, on the premise of reducing the overall amount of calculation, it is also beneficial to more accurately generate the estimated echo signal corresponding to the reference signal through the filter module, thereby better performing echo cancellation.
[0058] In some embodiments of the present application, for a reference channel, the instantaneous energy corresponding to the current frame reference signal and the average value of the instantaneous energies corresponding to multiple adjacent historical frame reference signals can be determined as the smoothed energy of the reference channel.
[0059] In some embodiments of the present application, for a reference channel, the instantaneous energy (denoted as P3) of the reference channel corresponding to the current frame reference signal can be calculated based on the instantaneous energy (denoted as P1) of the reference channel corresponding to the current frame reference signal, the smoothed energy (denoted as P2) of the reference channel corresponding to the previous frame reference signal, and the smoothing coefficient (denoted as α). For example, the smoothed energy of the reference channel corresponding to the current frame reference signal can be calculated using the formula α·P1+(1-α)·P2=P3. The smoothing coefficient can be determined based on actual conditions and is not limited here.
[0060] 103. Determine a plurality of target reference channels from the plurality of reference channels based on the smoothed energy of each reference channel, wherein the smoothed energy corresponding to each target reference channel is greater than or equal to an energy threshold.
[0061] The energy threshold can be determined according to actual conditions and is not limited here.
[0062] In some embodiments of the present disclosure, the energy threshold can be a fixed value determined based on multiple tests. In this case, the multiple target reference channels are all reference channels or any part of the reference channels among the multiple reference channels whose smoothed energy is greater than or equal to the energy threshold. The specific value can be determined based on actual conditions and is not limited here.
[0063] In some embodiments of the present disclosure, the energy threshold may be determined in real time based on actual conditions. For example, the smoothed energies corresponding to multiple reference channels may be arranged from largest to smallest, and the energy threshold may be the smoothed energy at the nth position (n is an integer greater than 2). In this case, the multiple target reference channels are the n reference channels with the largest smoothed energies among the multiple reference channels.
[0064] In some embodiments of the present disclosure, if the energy threshold is a fixed value, then in some embodiments of the present application, the above step 103 can be specifically implemented through the following steps 103a to 103c.
[0065] 103a . Determine, among the multiple reference channels, reference channels corresponding to the smoothed energy that is greater than or equal to the energy threshold as candidate reference channels, to obtain multiple candidate reference channels.
[0066] The multiple candidate reference channels are all reference channels whose corresponding smoothed energies are greater than or equal to the energy threshold among the multiple reference channels.
[0067] 103b: When the number of the multiple candidate reference channels is less than or equal to the number of target channels, determine the multiple candidate reference channels as the multiple target reference channels.
[0068] The target number of channels is the maximum number of channels after downmixing that meets the overall computational complexity requirement, which can be determined based on actual conditions and is not limited here.
[0069] 103c. When the number of the candidate reference channels is greater than the number of the target channels, determine the target reference channels from the candidate reference channels based on a priori rules, wherein the number of the target reference channels is equal to the number of the target channels.
[0070] In the disclosed embodiment, based on the relationship between the number of channels of multiple candidate reference channels and the number of target channels, as well as the set prior rules, multiple target reference channels are flexibly determined from multiple reference channels. In this way, the characteristics of different sound fields and sound effects, and the overall computational requirements can be combined to better determine the downmixing method for the current frame reference signal.
[0071] In some embodiments of the present application, the a priori rule includes at least one of the following: a candidate reference channel with a high priority is designated as the target reference channel; a candidate reference channel in a reference channel combination is designated as the target reference channel, and different reference channel combinations include different reference channels from the plurality of reference channels. The a priori rule may also include other content, which may be determined based on actual circumstances and is not limited herein.
[0072] It can be understood that when the prior rule includes a candidate reference channel with a high priority as the target reference channel, it means that the candidate reference channel with a higher priority has a more important effect on the production of the sound field and sound effects. When downmixing, the candidate reference channel with a higher priority is used as the target reference channel, which has less impact on the sound field and sound effects, that is, the downmixing (merging) effect is better. Therefore, the target channel with the highest priority among multiple candidate reference channels is used as multiple target reference channels for downmixing, and the downmixing effect is better.
[0073] It can be understood that when the prior rule includes a candidate reference channel in a reference channel merging combination as the target reference channel, it means that the reference channels belonging to a reference channel merging group have less impact on the sound field and sound effects when merging during downmixing, that is, the downmixing (merging) effect is better. Therefore, at least two candidate reference channels belonging to a reference channel merging group among multiple candidate reference channels can be merged during downmixing, that is, a candidate reference channel belonging to a reference channel merging group (which can be any candidate reference channel or a candidate reference channel with the highest priority) is used as the target reference channel (recorded as target reference channel 1), and the other candidate reference channels belonging to the reference channel merging group are used as reference channels to be merged with the target reference channel 1 during downmixing.
[0074] Example 1: Assume that the current frame reference signal is a microphone signal of 8 reference channels (reference channel 1 to reference channel 8). There are 5 candidate reference channels among the 8 reference channels (reference channel 1, reference channel 2, reference channel 4, reference channel 6, and reference channel 7), and the number of target channels is 3. Reference channel 1 and reference channel 2 belong to a reference channel merging group, reference channel 3 and reference channel 4 belong to a reference channel merging group, and reference channel 6, reference channel 7, and reference channel 8 belong to a reference channel merging group. Therefore, any one of reference channel 1 and reference channel 2 can be used as the target reference channel (for example, reference channel 1), reference channel 4 can be used as the target reference channel, and any one of reference channel 6 and reference channel 7 can be used as the target reference channel (for example, reference channel 7). Then, 3 target reference channels are obtained, and downmixing is performed based on the 3 target reference channels. The downmixing method includes: combining the signals corresponding to reference channel 1 and reference channel 2 and outputting them through reference channel 1, combining the signals corresponding to reference channel 3 and reference channel 4 and outputting them through reference channel 4, and combining the signals corresponding to reference channel 6, reference channel 7, and reference channel 8 and outputting them through reference channel 7.
[0075] It can be understood that when the prior rule includes a candidate reference channel with a high priority as the target reference channel, and a candidate reference channel in a reference channel combination as the target reference channel, the candidate reference channel with the highest priority in the reference channel combination is downmixed as the target reference channel.
[0076] In the embodiments of the present disclosure, a variety of a priori rules are provided, and the target reference channel can be better determined according to the a priori rules, thereby determining a more accurate downmixing method.
[0077] In some embodiments of the present disclosure, after the multiple candidate reference channels are determined in the above step 103a, the multiple candidate reference channels corresponding to the current frame reference signal can be compared to see whether they are the same as the multiple candidate reference channels corresponding to the previous frame reference signal. If they are the same, the multiple target reference channels corresponding to the previous frame reference signal can be directly determined as the multiple target reference channels corresponding to the current frame reference signal (so, the current frame reference signal can adopt the same downmixing method as the previous frame reference signal); if they are not the same, in one case, the multiple target reference channels corresponding to the current frame reference signal can be re-determined according to the above steps 103b and 103c (so, the current frame reference signal adopts a downmixing method different from the previous frame reference signal), and in another case, it can be determined whether the multiple candidate reference channels corresponding to the current frame reference signal include the previous frame reference signal. All reference channels among the multiple target reference channels corresponding to a frame reference signal; if the multiple candidate reference channels corresponding to the current frame reference signal include all reference channels among the multiple target reference channels corresponding to the previous frame reference signal, the multiple target reference channels corresponding to the previous frame reference signal may also be determined as the multiple target reference channels corresponding to the current frame reference signal (in this way, the current frame reference signal may adopt the same downmixing method as the previous frame reference signal); if the multiple candidate reference channels corresponding to the current frame reference signal only include some of the multiple target reference channels corresponding to the previous frame reference signal, the multiple target reference channels corresponding to the current frame reference signal may be re-determined according to the above steps 103b and 103c (in this way, the current frame reference signal adopts a different downmixing method from the previous frame reference signal).
[0078] It can be understood that, according to whether the reference channel that plays an important role in the current frame reference signal has changed relative to the reference channel that plays an important role in the previous frame reference signal, if there is no change, the previous downmixing method is directly used for downmixing; if there is a change, the downmixing method can be directly re-determined through the above steps 103b and 103c (i.e., determining multiple target reference channels), or it can be further determined according to whether the reference channel that plays an important role in the current frame reference signal includes all the reference channels in the multiple target reference channels corresponding to the previous frame reference signal. If included, the previous downmixing method is directly used for downmixing; if not included, the downmixing method is directly re-determined through the above steps 103b and 103c (i.e., determining multiple target reference channels). In this way, it is possible to avoid the downmixing method from changing too frequently, thereby ensuring that the filter module can converge quickly.
[0079] In the embodiment of the present disclosure, the above step 103 may also determine multiple target reference channels by other methods, which are not limited here.
[0080] 104. Downmix the current frame reference signal based on the multiple target reference channels to obtain downmixed reference signals corresponding to the multiple target reference channels.
[0081] After determining multiple target reference channels, a downmixing method is determined. The downmixing method can be determined based on historical experience or randomly. During the downmixing process, the weights of the multiple original reference channels to be merged in the target reference channel are calculated based on the downmixing method, and then the signals corresponding to the multiple original reference channels are weighted and output through the target reference channel. The downmixed reference signals corresponding to the multiple target reference channels are then obtained as system input.
[0082] 105. Input the downmixed reference signal into a filter module to generate an estimated echo signal corresponding to the current frame reference signal.
[0083] In the embodiment of the present application, the specific structure of the filter module is not limited, and can be determined by referring to the filter structure for echo estimation in related technologies and actual usage requirements.
[0084] For example, the filter module includes an adaptive filter. The filter module may include an adaptive FIR filter, an adaptive MDF filter, or a dual filter structure. The specific structure may be determined according to actual conditions and is not limited here.
[0085] The dual filter structure includes a foreground filter and an adaptive background filter.
[0086] 106. Perform echo cancellation on a near-end signal received by a microphone based on the estimated echo signal to obtain a target audio signal in the near-end signal.
[0087] It is understood that the near-end signal includes the target audio signal and the echo signal of the current frame reference signal, and may also include other noise signals, which is not limited here. In the embodiment of the present application, it is assumed that other noise signals in the near-end signal can be removed by other means.
[0088] In the embodiment of the present application, echo cancellation is performed by offsetting the estimated echo signal with the echo signal of the current frame reference signal in the near-end signal, and then obtaining the target audio signal.
[0089] In an embodiment of the present application, the instantaneous energy of each reference channel in a plurality of reference channels corresponding to a current frame reference signal is determined; based on the instantaneous energy of each reference channel, the smoothed energy of the corresponding reference channel is determined; based on the smoothed energy of each reference channel, a plurality of target reference channels are determined from the plurality of reference channels, and the smoothed energy corresponding to each target reference channel is greater than or equal to an energy threshold; based on the plurality of target reference channels, the current frame reference signal is downmixed to obtain downmixed reference signals corresponding to the plurality of target reference channels; the downmixed reference signal is input into a filter module to generate an estimated echo signal corresponding to the current frame reference signal; based on the estimated echo signal, the near-end signal received by the microphone is echo cancelled to obtain a target audio signal in the near-end signal. The reference channels that play an important role in the reference signals corresponding to different sound fields and sound effects are usually different, and the smoothing energy corresponding to the reference channels that play an important role is relatively high. Therefore, the present application determines multiple target reference channels from multiple reference channels based on the relationship between the smoothing energy and the energy threshold of multiple reference channels corresponding to the current frame reference signal, and then downmixes the current frame reference signal to obtain the downmixed reference signal corresponding to the multiple target reference channels, and then performs echo cancellation of the multi-reference channel reference signal based on the downmixed reference signal. In this way, in the process of multi-channel reference echo cancellation, it is possible to flexibly match the characteristics of different sound fields and sound effects to adaptively downmix the reference signal, which is conducive to giving full play to the importance of different reference channels and can effectively reduce the overall computational complexity.
[0090] In some embodiments of the present application, the filter module includes a foreground filter (Foreground Filter) and a background filter (Background Filter), and the background filter is an adaptive filter; Figure 1 ,like Figure 2 As shown, the above step 105 can be specifically implemented through the following steps 105a and 105b, and the above step 106 can be specifically implemented through the following steps 106a to 106c.
[0091] 105a: Input the downmixed reference signal into the foreground filter to obtain a first estimated echo signal.
[0092] 105b. Input the downmixed reference signal into the background filter to obtain a second estimated echo signal.
[0093] Among them, the background filter adopts an adaptive filtering algorithm, such as LMS, RLS, etc., which automatically adjusts the filter coefficient and outputs the echo signal estimated by the background filter.
[0094] 106a. Perform echo cancellation on the near-end signal based on the first estimated echo signal to obtain a first error signal.
[0095] 106b. Perform echo cancellation on the near-end signal based on the second estimated echo signal to obtain a second error signal.
[0096] 106c. Determine a smaller value between the first error signal and the second error signal as the target audio signal.
[0097] It can be understood that if the first error signal is smaller than the second error signal, the foreground filter's echo cancellation performance is superior; if the first error signal is greater than the second error signal, the background filter's echo cancellation performance is superior. This dual-filter structure comprises an iteratively updated adaptive background filter and a non-adaptive foreground filter. If the adaptive filter's performance deteriorates or even diverges, the AEC uses the foreground filter's echo cancellation results as the target audio signal. When the background filter's performance improves, the AEC uses the background filter's echo cancellation results as the target audio signal. This allows for better echo cancellation and a purer target audio signal.
[0098] For example, Figure 3As shown, a possible schematic diagram of an echo cancellation system including a dual filter structure is shown, wherein the energy calculation module 31 generates smoothed energies corresponding to multiple reference channels of the current frame reference signal, the downmixing mode determination module 32 determines multiple target reference channels based on the smoothed energies corresponding to the multiple reference channels, and then determines the downmixing mode. The reference channel compression module 33 downmixes the current frame reference signal based on the downmixing mode to obtain downmixed reference signals corresponding to the multiple target reference channels. The downmixed reference signal is input as a far-end input signal, and the downmixed reference signal is output through the near-end speaker module 34. The downmixed reference signal output by the speaker module 34 passes through the room impulse response module 35 (room impulse response module The near-end microphone module 36 collects the echo signal and the target audio signal (near-end audio signal) as the near-end signal. The downmixed reference signal passes through the foreground filter 37 and the background filter 38 respectively to generate a first estimated echo signal (yf in the figure) and a second estimated echo signal (yb in the figure). The first estimated echo signal and the echo signal in the near-end signal collected by the microphone module 36 are added and offset to obtain a first error signal (ef in the figure). The second estimated echo signal and the echo signal in the near-end signal collected by the microphone module 36 are added and offset to obtain a second error signal (eb in the figure). The first error signal and the second error signal are then compared, and the smaller value of the first error signal and the second error signal is output at the far end as the target audio signal. The coefficient updating module 39 in the figure is used to update the coefficients of the foreground filter 37 and the background filter 38 according to the magnitude of the first error signal and the second error signal.
[0099] In some embodiments of the present application, after the above step 106b, the echo cancellation method provided in the embodiment of the present application may further include the following steps 107 and 108.
[0100] 107. When the first error signal is smaller than the second error signal, update the filter coefficients of the background filter to the filter coefficients of the foreground filter.
[0101] 108. When the first error signal is greater than the second error signal, update the filter coefficients of the foreground filter to the filter coefficients of the background filter.
[0102] It can be understood that if the first error signal is smaller than the second error signal, it indicates that the performance of the adaptive background filter has deteriorated or even diverged. Resetting the background filter using the coefficients of the foreground filter can improve the performance of the background filter. If the first error signal is larger than the second error signal, it indicates that the performance of the background filter has improved. Resetting the foreground filter using the coefficients of the background filter can improve the performance of the foreground filter. This can result in a better estimated echo signal, better echo cancellation, and a purer target audio signal.
[0103] In some embodiments of the present application, after downmixing the current frame reference signal based on the multiple target reference channels to obtain the downmixed reference signals corresponding to the multiple target reference channels, and before inputting the downmixed reference signal into the background filter to obtain the second estimated echo signal, after the above step 104 and before the above step 105b, the echo cancellation method provided in the embodiment of the present application may further include the following steps 109 and 110.
[0104] 109. Record the downmixing mode label of the current frame reference signal.
[0105] The downmixing mode tag is used to indicate the downmixing mode of downmixing the multiple reference channels into the multiple target reference channels. The downmixing mode tag indicates which reference channels are downmixed into which reference channels.
[0106] Example 2 follows Example 1. The downmixing method labels are: downmix reference channel 1 and reference channel 2 into reference channel 1, downmix reference channel 3 and reference channel 4 into reference channel 4, and downmix reference channel 6, reference channel 7, and reference channel 8 into reference channel 7.
[0107] 110 . When the downmix mode label of the current frame reference signal is different from the downmix mode label of the previous frame reference signal, adaptively update the filter coefficients of the background filter based on the downmix mode label of the current frame reference signal.
[0108] In the embodiment of the present application, when the downmix mode flag of the current frame reference signal differs from the downmix mode flag of the previous frame reference signal, the filter coefficients of the background filter are adaptively updated based on the downmix mode flag of the current frame reference signal. This allows the coefficients of the background filter to be adaptively adjusted as the downmix mode flag changes, thereby improving the convergence speed of the background filter.
[0109] In some embodiments of the present application, when the downmixing mode label of the current frame reference signal is the same as the downmixing mode label of the previous frame reference signal, the downmixed reference signal is directly input into the background filter to obtain the second estimated echo signal.
[0110] In some embodiments of the present application, the above step 105b can be specifically implemented through the following step 105b1.
[0111] 105b1. When the downmix mode label of the current frame reference signal is different from the downmix mode label of the previous frame reference signal, in the process of adaptively filtering the downmixed reference signal through the background filter, adaptively adjust the iterative step size of the background filter based on the downmix mode label of the current frame reference signal.
[0112] In an embodiment of the present application, when the downmixing mode label of the current frame reference signal is different from the downmixing mode label of the previous frame reference signal, during the process of adaptively filtering the downmixed reference signal through the background filter, the iteration step size of the background filter is adaptively adjusted based on the downmixing mode label of the current frame reference signal (the iteration step size can be first increased and then adjusted back to the iteration step size before the increase). In this way, the convergence speed of the background filter can be improved and the convergence can be smoothed.
[0113] In some embodiments of the present application, when the downmixing mode label of the current frame reference signal is the same as the downmixing mode label of the previous frame reference signal, in the process of adaptively filtering the downmixed reference signal through the background filter, the iterative step size of the background filter may not be adaptively adjusted.
[0114] This application also provides an echo cancellation device, Figure 4 This is a schematic diagram of the structure of an echo cancellation device provided by this application, such as Figure 4 As shown, the echo cancellation device includes: a determination module 401, which is used to determine the instantaneous energy of each reference channel in a plurality of reference channels corresponding to the current frame reference signal; based on the instantaneous energy of each reference channel, respectively determining the smoothed energy of the corresponding reference channel; based on the smoothed energy of each reference channel, determining a plurality of target reference channels from the plurality of reference channels, wherein the smoothed energy corresponding to each target reference channel is greater than or equal to an energy threshold; a downmixing module 402, which is used to downmix the current frame reference signal based on the plurality of target reference channels to obtain downmixed reference signals corresponding to the plurality of target reference channels; a generation module 403, which is used to input the downmixed reference signal into a filter module to generate an estimated echo signal corresponding to the current frame reference signal; and an echo cancellation module 404, which is used to perform echo cancellation on a near-end signal received by a microphone based on the estimated echo signal to obtain a target audio signal in the near-end signal.
[0115] In some embodiments of the present application, the determination module 401 is specifically used to determine, among the multiple reference channels, the reference channels whose corresponding smoothed energy is greater than or equal to the energy threshold as candidate reference channels, to obtain multiple candidate reference channels; when the number of channels of the multiple candidate reference channels is less than or equal to the number of target channels, determine the multiple candidate reference channels as the multiple target reference channels; when the number of channels of the multiple candidate reference channels is greater than the number of target channels, determine the multiple target reference channels from the multiple candidate reference channels based on a priori rules, and the number of channels of the multiple target reference channels is equal to the number of target channels.
[0116] In some embodiments of the present application, the prior rule includes at least one of the following: a candidate reference channel with a high priority is the target reference channel; a candidate reference channel in a reference channel merging combination is the target reference channel, and different reference channel merging combinations include different reference channels among multiple reference channels.
[0117] In some embodiments of the present application, the filter module includes a foreground filter and a background filter, and the background filter is an adaptive filter; the generation module 403 is specifically configured to input the downmixed reference signal into the foreground filter to obtain a first estimated echo signal; and input the downmixed reference signal into the background filter to obtain a second estimated echo signal;
[0118] The echo cancellation module 404 is specifically configured to perform echo cancellation on the near-end signal based on the first estimated echo signal to obtain a first error signal; perform echo cancellation on the near-end signal based on the second estimated echo signal to obtain a second error signal; and determine the smaller value of the first error signal and the second error signal as the target audio signal.
[0119] In some embodiments of the present application, the device also includes: an updating module, which is used to update the filter coefficients of the background filter to the filter coefficients of the foreground filter when the first error signal is less than the second error signal after performing echo cancellation on the near-end signal based on the second estimated echo signal to obtain a second error signal; and when the first error signal is greater than the second error signal, update the filter coefficients of the foreground filter to the filter coefficients of the background filter.
[0120] In some embodiments of the present application, the device also includes: a recording module for recording a downmixing mode label of the current frame reference signal after downmixing the current frame reference signal based on the multiple target reference channels to obtain downmixed reference signals corresponding to the multiple target reference channels and before inputting the downmixed reference signal into the background filter to obtain a second estimated echo signal, wherein the downmixing mode label is used to indicate the downmixing mode of the multiple reference channels being downmixed into the multiple target reference channels; and an updating module for adaptively updating the filter coefficients of the background filter based on the downmixing mode label of the current frame reference signal when the downmixing mode label of the current frame reference signal is different from the downmixing mode label of the previous frame reference signal.
[0121] In some embodiments of the present application, the echo cancellation module 404 is specifically configured to adaptively adjust the iterative step size of the background filter based on the downmixing mode label of the current frame reference signal when the downmixing mode label of the current frame reference signal is different from the downmixing mode label of the previous frame reference signal during the process of adaptively filtering the downmixed reference signal through the background filter.
[0122] It should be noted that the above-mentioned echo cancellation device can be the electronic device in the above-mentioned method embodiment of this application, or it can be a functional module and / or functional entity in the electronic device that can realize the functions of the device embodiment, and the embodiment of this application does not limit it.
[0123] In the embodiment of the present application, each module can implement the echo cancellation method provided by the above method embodiment and can achieve the same technical effect. To avoid repetition, it will not be described here.
[0124] Figure 5 The structural diagram of an electronic device provided in the embodiment of the present application is used to exemplarily illustrate an electronic device that implements any echo cancellation method in the embodiment of the present application, and should not be understood as a specific limitation on the embodiment of the present application.
[0125] like Figure 5 As shown, the electronic device 500 may include a processor (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. Various programs and data required for the operation of the electronic device 500 are also stored in the RAM 503. The processor 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0126] Typically, the following devices may be connected to the I / O interface 505: an input device 506 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 508 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 509. The communication device 509 may allow the electronic device 500 to communicate with other devices wirelessly or by wire to exchange data. Although the electronic device 500 is shown as having various devices, it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have alternatively.
[0127] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication device 509, or installed from the storage device 508, or installed from the ROM 502. When the computer program is executed by the processor 501, the functions defined in any echo cancellation method provided in the embodiment of the present application can be performed.
[0128] It should be noted that the computer-readable medium mentioned above in this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. Computer-readable storage media can be, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or devices, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. This propagated data signal can take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0129] In some embodiments, the client and server can communicate using any currently known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.
[0130] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0131] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device: determines the instantaneous energy of each reference channel in the multiple reference channels corresponding to the current frame reference signal; based on the instantaneous energy of each reference channel, determines the smoothed energy of the corresponding reference channel; based on the smoothed energy of each reference channel, determines multiple target reference channels from the multiple reference channels, and the smoothed energy corresponding to each target reference channel is greater than or equal to an energy threshold; downmixes the current frame reference signal based on the multiple target reference channels to obtain downmixed reference signals corresponding to the multiple target reference channels; inputs the downmixed reference signal into a filter module to generate an estimated echo signal corresponding to the current frame reference signal; based on the estimated echo signal, performs echo cancellation on the near-end signal received by the microphone to obtain the target audio signal in the near-end signal.
[0132] In an embodiment of the present application, a computer program code for performing the operations of the present application can be written in one or more programming languages or a combination thereof, and the above-mentioned programming languages include but are not limited to object-oriented programming languages, such as Java, Smalltalk, C++, and also include conventional procedural programming languages, such as "C" language or similar programming languages. The program code can be executed entirely on the computer, partially on the computer, as an independent software package, partially on the computer and partially on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer can be connected to the computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, using an Internet service provider to connect through the Internet).
[0133] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.
[0134] The units involved in the embodiments described in this application may be implemented by software or hardware, wherein the name of a unit does not, in some cases, constitute a limitation on the unit itself.
[0135] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0136] In the context of the present application, computer-readable medium can be a tangible medium that can contain or store a program for use by an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. Computer-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. More specific examples of computer-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0137] An embodiment of the present application further provides a vehicle, comprising: the above-mentioned echo cancellation device, or the above-mentioned electronic device, or the above-mentioned computer-readable storage medium.
[0138] The above description is merely a preferred embodiment of the present application and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure in this application is not limited to the technical solutions formed by a specific combination of the above-mentioned technical features, but also encompasses other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this application.
[0139] In addition, although adopting specific order to describe each operation, this should not be interpreted as requiring these operations to be executed in the specific order shown or in sequential order.Under certain environment, multitasking and parallel processing may be advantageous.Similarly, although comprising some specific implementation details in the above discussion, these should not be interpreted as limiting the scope of the application.Some features described in the context of separate embodiment can also be implemented in a single embodiment in combination.On the contrary, the various features described in the context of a single embodiment also can be implemented in multiple embodiments individually or in the mode of any suitable subcombination.
[0140] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. An echo cancellation method, characterized in that: The method comprises: Determining the instantaneous energy of each reference channel in a plurality of reference channels respectively corresponding to the current frame reference signal; Based on the instantaneous energy of each reference channel, respectively determine the smoothed energy of the corresponding reference channel; Determining a plurality of target reference channels from the plurality of reference channels based on the smoothed energy of each reference channel, wherein the smoothed energy corresponding to each target reference channel is greater than or equal to an energy threshold; Downmixing the current frame reference signal based on the multiple target reference channels to obtain downmixed reference signals corresponding to the multiple target reference channels; Inputting the downmixed reference signal into a filter module to generate an estimated echo signal corresponding to the current frame reference signal; Based on the estimated echo signal, echo cancellation is performed on the near-end signal received by the microphone to obtain a target audio signal in the near-end signal.
2. The method according to claim 1, characterized in that The determining a plurality of target reference channels from the plurality of reference channels based on the smoothed energy of each reference channel comprises: Determine, among the multiple reference channels, reference channels corresponding to the smoothed energy greater than or equal to the energy threshold as candidate reference channels, to obtain multiple candidate reference channels; In a case where the number of channels of the multiple candidate reference channels is less than or equal to the number of target channels, determining the multiple candidate reference channels as the multiple target reference channels; In a case where the number of channels of the multiple candidate reference channels is greater than the number of target channels, the multiple target reference channels are determined from the multiple candidate reference channels based on a priori rules, and the number of channels of the multiple target reference channels is equal to the number of target channels.
3. The method according to claim 2, characterized in that The prior rules include at least one of the following: The candidate reference channel with the highest priority is the target reference channel; A candidate reference channel in a reference channel merging combination is a target reference channel, and different reference channel merging combinations include different reference channels among the multiple reference channels.
4. The method according to claim 1, wherein The filter module includes a foreground filter and a background filter, and the background filter is an adaptive filter; The step of inputting the downmixed reference signal into a filter module to generate an estimated echo signal corresponding to the current frame reference signal includes: inputting the downmixed reference signal into the foreground filter to obtain a first estimated echo signal; Inputting the downmixed reference signal into the background filter to obtain a second estimated echo signal; The step of performing echo cancellation on a near-end signal received by a microphone based on the estimated echo signal to obtain a target audio signal in the near-end signal includes: performing echo cancellation on the near-end signal based on the first estimated echo signal to obtain a first error signal; performing echo cancellation on the near-end signal based on the second estimated echo signal to obtain a second error signal; A smaller value between the first error signal and the second error signal is determined as the target audio signal.
5. The method according to claim 4, characterized in that After performing echo cancellation on the near-end signal based on the second estimated echo signal to obtain a second error signal, the method further includes: When the first error signal is less than the second error signal, updating the filter coefficients of the background filter to the filter coefficients of the foreground filter; When the first error signal is greater than the second error signal, the filter coefficients of the foreground filter are updated to the filter coefficients of the background filter.
6. The method according to claim 4, characterized in that After downmixing the current frame reference signal based on the multiple target reference channels to obtain downmixed reference signals corresponding to the multiple target reference channels, and before inputting the downmixed reference signal into the background filter to obtain the second estimated echo signal, the method further includes: Recording a downmixing mode tag of the current frame reference signal, where the downmixing mode tag is used to indicate a downmixing mode in which the multiple reference channels are downmixed into the multiple target reference channels; When the downmix mode label of the current frame reference signal is different from the downmix mode label of the previous frame reference signal, the filter coefficients of the background filter are adaptively updated based on the downmix mode label of the current frame reference signal.
7. The method according to claim 6, characterized in that Inputting the downmixed reference signal into the background filter to obtain a second estimated echo signal includes: When the downmix mode label of the current frame reference signal is different from the downmix mode label of the previous frame reference signal, in a process of adaptively filtering the downmixed reference signal through the background filter, an iterative step size of the background filter is adaptively adjusted based on the downmix mode label of the current frame reference signal.
8. An echo cancellation device, characterized in that: include: A determination module, configured to determine the instantaneous energy of each reference channel in a plurality of reference channels corresponding to the current frame reference signal; Based on the instantaneous energy of each reference channel, respectively determine the smoothed energy of the corresponding reference channel; Determining a plurality of target reference channels from the plurality of reference channels based on the smoothed energy of each reference channel, wherein the smoothed energy corresponding to each target reference channel is greater than or equal to an energy threshold; a downmixing module, configured to downmix the current frame reference signal based on the multiple target reference channels to obtain downmixed reference signals corresponding to the multiple target reference channels; a generating module, configured to input the downmixed reference signal into a filter module to generate an estimated echo signal corresponding to the current frame reference signal; The echo cancellation module is configured to perform echo cancellation on a near-end signal received by a microphone based on the estimated echo signal to obtain a target audio signal in the near-end signal.
9. An electronic device, characterized in that: include: A processor, wherein the processor is configured to execute a computer program stored in a memory, wherein the computer program, when executed by the processor, implements the steps of the echo cancellation method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the echo cancellation method according to any one of claims 1 to 7 are implemented.
11. A vehicle, characterized in that: include: The echo cancellation device according to claim 8, or the electronic device according to claim 9, or comprising the computer-readable storage medium according to claim 10.