Method and electronic device for echo cancellation

By employing a dual-filter technique, the first filter eliminates echoes in a steady state, while the second filter rapidly tracks path changes and adjusts the aggressiveness of echo suppression and the convergence speed. This solves the echo leakage problem when the echo path changes abruptly, thus improving the stability and effectiveness of echo cancellation.

CN116013345BActive Publication Date: 2026-02-27ZHEJIANG DAHUA TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211672640.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-26
Publication Date
2026-02-27
Estimated Expiration
2042-12-26

AI Technical Summary

Technical Problem

Existing echo cancellation algorithms deteriorate drastically when the echo path changes abruptly, leading to widespread echo leakage and affecting call quality and user experience.

Method used

The dual-filter technology is adopted. The first filter efficiently eliminates echoes in steady state, while the second filter quickly tracks changes in the echo path. By adjusting the echo suppression aggressiveness and convergence speed, the echoes are quickly suppressed, avoiding large-area echo leakage.

Benefits of technology

It effectively suppresses echoes when the echo path changes abruptly, avoiding echo leakage and improving call quality and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116013345B_ABST
    Figure CN116013345B_ABST
Patent Text Reader

Abstract

The application discloses an echo cancellation method and electronic equipment, which is used for better eliminating echo in a steady state and avoiding large-area echo leakage when a path mutates. The method comprises the following steps: acquiring an audio signal, wherein the audio signal comprises a near-end signal and a far-end signal; filtering and processing the far-end signal by using a first filter and a second filter respectively to obtain respective corresponding echo information, wherein the first filter has higher echo elimination capability than the second filter when an echo path is stable, and the second filter has a convergence speed greater than the first filter when the echo path changes; adjusting echo suppression aggressiveness according to the respective corresponding echo information, and determining a first residual signal by using the echo information corresponding to the first filter and the near-end signal; and suppressing the first residual signal by using the adjusted echo suppression aggressiveness to obtain an audio signal after echo cancellation.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of audio signal processing, in particular to an echo cancellation method and electronic equipment. BACKGROUND

[0002] In instant messaging applications, real-time voice communication between two or more parties is required. In high-demand situations, external speakers are usually used for playback, which will inevitably produce echo, i.e., after one party speaks, the sound is played back through the other party's speaker, and then picked up by the other party's microphone and transmitted back to the speaker. If the echo is not processed, it will affect the call quality and user experience, and even worse, it will cause oscillation and howling.

[0003] Acoustic echo cancellation is a processing method that prevents the return of the far-end sound by canceling or removing the far-end audio signal picked up by the local microphone. After the microphone picks up the sound, the sound played by the local speaker is removed from the microphone's sound data, so that the microphone records only the local user's voice.

[0004] However, the echo cancellation algorithm currently used will deteriorate sharply in the case of echo path mutation, and the echo cancellation effect will also deteriorate sharply compared to the steady-state echo path, and large-area echo leakage will occur. SUMMARY

[0005] The present application provides an echo cancellation method and electronic equipment for utilizing the different characteristics of double filters in the steady state and convergence process to achieve better echo cancellation in the steady state and quickly suppress echo to avoid large-area echo leakage in the case of path mutation.

[0006] In a first aspect, the present application provides an echo cancellation method, which comprises:

[0007] Obtaining an audio signal, the audio signal comprising a near-end signal and a far-end signal, the near-end signal and the far-end signal being distinguished based on the propagation mode of the audio signal;

[0008] Filtering the far-end signal using a first filter and a second filter respectively to obtain respective corresponding echo information, wherein the first filter has higher echo cancellation capability than the second filter when the echo path is stable, and the second filter has greater convergence speed than the first filter when the echo path changes;

[0009] Adjusting the echo suppression aggressiveness according to the respective corresponding echo information, and determining a first residual signal using the first filter corresponding echo information and the near-end signal;

[0010] The first residual signal is suppressed by using the adjusted echo suppression aggressiveness, to obtain an echo-canceled audio signal.

[0011] The embodiment uses a second filter to track the echo path, and adjusts the echo suppression aggressiveness when detecting a change in the echo path, adopts aggressive echo suppression measures, and reduces the situation of large-area echo leakage.

[0012] In a second aspect, an electronic device is provided, including a processor and a memory, the memory is used to store programs executable by the processor, and the processor is used to read the programs in the memory and perform the following steps:

[0013] An audio signal is obtained, the audio signal includes a near-end signal and a far-end signal, the near-end signal and the far-end signal are distinguished based on a propagation mode of the audio signal;

[0014] The far-end signal is filtered by using a first filter and a second filter respectively, to obtain respective corresponding echo information, wherein the first filter has higher echo cancellation capability than the second filter when the echo path is stable, and the second filter has a greater convergence speed than the first filter when the echo path changes;

[0015] The echo suppression aggressiveness is adjusted according to the respective corresponding echo information, and a first residual signal is determined by using the echo information corresponding to the first filter and the near-end signal;

[0016] The first residual signal is suppressed by using the adjusted echo suppression aggressiveness, to obtain an echo-canceled audio signal.

[0017] As an optional implementation, when the echo information includes an echo path, the processor is specifically configured to perform:

[0018] The similarity between the respective corresponding echo paths is determined;

[0019] The echo suppression aggressiveness is adjusted according to the similarity, wherein the echo suppression aggressiveness increases with the increase of the similarity.

[0020] As an optional implementation, when the echo information includes an echo signal, the processor is specifically configured to perform:

[0021] A second residual signal is determined according to the echo information corresponding to the second filter and the near-end signal;

[0022] An energy ratio is determined according to the power of the first residual signal and the power of the second residual signal;

[0023] adjusting an echo suppression aggressiveness according to the energy ratio, wherein the echo suppression aggressiveness increases as the energy ratio increases.

[0024] As an optional implementation, the processor is specifically configured to perform:

[0025] determining a state of a first filter according to the respective corresponding echo information, and adjusting an echo suppression aggressiveness according to the state of the first filter.

[0026] As an optional implementation, the processor is specifically further configured to perform:

[0027] adjusting a convergence speed of the first filter according to the respective corresponding echo information;

[0028] filtering the far-end signal by using the adjusted first filter to obtain corresponding new echo information, and continuing to adjust the echo suppression aggressiveness according to the new echo information and the corresponding echo information of the second filter.

[0029] As an optional implementation, when the echo information comprises echo paths, the processor is specifically configured to perform:

[0030] determining a similarity between the respective corresponding echo paths;

[0031] adjusting a convergence speed of the first filter according to the similarity, wherein the convergence speed decreases as the similarity increases.

[0032] As an optional implementation, when the echo information comprises echo signals, the processor is specifically configured to perform:

[0033] determining a second residual signal according to the corresponding echo information of the second filter and the near-end signal;

[0034] determining an energy ratio according to a power of the first residual signal and a power of the second residual signal;

[0035] adjusting a convergence speed of the first filter according to the energy ratio, wherein the convergence speed increases as the energy ratio increases.

[0036] As an optional implementation, the processor is specifically configured to perform:

[0037] adjusting a noise covariance matrix of the first filter according to the respective corresponding echo information, wherein the noise covariance matrix is used to represent a degree of deviation of echo path estimation by the first filter;

[0038] According to the noise covariance matrix, the convergence speed of the first filter is adjusted.

[0039] As an optional implementation, the processor is specifically configured to perform:

[0040] According to the respective corresponding echo information, the convergence state of the first filter is determined, and the convergence speed of the first filter is adjusted according to the convergence state of the first filter.

[0041] As an optional implementation, the processor is specifically configured to perform:

[0042] The steady-state error of the first filter is smaller than the steady-state error of the second filter, and / or the convergence speed of the second filter is greater than the convergence speed of the first filter.

[0043] In a third aspect, the embodiments of the present application also provide an echo cancellation device, which comprises:

[0044] An audio acquisition module is configured to acquire an audio signal, wherein the audio signal comprises a near-end signal and a far-end signal, and the near-end signal and the far-end signal are distinguished based on a propagation mode of the audio signal.

[0045] A double-filtering module is configured to filter the far-end signal by using a first filter and a second filter respectively to obtain respective corresponding echo information, wherein the first filter has a higher ability to eliminate echo than the second filter when the echo path is stable, and the second filter has a greater convergence speed than the first filter when the echo path changes.

[0046] A residual calculation module is configured to adjust an echo suppression aggressiveness according to the respective corresponding echo information, and determine a first residual signal by using the corresponding echo information of the first filter and the near-end signal.

[0047] An echo suppression module is configured to suppress the first residual signal by using the adjusted echo suppression aggressiveness to obtain an echo-canceled audio signal.

[0048] As an optional implementation, when the echo information comprises an echo path, the residual calculation module is specifically configured to:

[0049] Determine a similarity between the respective corresponding echo paths.

[0050] Adjust the echo suppression aggressiveness according to the similarity, wherein the echo suppression aggressiveness increases with an increase of the similarity.

[0051] As an optional implementation, when the echo information comprises an echo signal, the residual calculation module is specifically configured to:

[0052] determining a second residual signal according to the echo information corresponding to the second filter and the near-end signal;

[0053] determining an energy ratio according to the power of the first residual signal and the power of the second residual signal;

[0054] adjusting an echo suppression aggressiveness according to the energy ratio, wherein the echo suppression aggressiveness increases as the energy ratio increases.

[0055] As an optional implementation, the calculating residual module is specifically configured to:

[0056] determining a state of the first filter according to the respective corresponding echo information, and adjusting an echo suppression aggressiveness according to the state of the first filter.

[0057] As an optional implementation, the adjusting convergence speed module is specifically configured to:

[0058] adjusting a convergence speed of the first filter according to the respective corresponding echo information;

[0059] filtering the far-end signal by using the adjusted first filter to obtain corresponding new echo information, and continuing to adjust the echo suppression aggressiveness according to the new echo information and the echo information corresponding to the second filter.

[0060] As an optional implementation, when the echo information includes an echo path, the adjusting convergence speed module is specifically configured to:

[0061] determining a similarity between the respective corresponding echo paths;

[0062] adjusting a convergence speed of the first filter according to the similarity, wherein the convergence speed decreases as the similarity increases.

[0063] As an optional implementation, when the echo information includes an echo signal, the adjusting convergence speed module is specifically configured to:

[0064] determining a second residual signal according to the echo information corresponding to the second filter and the near-end signal;

[0065] determining an energy ratio according to the power of the first residual signal and the power of the second residual signal;

[0066] adjusting a convergence speed of the first filter according to the energy ratio, wherein the convergence speed increases as the energy ratio increases.

[0067] As an optional implementation, the adjustment convergence speed module is specifically used for:

[0068] adjusting a noise covariance matrix of the first filter according to the respective corresponding echo information, wherein the noise covariance matrix is used to represent a degree of deviation of echo path estimation of the first filter;

[0069] adjusting a convergence speed of the first filter according to the noise covariance matrix.

[0070] As an optional implementation, the adjustment convergence speed module is specifically used for:

[0071] determining a convergence state of the first filter according to the respective corresponding echo information, and adjusting a convergence speed of the first filter according to the convergence state of the first filter.

[0072] As an optional implementation, a steady-state error of the first filter is less than a steady-state error of the second filter, and / or, a convergence speed of the second filter is greater than a convergence speed of the first filter.

[0073] In a fourth aspect, the embodiments of the present application further provide a computer storage medium, which has a computer program stored thereon, and the computer program is executed by a processor to implement the steps of the method in the first aspect.

[0074] These and other aspects of the present application will become apparent from the following description of the embodiments taken in conjunction with the accompanying drawings, although variations of these embodiments may be understood to exist in the field and may be practiced and carried out in various ways based on the teachings disclosed herein. BRIEF DESCRIPTION OF DRAWINGS

[0075] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can be obtained from these drawings without any creative effort.

[0076] Figure 1 An application scenario of echo cancellation provided by the embodiments of the present application;

[0077] Figure 2 A specific implementation flowchart of an echo cancellation method provided by the embodiments of the present application;

[0078] Figure 3 A specific implementation method flowchart of echo cancellation provided by the embodiments of the present application;

[0079] Figure 4 A schematic diagram of an electronic device provided by the embodiments of the present application;

[0080] Figure 5 A schematic diagram of an echo cancellation device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0081] In order to make the objects, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.

[0082] In the embodiments of the present application, the term "and / or" describes the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can mean that A exists alone, A and B exist together, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after it.

[0083] The application scenarios described in the embodiments of the present application are used to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those of ordinary skill in the art can know that, as new application scenarios appear, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems. In the description of the present application, unless otherwise specified, the meaning of "multiple" is two or more.

[0084] In a video conference, people usually use a hands-free or video conference terminal to talk, and the sound played by a loudspeaker and the near-end speech will be collected by a microphone at the same time, such as Figure 1As shown in the illustration, this embodiment provides an application scenario diagram for echo cancellation. The microphone collects both the user's voice signal and the sound played from a distant speaker, thus requiring echo cancellation. The performance of acoustic echo cancellation, as a crucial indicator in voice interaction systems, significantly impacts the communication experience between the user and the device, or between users themselves. Traditional acoustic echo cancellation typically includes two main modules: linear echo cancellation and echo post-processing. Linear echo cancellation can usually employ algorithms such as Normalized Least Mean Square (NLMS), Affine projection (AP), Recursive Least Squares (RLS), and Kalman filtering to obtain a linear residual signal. In cases of abrupt changes in the echo path, the adaptive algorithm needs to reconverge to estimate the current true echo path. During convergence, the linear echo cancellation effect deteriorates sharply compared to the steady-state condition, resulting in large-area echo leakage. While current echo cancellation algorithms can remove echo components from the acquired signal to some extent, their performance deteriorates drastically when the echo path changes abruptly. The echo cancellation effect also deteriorates sharply compared to the steady-state condition, resulting in large-area echo leakage.

[0085] Echo-path change detector (ECD) is a crucial step in improving the robustness of echo cancellation algorithms. The echo cancellation method provided in this embodiment extracts the different characteristics exhibited by the dual filters during steady-state and convergence processes. Based on the corresponding echo information, it detects whether a sudden change in the echo path has occurred in the current environment. When such a change occurs, it takes aggressive echo suppression measures to quickly suppress the echo and avoid large-scale echo leakage.

[0086] like Figure 2 As shown in the figure, the specific implementation process of the echo cancellation method provided in this embodiment is as follows:

[0087] Step 200: Acquire an audio signal, which includes a near-end signal and a far-end signal, and the near-end signal and the far-end signal are distinguished based on the propagation mode of the audio signal;

[0088] In implementation, the propagation mode includes a signal propagated through a playing device such as a speaker / loudspeaker, or a signal propagated through a non-playing device such as a signal propagated through air. The far-end signal includes a voice signal played through the speaker / loudspeaker, and the near-end signal includes a voice signal in the environment other than the voice signal played through the speaker / loudspeaker, such as user voice, ambient noise, and the like. The near-end signal includes but is not limited to a voice signal collected by an audio collection device and containing echo, and the far-end signal includes but is not limited to a far-end played voice signal. For example, the near-end voice signal can be understood as user voice picked up by a microphone, and the far-end voice signal can be understood as a sound played through a loudspeaker / speaker of the near-end device and transmitted to the near-end device through a network or the like.

[0089] Step 201, filtering the far-end signal by using a first filter and a second filter respectively to obtain respective corresponding echo information, wherein the first filter has higher ability to eliminate echo when the echo path is stable than the second filter, and the second filter has greater convergence speed when the echo path changes than the first filter.

[0090] In some embodiments, before filtering the far-end signal by using the first filter and the second filter, the acquired audio signal can be converted from time domain to frequency domain. For example, the conversion process is as follows for a microphone collected audio signal:

[0091] First, the near-end signal collected by the microphone in the echo cancellation is denoted as d(n), and the far-end signal played through the loudspeaker is denoted as x(n). Second, the near-end signal is framed and short-time Fourier transformed to obtain a frequency domain near-end signal, and the far-end signal is also framed and short-time Fourier transformed to obtain a frequency domain far-end signal.

[0092] In this embodiment, the framing method can adopt an overlapping segmentation method, and the overlap between frames is a frame shift, where the frame shift value is half of the frame length. The near-end signal d(n) and the far-end signal x(n) in the current frame audio signal collected by the microphone are converted from time domain signals to frequency domain signals by using the short-time Fourier transform method to obtain frequency domain signals of the current frame audio signal. The frequency domain near-end signal in the converted frequency domain signal of the current frame audio signal is denoted as D(k), and the frequency domain far-end signal in the converted frequency domain signal of the current frame audio signal is denoted as X(k).

[0093] In implementation, the frequency domain far-end signal of the far-end signal is filtered by using the first filter and the second filter respectively to obtain respective corresponding echo information.

[0094] In some embodiments, the echo information in the present embodiment includes, but is not limited to, at least one of an echo path and an echo signal, wherein the echo signal can be calculated through the echo path and a frequency domain far-end signal, and specifically can be determined through the following formula:

[0095] Y shadow (k)=W shadow (k, l)X(k) Formula (1);

[0096] wherein, W shadow (k, l) represents an echo path (a frequency domain signal), Y shadow (k) represents an echo signal (a frequency domain signal), and X(k) represents a frequency domain far-end signal; k represents a frequency point, and l represents a time frame serial number.

[0097] In the present embodiment, a double filter is used, and the double filter has different characteristics. Both the first filter and the second filter can perform linear filtering processing on the far-end signal, but the first filter is used to output the final linear filtering result, and the second filter is used to perform foreground prediction and timely track the rapid path change of the echo path. By using the different characteristics of the double filter, the echo can be better eliminated in a steady state, the echo can be quickly suppressed when the path suddenly changes, and large-area echo leakage can be avoided.

[0098] In some embodiments, the steady-state error of the first filter is smaller than the steady-state error of the second filter, and / or the convergence speed of the second filter is greater than the convergence speed of the first filter.

[0099] It should be noted that the convergence speed and the steady-state error of the filter are usually contradictory. When the filter has a faster convergence speed, it usually brings a larger steady-state error, and vice versa. When the filter has a lower steady-state error, the convergence speed is usually slower. In the present embodiment, the faster convergence speed of the second filter is used to quickly track the echo path, and the lower steady-state error of the first filter is used as the final linear output.

[0100] In some embodiments, the filtering algorithm used by the first filter / second filter in the present embodiment includes, but is not limited to, at least one of an LMS (Least Mean Square) algorithm, an NLMS (Normalized LMS) algorithm, an RLS (Recursive Least Square) algorithm, and a KALMAN (Kalman) algorithm.

[0101] Step 202, adjusting the echo suppression aggressiveness according to the respective corresponding echo information, and determining a first residual signal by using the echo information corresponding to the first filter and the near-end signal;

[0102] In some embodiments, when the echo information comprises echo paths, the echo suppression aggressiveness can be adjusted according to the similarity of the echo paths, specifically:

[0103] determining the similarity between the respective corresponding echo paths; and adjusting the echo suppression aggressiveness according to the similarity, wherein the echo suppression aggressiveness increases with the increase of the similarity.

[0104] In some embodiments, since the main function of the second filter in the embodiment is to track the change of the fast echo path in time, a filter algorithm with faster convergence is used, for example, the NLMS algorithm. Taking the NLMS algorithm with a larger step size as an example, the frequency domain form of the echo path estimated by the second filter is denoted as W shadoω (k, l).

[0105] Since the first filter in the embodiment is the final linear output result filter, its function is to eliminate echo as much as possible, therefore, a filter with smaller steady-state error is used, for example, the KALMAN filter, which has a lower steady-state error after convergence and can better remove echo. Taking the first filter as the KALMAN filter as an example, the frequency domain form of the echo path estimated by the first filter is denoted as W main (k, l).

[0106] Optionally, the similarity between the echo paths is determined by the following formula:

[0107] similarity(k) = Distance(W main (k, l), W shadow (k, l)) Formula (2).

[0108] wherein, similarity(k) represents the similarity, Distance() represents the distance between the echo paths, W shadow (k, l) represents the echo path corresponding to the second filter, and W main (k, l) represents the echo path corresponding to the first filter. The distance calculation method in the embodiment can use the cosine similarity, cepstrum distance, KL divergence, etc. The embodiment does not make too many limitations on this.

[0109] In some embodiments, the state of the first filter is determined according to the similarity between the respective corresponding echo paths, and the echo suppression aggressiveness is adjusted according to the state of the first filter. Wherein, the similarity changes between [0, 1], the closer to 0 the similarity is, the greater the difference between W shadow (k, l) and W main (k, l), the closer to 1 the similarity is, the smaller the difference between W shadow (k, l) and W mainThe smaller the difference is, the more similar (k, l) is; the greater the similarity is, the better the convergence state of the first filter is; the smaller the similarity is, the more divergent the state of the first filter is; the better the convergence state of the first filter is, the smaller the aggressiveness of the echo suppression is, the more divergent the first filter is, the greater the aggressiveness of the echo suppression is, wherein the aggressiveness of the echo suppression varies within a certain range.

[0110] In implementation, the relationship between the state of the first filter and the similarity in the embodiment is as follows:

[0111]

[0112] wherein L1 can be 0, when the similarity tends to 0, the first filter is in the convergence state, when the similarity is greater than 0 and less than L2, the first filter is in the under-filtering state, and when the similarity is greater than L2, the first filter is in the divergent state.

[0113] Optionally, the aggressiveness of the echo suppression is adjusted according to the state of the first filter in the following manner:

[0114]

[0115] wherein gamma represents the aggressiveness of the echo suppression, the value range of gamma is [0, 1], the higher the value of gamma is, the greater the suppression of the echo is, and the more serious the damage to the near-end speech is. The aggressiveness of the echo suppression is adjusted according to the state of the first filter, for adapting to the current environment, wherein grade_1, grade_2 and grade_3 respectively correspond to different aggressiveness of the echo suppression, and in general, grade_1 < grade_2 < grade_3.

[0116] In some embodiments, when the echo information comprises an echo signal, the aggressiveness of the echo suppression can be adjusted according to the energy ratio of the residual signal, specifically as follows:

[0117] Step 1) determining a second residual signal according to the echo information corresponding to the second filter and the near-end signal;

[0118] In implementation, the second residual signal is determined according to the difference between the near-end signal and the echo signal corresponding to the second filter, specifically by the following formula:

[0119] e shadow (n) = d(n) - y shadow (k) of formula (3);

[0120] wherein Y shadow (k) represents the echo signal corresponding to the second filter, y shadow (n) represents Y shadow(k) is a time domain form of Y shadow (n) represents the second residual signal.

[0121] Step 2) determining an energy ratio according to a power of the first residual signal and a power of the second residual signal.

[0122] In an implementation, the first residual signal is determined by using a difference between the near-end signal and echo information corresponding to the first filter, specifically by the following formula:

[0123] Y main (k) = W main (k, l)X(k) Formula (4).

[0124] e main (n) = d(n) - y main (n) Formula (5).

[0125] In Formula (4), W main (k, l) represents an estimated echo path of the first filter (a frequency domain signal), X(k) represents a frequency domain far-end signal, Y main (k) represents an echo signal corresponding to the first filter (a frequency domain signal); k represents a frequency point, and l represents a time frame serial number.

[0126] In Formula (5), y main (n) represents a time domain form of Y main (k), d(n) represents an obtained near-end signal (a time domain signal), e main (n) represents the first residual signal (a time domain signal).

[0127] Optionally, an energy ratio is determined according to a ratio of a power of the first residual signal and a power of the second residual signal.

[0128] Step 3) adjusting an echo suppression aggressiveness according to the energy ratio, wherein the echo suppression aggressiveness increases with an increase of the energy ratio. Wherein the echo suppression aggressiveness increases with the increase of the energy ratio within a certain range.

[0129] Optionally, the echo suppression aggressiveness can be adaptively adjusted based on the energy ratio, for example, a plurality of intervals of energy ratio can be set, and each interval of energy ratio corresponds to an echo suppression aggressiveness, when the calculated energy ratio is in a certain interval, the echo suppression aggressiveness corresponding to the interval is taken as the final adjustment target, and the current echo suppression aggressiveness is adjusted to the adjustment target. A curve relationship between the energy ratio and the echo suppression aggressiveness can also be set, for example, it can be a linear relationship, a nonlinear relationship, etc., as long as the curve relationship satisfies the condition that the echo suppression aggressiveness increases with the increase of the energy ratio. How to adjust the echo suppression aggressiveness based on the energy ratio can be adjusted according to actual needs, and the embodiment does not make too many limitations on this.

[0130] When the energy ratio exceeds the threshold value, it indicates that the filter is in a divergent state, and the first filter can be reset.

[0131] In some embodiments, the state of the first filter is determined according to the respective corresponding echo information, and the echo suppression aggressiveness is adjusted according to the state of the first filter.

[0132] Optionally, the state of the first filter is determined according to the energy ratio, and the echo suppression aggressiveness is adjusted according to the state of the first filter. The smaller the energy ratio, the better the convergence state of the first filter, and the closer the echo path to the real echo path; the better the convergence state of the first filter, the smaller the echo suppression aggressiveness, and the echo suppression aggressiveness changes within a certain range.

[0133] In the implementation, the power of the first residual signal is calculated by the following formula:

[0134]

[0135] Wherein, β is a preset value, representing a smoothing factor; e main (i) represents the first residual signal, P main (n) represents the power of the first residual signal.

[0136]

[0137] Wherein, β is a preset value, representing a smoothing factor; e shadow (i) represents the second residual signal, P shadow (n) represents the power of the second residual signal.

[0138] In the implementation, P main / P shadow represents the energy ratio, wherein when the first filter is in a steady state, P main / P shadowGenerally less than 1. The state of the first filter can be determined according to the energy ratio in the following way:

[0139]

[0140] wherein T1 represents a value greater than 1 and close to 1, T2 > 1 and T2 > T1. When the energy ratio is close to 1, the first filter is in a converging state, when the energy ratio is greater than 1 and less than T2, the first filter is in an under-filtering state, and when the energy ratio is greater than or equal to T2, the first filter is in a diverging state.

[0141] It should be noted that when the environment is relatively stable, the first filter converges after a sufficient time and generally enters the converging state, at which time the energy of the first residual signal corresponding to the first filter is less than the energy of the second residual signal corresponding to the second filter. When there are slight changes in the environment such as people walking, moving tables and chairs, etc., the filter generally enters the under-filtering state, at which time the energy of the first residual signal corresponding to the first filter is slightly less than or equal to the energy of the second residual signal corresponding to the second filter. When the collection device is moved, the conference room door is opened and closed, the loudspeaker is blocked, etc., the first filter can enter the diverging state.

[0142] The threshold values T1 and T2 can be set according to actual use. In general, the first filter is in the converging state or in part of the under-converging state when P main (n) < P shadow (n).

[0143] After determining the state of the filter according to the energy ratio, the aggressiveness of the echo suppression can also be adjusted according to the state of the first filter, as shown below:

[0144]

[0145] wherein gamma represents the aggressiveness of the echo suppression, gamma takes a value in the range [0, 1], the higher the value of gamma, the greater the suppression of the echo, and the more serious the damage to the near-end speech. The aggressiveness of the echo suppression is adjusted according to the state of the first filter described above, which is used to adapt to the current environment, wherein grade_1, grade_2 and grade_3 correspond to different aggressiveness of the echo suppression, and in general grade_1 < grade_2 < grade_3.

[0146] In some embodiments, when adjusting the aggressiveness of the echo suppression according to the respective corresponding echo information, the state of the first filter can be determined according to the respective corresponding echo information first, and the aggressiveness of the echo suppression is adjusted according to the state of the first filter. The more converging the first filter is, the smaller the aggressiveness of the echo suppression is; the more diverging the first filter is, the greater the aggressiveness of the echo suppression is.

[0147] Step 203, suppress the first residual signal by using the adjusted echo suppression aggressiveness, to obtain the audio signal after echo cancellation.

[0148] In some embodiments, the present embodiment can adjust the convergence speed of the first filter according to the respective corresponding echo information while adjusting the echo suppression aggressiveness according to the respective corresponding echo information; filter the far-end signal by using the adjusted first filter to obtain corresponding new echo information, continue to adjust the echo suppression aggressiveness according to the new echo information and the echo information corresponding to the second filter, and suppress the first residual signal by using the continuously adjusted echo suppression aggressiveness to obtain the audio signal after echo cancellation.

[0149] It should be noted that the present embodiment performs real-time echo cancellation for each frame of received audio signal in the process of echo cancellation, and after the echo in the current frame of audio signal is suppressed by using the adjusted echo suppression aggressiveness, the next frame of audio signal is filtered by using the first filter with adjusted convergence speed, and then the echo information obtained after filtering and the residual signal of the near-end signal are suppressed by using the echo suppression aggressiveness adjusted again to obtain the next frame of audio signal after echo cancellation.

[0150] In the present embodiment, a double filter is used. Since the double filter has different characteristics, the first filter and the second filter can both perform linear filtering on the far-end signal, but the first filter is used to output the final linear filtering result, and the second filter is used for foreground prediction and timely tracking of rapid path changes of the echo path. By using the different characteristics of the double filter, the echo can be better eliminated in the steady state, the first filter can quickly converge when the path suddenly changes, the echo can be quickly suppressed, and large-area echo leakage can be avoided.

[0151] It should be noted that the convergence speed of the filter is usually contradictory to the steady-state error. When the filter has a faster convergence speed, it usually brings a larger steady-state error, and vice versa. When the filter has a lower steady-state error, the convergence speed is usually slower. In the present embodiment, when the path mutation is detected, the convergence speed of the first filter is adjusted to quickly suppress the echo and avoid large-area echo leakage.

[0152] In some embodiments, when the echo information includes an echo path, the convergence speed of the first filter is adjusted according to the similarity of the echo paths, specifically:

[0153] Determine the similarity between the respective corresponding echo paths; adjust the convergence speed of the first filter according to the similarity, wherein the convergence speed decreases as the similarity increases.

[0154] In the implementation, the first filter is taken as an example of KALMAN filter, and the echo path estimated by the first filter is denoted as W main (k, l) in frequency domain. The second filter is taken as an example of NLMS algorithm with large step size, and the echo path estimated by the second filter is denoted as W shadow (k, l) in frequency domain.

[0155] In the implementation, W main (k, l) can be iteratively updated according to the first residual signal and the planned step size, and the update rule is as follows:

[0156] W main (k, l+1) = AW + (k, l) in formula (7);

[0157] Wherein, A is a preset value, and A is usually defined as a value close to 1 when the echo path is assumed to be slowly changing. W + (k, l) represents an intermediate variable for calculating the echo path of the far-end signal of the current audio signal frame.

[0158] W+m ain (k, l) = W main (k, l) + K(k)E main (k) in formula (8);

[0159] Wherein, W main (k, l) represents the frequency domain form of the echo path estimated by the first filter, E main (k) represents the frequency domain form of the first residual signal, and K(k) represents the Kalman gain, which can be understood as a form of step size in the echo path estimation process.

[0160] P(k, l+1) = A 2 P + (k, l) + ψ ΔΔ (k) in formula (9);

[0161] Wherein, P(k, l+1) represents the noise covariance matrix, P + (k, l) represents an intermediate variable for calculating the noise covariance matrix, A is a preset value, and ψ ΔΔ (k) represents the process noise covariance matrix, which can be understood as the degree of deviation between the current echo path estimated by the first filter and the real echo path obtained in the estimation process. The process noise covariance matrix has a greater impact on the overall algorithm effect, and can be calculated by using various estimation methods. The present embodiment does not make too many limitations, and one of the estimation methods can be represented as: ψ ΔΔ (k) = (1-A 2 )E[W main (k, l)WH main (k, l)].

[0162]

[0163] wherein, I M denotes a unit diagonal matrix, C(k) denotes a fixed coefficient matrix, denotes an adjusted noise covariance matrix; K(k) denotes a Kalman gain, k denotes a frequency point, and l denotes a time frame serial number.

[0164]

[0165] wherein, P(k, l) denotes a noise covariance matrix of the first filter, and a denotes an adjustment parameter, the value of a directly affects the echo path tracking ability and the steady-state error size, and thus the value size can be monitored in real time to control the value size, so that the filter converges quickly when the path changes, and the steady-state error is as low as possible.

[0166]

[0167] wherein, K(k) denotes a Kalman gain, P(k, l) denotes a noise covariance matrix of the first filter, C(k) denotes a fixed coefficient matrix, and ss (k) denotes an observation noise in Kalman filtering, wherein ss (k) can be expressed as SS (k) = E[E main (k)E main H (k)], E main (k) denotes a frequency domain form of the first residual signal. B denotes a block number of frequency domain division, and b denotes an index value of the block number.

[0168] In the implementation, W shadow (k, l) can be iteratively updated according to the second residual signal and the planning step size, and the update rule is as follows:

[0169]

[0170] wherein, X(k) denotes a frequency domain far-end signal, μ(k) denotes a planning step size, μ and Δ are preset fixed values, μ is a step size, which can be set to 0.5, and Δ is a variable for preventing the denominator from being 0, which can be set to a small value, for example, 1e-10.

[0171] W shadow (k, l+1) = W shadow (k, l) + μ(k)X H (k)E(k) formula (14);

[0172] wherein X(k) represents a frequency domain far-end signal, k represents a frequency point, I represents a time frame number, μ(k) represents a planning step, E(k) represents a second residual error signal e shadow (n) in a frequency domain; W shadow (k, I) represents an echo path of the Ith frame far-end signal filtered by the second filter, W shadow (k, I+1) represents an echo path of the (I+1)th frame far-end signal filtered by the second filter.

[0173] Optionally, the similarity between the echo paths is determined by the above formula (2).

[0174] In some embodiments, a state of the first filter is determined according to the similarity between the respective corresponding echo paths, and a convergence speed of the first filter is adjusted according to the state of the first filter.

[0175] In some embodiments, the adjusting the convergence speed of the first filter according to the respective corresponding echo information comprises:

[0176] adjusting a noise covariance matrix of the first filter according to the respective corresponding echo information, wherein the noise covariance matrix is used to represent a degree of deviation of the first filter in echo path estimation; and determining the convergence speed of the first filter according to the adjusted noise covariance matrix.

[0177] In some embodiments, a state of the first filter is first determined according to the similarity between the respective corresponding echo paths, a noise covariance matrix of the first filter is adjusted according to the state of the first filter, and the convergence speed of the first filter is determined according to the adjusted noise covariance matrix. The smaller the similarity, the more convergent the first filter; the greater the similarity, the more divergent the first filter; and the greater the similarity, the more stable the current path, and the echo cancellation capability in a steady state can be improved by reducing the convergence speed.

[0178] Step 1) determining a state of the first filter according to the similarity between the respective corresponding echo paths.

[0179] The relationship between the state of the first filter and the similarity in the embodiment is as follows:

[0180]

[0181] wherein L1 can be 0, the first filter is in a convergent state when the similarity tends to 0, the first filter is in an under-filtered state when the similarity is greater than 0 and less than L2, and the first filter is in a divergent state when the similarity is greater than L2.

[0182] Step 2) adjusting the noise covariance matrix of the first filter according to the state of the first filter.

[0183] Optionally, the noise covariance matrix of the first filter is adjusted according to the state of the first filter in the following manner:

[0184]

[0185] wherein the adjustment parameter a is determined according to the state of the first filter, and the noise covariance matrix P(k) of the first filter is adjusted by using the adjustment parameter a, i.e. the adjusted noise covariance matrix is obtained according to the above formula (11)

[0186] Step 3) adjusting the convergence speed of the first filter according to the noise covariance matrix.

[0187] The size of the noise covariance matrix in the first filter is corrected by using the adjustment parameter a. When the filter is in a converging state, T3 takes a smaller value to ensure a lower steady-state error. When the filter needs to converge, the value of a is increased to achieve the purpose of quickly tracking the echo path. Generally, T3 < T4 < T5 can be set.

[0188] In some embodiments, when the echo information comprises an echo signal, the convergence speed of the first filter can be adjusted according to the energy ratio of the residual error signal, specifically:

[0189] Step 1) determining a second residual error signal according to the echo information corresponding to the second filter and the near-end signal;

[0190] In implementation, the second residual error signal e shadow (n) can be determined by the above formula (3), and the first residual error signal e main (n) can be determined according to formula (5), which will not be described here.

[0191] Step 2) determining an energy ratio according to the power of the first residual error signal and the power of the second residual error signal;

[0192] Optionally, the energy ratio is determined according to the ratio of the power of the first residual error signal and the power of the second residual error signal.

[0193] In some embodiments, the energy ratio can also be determined by the absolute value of the difference between the power of the first residual error signal and the power of the second residual error signal. The energy ratio in this embodiment represents the energy of the residual error signal, and the manner of determining the energy ratio based on the residual error signal is not limited in this embodiment.

[0194] Step 3) adjusting the convergence speed of the first filter according to the energy ratio, wherein the convergence speed increases with the increase of the energy ratio.

[0195] Optionally, the convergence speed can be adaptively adjusted based on the energy ratio. For example, a plurality of intervals of energy ratio can be set, and each interval of energy ratio corresponds to a convergence speed. When the calculated energy ratio is in a certain interval, the convergence speed corresponding to the interval is taken as the final adjustment target, and the convergence speed of the current first filter is adjusted to the adjustment target. A curve relationship between the energy ratio and the convergence speed can also be set, which can be a linear relationship, a nonlinear relationship, etc., as long as the curve relationship satisfies the condition that the convergence speed increases with the increase of the energy ratio. How to adjust the convergence speed based on the energy ratio can be adjusted according to actual needs, and this embodiment does not make too many limitations.

[0196] In some embodiments, the convergence state of the first filter is determined according to the respective corresponding echo information, and the convergence speed of the first filter is adjusted according to the convergence state of the first filter.

[0197] Optionally, the state of the first filter is determined according to the energy ratio, the noise covariance matrix of the first filter is adjusted according to the state of the first filter, and the convergence speed of the first filter is determined according to the adjusted noise covariance matrix of the first filter. The smaller the energy ratio, the more convergent the filter, the larger the energy ratio, the more divergent the filter, the smaller the noise covariance matrix, the more convergent the filter, and the larger the noise covariance matrix, the more divergent the filter.

[0198] In the implementation, the power of the first residual signal can be calculated according to the above formula (6), and the power of the second residual signal can be obtained according to the above formula (7).

[0199] In the implementation, P main / P shadow represents the energy ratio, wherein when the first filter is in a steady state, P main / P shadow is usually a value less than 1. The state of the first filter can be determined according to the energy ratio in the following manner:

[0200]

[0201] wherein T1 represents a value greater than 1 and close to 1, and T2>1 and T2>T1. When the energy ratio is close to 1, the first filter is in a convergent state, when the energy ratio is greater than 1 and less than T2, the first filter is in an under-filtered state, and when the energy ratio is greater than or equal to T2, the first filter is in a divergent state.

[0202] After determining the state of the filter according to the energy ratio, the noise covariance matrix of the first filter can be adjusted according to the state of the first filter:

[0203]

[0204] wherein the adjustment parameter a is determined according to the state of the first filter, and the noise covariance matrix P(k) of the first filter is adjusted by using the adjustment parameter a, that is, the adjusted noise covariance matrix P'(k) is obtained according to the above formula (11) the convergence speed of the first filter is determined according to the adjusted noise covariance matrix P'(k)

[0205] The embodiment adopts the design of the double-filter architecture, takes the first filter with the variable step as the main filter, which has good echo suppression capability in the steady state; and takes the second filter with the large step as the auxiliary filter, which has a fast convergence speed. In the stable echo environment and the changing echo environment, the main and auxiliary filters show different characteristics, and the estimated echo path similarity or the energy difference of the residual signal can be selected as the feature to determine the state of the first filter. When the path change is detected, the first filter needs to converge as soon as possible, and the noise covariance matrix correction coefficient a in the first filter is increased, so that the first filter can converge quickly, the echo suppression aggressiveness is increased, and a large area of echo residue is avoided. When the steady state is detected, the adjustment parameter a and the echo suppression aggressiveness are reduced, and the degree of echo cancellation and the effect of the near-end speech are ensured.

[0206] As shown in Figure 3 The embodiment also provides a specific implementation method of echo cancellation, and the implementation process of the method is as follows:

[0207] Step 300: obtaining a plurality of audio signal frames, each audio signal frame comprising a near-end signal and a far-end signal;

[0208] Step 301: performing filtering processing on the far-end signal contained in the current audio signal frame by using the first filter and the second filter respectively to obtain respective corresponding echo information;

[0209] Step 302: adjusting the echo suppression aggressiveness and the convergence speed of the first filter according to the respective corresponding echo information;

[0210] Optionally, the state of the first filter is determined according to the respective corresponding echo information, and the echo suppression aggressiveness and the convergence speed of the first filter are adjusted according to the state of the first filter. In the implementation, the state of the first filter can be determined according to the similarity between the respective corresponding echo paths, or the state of the first filter can be determined according to the energy ratio of the residual signal. The specific determination process is shown in the above steps, and will not be described here.​

[0211] The relationship between the state of the filter and the corresponding echo suppression aggressiveness, the convergence speed of the first filter is as follows:

[0212]

[0213] Wherein, gamma represents the echo suppression aggressiveness, gamma takes the value interval range [0, 1], the higher the value of gamma, the greater the echo suppression, and the more serious the damage to the near-end speech. The size of the noise covariance matrix in the first filter is corrected by using the adjustment parameter a, when the filter is in the convergent state, T3 takes a smaller value to ensure a lower steady-state error, when the filter needs to converge, the value of a is increased to achieve the purpose of quickly tracking the echo path. The noise covariance matrix P(k) of the first filter is adjusted by using the adjustment parameter a, that is, the adjusted noise covariance matrix is obtained according to the above formula (11) According to the noise covariance matrix, the convergence speed of the first filter is adjusted.

[0214] Step 303, using the echo information corresponding to the first filter and the near-end signal contained in the current audio signal frame, determining the first residual signal;

[0215] Step 304, using the adjusted echo suppression aggressiveness to suppress the first residual signal to obtain the current audio signal frame after echo cancellation;

[0216] Step 305, using the first filter and the second filter with adjusted convergence speed to filter the far-end signal contained in the next audio signal frame respectively to obtain the respective corresponding echo information;

[0217] Step 306, adjusting the echo suppression aggressiveness and the convergence speed of the first filter again according to the respective corresponding echo information;

[0218] Step 307, using the echo information corresponding to the first filter and the near-end signal contained in the next audio signal frame, determining the first residual signal;

[0219] Step 308, using the echo suppression aggressiveness adjusted again to suppress the first residual signal to obtain the next audio signal frame after echo cancellation.

[0220] The embodiment adopts a double-filter structure, the first filter can achieve a higher degree of echo suppression in the steady state, and the second filter can quickly track the changes of the echo path. In combination with the different characteristics of the double filters in the steady state and when the echo path changes, the state of the current first filter is judged, and the convergence speed of the filter and the echo suppression aggressiveness are adjusted according to the state of the filter, so that fast convergence is achieved while avoiding large-area echo leakage.

[0221] Based on the same inventive concept, the embodiments of the present application also provide an electronic device. Since the electronic device is the electronic device in the method of the embodiments of the present application, and the principle of solving the problem of the electronic device is similar to that of the method, the implementation of the electronic device can be referred to the implementation of the method, and the repeated parts will not be described herein.

[0222] As shown in Figure 4 The electronic device includes a processor 400 and a memory 401, the memory 401 is used to store programs executable by the processor 400, and the processor 400 is used to read the programs in the memory 401 and perform the following steps:

[0223] Obtain an audio signal, the audio signal includes a near-end signal and a far-end signal, the near-end signal and the far-end signal are distinguished based on the propagation mode of the audio signal;

[0224] Filter the far-end signal using a first filter and a second filter respectively to obtain respective corresponding echo information, wherein the first filter has higher ability to eliminate echo than the second filter when the echo path is stable, and the second filter has greater convergence speed than the first filter when the echo path changes;

[0225] Adjust the echo suppression aggressiveness according to the respective corresponding echo information, and determine a first residual signal using the echo information corresponding to the first filter and the near-end signal;

[0226] Suppress the first residual signal using the adjusted echo suppression aggressiveness to obtain an echo-canceled audio signal.

[0227] As an optional implementation, when the echo information includes an echo path, the processor 400 is specifically configured to perform:

[0228] Determine the similarity between the respective corresponding echo paths;

[0229] Adjust the echo suppression aggressiveness according to the similarity, wherein the echo suppression aggressiveness increases with the increase of the similarity.

[0230] As an optional implementation, when the echo information includes an echo signal, the processor 400 is specifically configured to perform:

[0231] Determine a second residual signal according to the echo information corresponding to the second filter and the near-end signal;

[0232] Determine an energy ratio according to the power of the first residual signal and the power of the second residual signal;

[0233] adjusting an echo suppression aggressiveness according to the energy ratio, wherein the echo suppression aggressiveness increases as the energy ratio increases.

[0234] As an optional implementation, the processor 400 is specifically configured to perform:

[0235] determining a state of a first filter according to the respective corresponding echo information, and adjusting an echo suppression aggressiveness according to the state of the first filter.

[0236] As an optional implementation, the processor 400 is specifically further configured to perform:

[0237] adjusting a convergence speed of the first filter according to the respective corresponding echo information;

[0238] filtering the far-end signal by using the adjusted first filter to obtain corresponding new echo information, and continuing to adjust the echo suppression aggressiveness according to the new echo information and the echo information corresponding to the second filter.

[0239] As an optional implementation, when the echo information comprises echo paths, the processor 400 is specifically configured to perform:

[0240] determining a similarity between the respective corresponding echo paths;

[0241] adjusting a convergence speed of the first filter according to the similarity, wherein the convergence speed decreases as the similarity increases.

[0242] As an optional implementation, when the echo information comprises echo signals, the processor 400 is specifically configured to perform:

[0243] determining a second residual signal according to the echo information corresponding to the second filter and the near-end signal;

[0244] determining an energy ratio according to a power of the first residual signal and a power of the second residual signal;

[0245] adjusting a convergence speed of the first filter according to the energy ratio, wherein the convergence speed increases as the energy ratio increases.

[0246] As an optional implementation, the processor 400 is specifically configured to perform:

[0247] adjusting a noise covariance matrix of the first filter according to the respective corresponding echo information, wherein the noise covariance matrix is used to represent a degree of deviation of echo path estimation by the first filter;

[0248] adjust the convergence speed of the first filter according to the noise covariance matrix.

[0249] As an optional implementation, the processor 400 is specifically configured to perform:

[0250] determine the convergence state of the first filter according to the respective corresponding echo information, and adjust the convergence speed of the first filter according to the convergence state of the first filter.

[0251] As an optional implementation, the processor 400 is specifically configured to perform:

[0252] The steady-state error of the first filter is smaller than the steady-state error of the second filter, and / or the convergence speed of the second filter is greater than the convergence speed of the first filter.

[0253] Based on the same inventive concept, the embodiment of the present application also provides a device for echo cancellation. Since the device is the device in the method of the embodiment of the present application, and the principle of solving problems of the device is similar to that of the method, the implementation of the device can be referred to the implementation of the method, and the repeated parts will not be described here.

[0254] As shown in Figure 5 the device comprises:

[0255] The audio acquisition module 500 is configured to acquire an audio signal, wherein the audio signal comprises a near-end signal and a far-end signal, and the near-end signal and the far-end signal are distinguished based on a propagation mode of the audio signal.

[0256] The double-filtering module 501 is configured to filter the far-end signal using a first filter and a second filter respectively to obtain respective corresponding echo information, wherein the first filter has higher echo cancellation capability than the second filter when the echo path is stable, and the second filter has a greater convergence speed than the first filter when the echo path changes.

[0257] The residual calculation module 502 is configured to adjust an echo suppression aggressiveness according to the respective corresponding echo information, and determine a first residual signal using the echo information corresponding to the first filter and the near-end signal.

[0258] The echo suppression module 503 is configured to suppress the first residual signal using the adjusted echo suppression aggressiveness to obtain an echo-canceled audio signal.

[0259] As an optional implementation, when the echo information comprises an echo path, the residual calculation module 502 is specifically configured to:

[0260] determine a similarity between the respective corresponding echo paths;

[0261] adjust an echo suppression aggressiveness according to the similarity, wherein the echo suppression aggressiveness increases as the similarity increases.

[0262] As an optional implementation, when the echo information comprises echo signals, the calculating residual module 502 is specifically configured to:

[0263] determine a second residual signal according to the echo information corresponding to the second filter and the near-end signal;

[0264] determine an energy ratio according to the power of the first residual signal and the power of the second residual signal;

[0265] adjust an echo suppression aggressiveness according to the energy ratio, wherein the echo suppression aggressiveness increases as the energy ratio increases.

[0266] As an optional implementation, the calculating residual module 502 is specifically configured to:

[0267] determine a state of the first filter according to the respective corresponding echo information, and adjust an echo suppression aggressiveness according to the state of the first filter.

[0268] As an optional implementation, the method further comprises adjusting a convergence speed module specifically configured to:

[0269] adjust a convergence speed of the first filter according to the respective corresponding echo information;

[0270] filter the far-end signal by using the adjusted first filter to obtain corresponding new echo information, and continue adjusting the echo suppression aggressiveness according to the new echo information and the echo information corresponding to the second filter.

[0271] As an optional implementation, when the echo information comprises echo paths, the adjusting convergence speed module is specifically configured to:

[0272] determine a similarity between the respective corresponding echo paths;

[0273] adjust a convergence speed of the first filter according to the similarity, wherein the convergence speed decreases as the similarity increases.

[0274] As an optional implementation, when the echo information comprises echo signals, the adjusting convergence speed module is specifically configured to:

[0275] determine a second residual signal according to the echo information corresponding to the second filter and the near-end signal;

[0276] determine an energy ratio according to the power of the first residual signal and the power of the second residual signal;

[0277] adjust the convergence speed of the first filter according to the energy ratio, wherein the convergence speed increases with the increase of the energy ratio.

[0278] As an optional implementation, the adjustment convergence speed module is specifically configured to:

[0279] adjust a noise covariance matrix of the first filter according to the respective corresponding echo information, wherein the noise covariance matrix is used to represent the degree of deviation of the echo path estimation of the first filter;

[0280] adjust the convergence speed of the first filter according to the noise covariance matrix.

[0281] As an optional implementation, the adjustment convergence speed module is specifically configured to:

[0282] determine the convergence state of the first filter according to the respective corresponding echo information, and adjust the convergence speed of the first filter according to the convergence state of the first filter.

[0283] As an optional implementation, the steady-state error of the first filter is less than the steady-state error of the second filter, and / or the convergence speed of the second filter is greater than the convergence speed of the first filter.

[0284] Based on the same inventive concept, the embodiments of the present disclosure provide a computer storage medium, which includes computer program code, when the computer program code is run on a computer, the computer program code causes the computer to execute the method of echo cancellation as any one of the foregoing. Since the principle of solving problems of the above computer storage medium is similar to the method of echo cancellation, the implementation of the above computer storage medium can be referred to the implementation of the method, and the repeated parts will not be described here.

[0285] In the specific implementation process, the computer storage medium can include a universal serial bus flash drive (USB, Universal Serial Bus Flash Drive), a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various storage media that can store program codes.

[0286] Based on the same inventive concept, this disclosure also provides a computer program product, which includes computer program code that, when executed on a computer, causes the computer to perform any of the echo cancellation methods discussed above. Since the principle by which the above-described computer program product solves the problem is similar to that of the echo cancellation method, the implementation of the above-described computer program product can be referred to the implementation of the method, and repeated details will not be elaborated further.

[0287] Computer program products may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0288] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage and optical storage) containing computer-usable program code.

[0289] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 Devices that specify the functions in one or more boxes.

[0290] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including an instruction device, which is implemented in a processFigure 1 one or more processes and / or functions described in the one or more blocks. Figure 1 one or more blocks or multiple blocks.

[0291] These computer program instructions can also be loaded into computer or other programmable data processing devices, so that a series of operation steps are performed on the computer or other programmable devices to generate computer-implemented processes, thus the instructions executed on the computer or other programmable devices provide a process for implementing the functions described in the flowchart Figure 1 one or more processes and / or functions described in the one or more blocks. Figure 1 one or more blocks or multiple blocks.

[0292] Obviously, those skilled in the art can make various modifications and variations to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application belong to the scope of the claims of the present application and their equivalent technologies, the present application also intends to include these modifications and variations.

Claims

1. A method for echo cancellation, characterized in that, The method includes: Acquire an audio signal, the audio signal including a near-end signal and a far-end signal, the near-end signal and the far-end signal being distinguished based on the propagation mode of the audio signal; The far-end signal is filtered by a first filter and a second filter respectively to obtain their respective echo information. The first filter has a higher ability to eliminate echo when the echo path is stable than the second filter, and the second filter has a higher convergence speed when the echo path changes than the first filter. The echo suppression aggression is adjusted according to the respective echo information, and a first residual signal is determined using the echo information corresponding to the first filter and the near-end signal; when the echo information includes an echo signal, the step of adjusting the echo suppression aggression according to the respective echo information includes: determining a second residual signal according to the echo information corresponding to the second filter and the near-end signal; determining an energy ratio according to the power of the first residual signal and the power of the second residual signal; adjusting the echo suppression aggression according to the energy ratio, wherein the echo suppression aggression increases as the energy ratio increases; The first residual signal is suppressed by using the adjusted echo suppression aggressiveness to obtain the echo-cancelled audio signal.

2. The method according to claim 1, characterized in that, When the echo information includes an echo path, adjusting the echo suppression aggressiveness based on the corresponding echo information includes: Determine the similarity between the respective echo paths; The echo suppression aggressiveness is adjusted based on the similarity, wherein the echo suppression aggressiveness increases as the similarity increases.

3. The method according to claim 1 or 2, characterized in that, The step of adjusting the echo suppression aggressiveness according to the corresponding echo information includes: The state of the first filter is determined based on the corresponding echo information, and the echo suppression intensity is adjusted according to the state of the first filter.

4. The method according to claim 1 or 2, characterized in that, The method further includes: The convergence speed of the first filter is adjusted according to the corresponding echo information. The far-end signal is filtered using the adjusted first filter to obtain corresponding new echo information. The echo suppression intensity is then further adjusted based on the new echo information and the echo information corresponding to the second filter.

5. The method according to claim 4, characterized in that, When the echo information includes an echo path, adjusting the convergence speed of the first filter according to the respective corresponding echo information includes: Determine the similarity between the respective echo paths; The convergence speed of the first filter is adjusted according to the similarity, wherein the convergence speed decreases as the similarity increases.

6. The method according to claim 4, characterized in that, When the echo information includes echo signals, adjusting the convergence speed of the first filter according to the respective corresponding echo information includes: The second residual signal is determined based on the echo information corresponding to the second filter and the near-end signal; The energy ratio is determined based on the power of the first residual signal and the power of the second residual signal; The convergence speed of the first filter is adjusted according to the energy ratio, wherein the convergence speed increases as the energy ratio increases.

7. The method according to claim 4, characterized in that, The step of adjusting the convergence speed of the first filter according to the corresponding echo information includes: The noise covariance matrix of the first filter is adjusted according to the respective echo information, wherein the noise covariance matrix is ​​used to represent the degree of deviation of the first filter in echo path estimation. The convergence speed of the first filter is adjusted based on the noise covariance matrix.

8. The method according to claim 4, characterized in that, The step of adjusting the convergence speed of the first filter according to the corresponding echo information includes: The convergence state of the first filter is determined based on the corresponding echo information, and the convergence speed of the first filter is adjusted based on the convergence state of the first filter.

9. The method according to claim 1, characterized in that, The steady-state error of the first filter is less than the steady-state error of the second filter, and / or the convergence speed of the second filter is greater than the convergence speed of the first filter.

10. An electronic device, characterized in that, The electronic device includes a processor and a memory, the memory being used to store a program executable by the processor, and the processor being used to read the program in the memory and execute the steps of the method according to any one of claims 1 to 9.

11. A computer storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the method as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Echo canceller

    US20040161101A1