Echo suppression method, device, electronic device and storage medium

By acquiring the audio signal, performing linear filtering and voice activity detection, combined with similarity analysis, the audio status information is accurately determined, and the problem of inaccurate echo suppression radical factor in video conferences is solved, achieving better echo suppression effect.

CN115440236BActive Publication Date: 2025-08-29ZHEJIANG DAHUA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211061153.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-31
Publication Date
2025-08-29
Estimated Expiration
2042-08-31

AI Technical Summary

Technical Problem

The prior art cannot accurately determine whether the audio frame in a video conference is a single-speaking state or a double-speaking state, resulting in poor accuracy of the echo suppression radical factor and poor echo suppression effect.

Method used

By acquiring the audio signals of the audio acquisition device and the audio output device, linear filtering is performed to determine the estimated echo signal, combined with voice activity detection and similarity analysis, audio status information is determined, and the echo suppression progress factor is adjusted according to the status information for processing.

Benefits of technology

It improves the accuracy of the echo suppression radical factor, improves the echo suppression effect, and ensures the clarity and consistency of the audio signal.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115440236B_ABST
    Figure CN115440236B_ABST
Patent Text Reader

Abstract

The present application discloses an echo suppression method, apparatus, electronic device, and storage medium. The method obtains a first audio signal captured by an audio capture device and a second audio signal output by an audio output device, determines an estimated echo signal corresponding to the second audio signal through linear filtering, and determines a residual signal based on the first audio signal and the estimated echo signal. Voice activity detection is performed on the first audio signal and the second audio signal, respectively, and the similarity between the first audio signal and the estimated echo signal is determined. The audio state information is determined by combining the detection results and similarity of the voice activity detection performed on the first audio signal and the second audio signal, thereby making the determination of the audio state information more accurate, and thereby making the determination of the echo suppression aggressiveness factor more accurate, resulting in a better echo suppression effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of audio signal processing, and in particular to an echo suppression method, device, electronic device, and storage medium. Background Art

[0002] Echo suppression is essential in video conferencing. As a crucial metric in voice interaction systems, acoustic echo suppression performance significantly impacts the communication experience between users and devices, and between users themselves. Acoustic echo cancellation typically involves two steps: linear echo cancellation and echo post-processing. Linear echo cancellation typically uses a filtering algorithm to generate a residual signal, determines an echo suppression aggressiveness factor, and then performs echo post-processing on the residual signal based on the echo suppression aggressiveness factor to produce the final signal.

[0003] Video conferencing generally includes single-speaker and dual-speaker states. The echo suppression aggressiveness factors determined in the single-speaker and dual-speaker states are different. The problem with related technologies is that it is impossible to accurately determine whether the audio frame is in the single-speaker or dual-speaker state, which makes the determined echo suppression aggressiveness factor less accurate and the echo suppression effect poor. Summary of the Invention

[0004] The embodiments of the present application provide an echo suppression method, apparatus, electronic device, and storage medium to address the problem that related technologies cannot accurately determine audio status information, resulting in poor accuracy of the determined echo suppression aggressiveness factor and poor echo suppression effect.

[0005] The present application provides an echo suppression method, the method comprising:

[0006] Acquire a first audio signal collected by an audio acquisition device and a second audio signal output by an audio output device; determine an estimated echo signal corresponding to the second audio signal through linear filtering, and determine a residual signal based on the first audio signal and the estimated echo signal;

[0007] Perform voice activity detection on the first audio signal and the second audio signal, respectively, and determine the similarity between the first audio signal and the estimated echo signal; determine audio state information based on the voice activity detection result and the similarity, determine an echo suppression aggressiveness factor based on the audio state information, and perform echo suppression processing on the residual signal based on the echo suppression aggressiveness factor.

[0008] Further, the determining the similarity between the first audio signal and the estimated echo signal includes:

[0009] A first autopower spectrum of the first audio signal, a second autopower spectrum of the estimated echo signal, and a cross-power spectrum between the first audio signal and the estimated echo signal are respectively determined, and a similarity between the first audio signal and the estimated echo signal is determined based on the first autopower spectrum, the second autopower spectrum, and the cross-power spectrum.

[0010] Furthermore, the determining the similarity between the first audio signal and the estimated echo signal according to the first auto power spectrum, the second auto power spectrum, and the cross power spectrum includes:

[0011] Obtaining and determining the sub-similarity corresponding to each frequency point in a preset frequency band respectively based on the first auto-power spectrum, the second auto-power spectrum, and the cross-power spectrum corresponding to each frequency point;

[0012] The similarity between the first audio signal and the estimated echo signal is determined according to the sub-similarity corresponding to each frequency point.

[0013] Furthermore, the audio status information includes at least one of single-talk status information, double-talk status information and full-talk status information.

[0014] Furthermore, the determining of the audio state information according to the voice activity detection result and the similarity includes:

[0015] If it is determined that the voice activity detection result of the first audio signal is no voice activity detected, and the voice activity detection result of the second audio signal is detected voice activity, and the similarity is less than a preset first similarity threshold, determining that the audio state information is single-speaking state information;

[0016] If it is determined that the voice activity detection results of the first audio signal and the second audio signal are both voice activity detected, and the similarity is greater than the preset first similarity threshold, determining that the audio state information is dual-talk state information;

[0017] If the voice activity detection result of the second audio signal is determined to be no voice activity detected, the audio state information is determined to be all-through state information.

[0018] Furthermore, the process of determining the preset first similarity threshold includes:

[0019] determining a first ratio of a second autopower spectrum of the estimated echo signal to a first autopower spectrum of the first audio signal;

[0020] The preset first similarity threshold is determined according to the first ratio and a preset correction coefficient.

[0021] Furthermore, the audio state information includes at least single-talk state information and dual-talk state information. After determining the audio state information based on the voice activity detection result and the similarity, and before determining the echo suppression aggressiveness factor based on the audio state information, the method further includes:

[0022] Determining a third autopower spectrum of the residual signal, and a second ratio of the third autopower spectrum to the first autopower spectrum;

[0023] If the audio state information is determined to be single-speaking state information based on the voice activity detection result and the similarity, and the second ratio is greater than a preset second similarity threshold, the single-speaking state information is corrected to dual-speaking state information;

[0024] If the audio state information is determined to be dual-talk state information based on the voice activity detection result and the similarity, and the third autopower spectrum is less than a preset third similarity threshold, the dual-talk state information is corrected to single-talk state information.

[0025] Furthermore, the audio state information includes at least the single-talk state information, and determining the echo suppression aggressiveness factor according to the audio state information includes:

[0026] In response to the audio state information being the single-talk state information, obtaining a first echo suppression aggressiveness factor of a previous audio frame;

[0027] The echo suppression aggressiveness factor is determined according to the sum of the first echo suppression aggressiveness factor and a preset first step length factor.

[0028] Furthermore, the audio state information includes at least the double-talk state information or the full-talk state information, and determining the echo suppression aggressiveness factor according to the audio state information includes:

[0029] In response to the audio state information being the double-talk state information or the all-talk state information, obtaining a first echo suppression aggressiveness factor of a previous audio frame;

[0030] The echo suppression aggressiveness factor is determined according to a difference between the first echo suppression aggressiveness factor and a preset second step size factor.

[0031] In another aspect, the present application provides an echo suppression device, comprising:

[0032] a determination module configured to obtain a first audio signal collected by an audio collection device and a second audio signal output by an audio output device; determine an estimated echo signal corresponding to the second audio signal through linear filtering, and determine a residual signal based on the first audio signal and the estimated echo signal;

[0033] An echo suppression module is configured to perform voice activity detection on the first audio signal and the second audio signal, respectively, and determine a similarity between the first audio signal and the estimated echo signal; determine audio state information based on the voice activity detection result and the similarity, determine an echo suppression aggressiveness factor based on the audio state information, and perform echo suppression processing on the residual signal based on the echo suppression aggressiveness factor.

[0034] Furthermore, the echo suppression module is specifically used to respectively determine a first autopower spectrum of the first audio signal, a second autopower spectrum of the estimated echo signal, and a cross-power spectrum between the first audio signal and the estimated echo signal, and determine the similarity between the first audio signal and the estimated echo signal based on the first autopower spectrum, the second autopower spectrum, and the cross-power spectrum.

[0035] Furthermore, the echo suppression module is specifically used to respectively obtain and determine the sub-similarity corresponding to each frequency point in a preset frequency band based on the first auto-power spectrum, the second auto-power spectrum and the cross-power spectrum corresponding to each frequency point; and determine the similarity between the first audio signal and the estimated echo signal based on the sub-similarity corresponding to each frequency point.

[0036] Furthermore, the echo suppression module is specifically configured to, if it is determined that the voice activity detection result of the first audio signal is no voice activity detected, and the voice activity detection result of the second audio signal is detected voice activity, and the similarity is less than a preset first similarity threshold, determine that the audio state information is single-talk state information; if it is determined that the voice activity detection results of the first audio signal and the second audio signal are both voice activity detected, and the similarity is greater than the preset first similarity threshold, determine that the audio state information is dual-talk state information; and if it is determined that the voice activity detection result of the second audio signal is no voice activity detected, determine that the audio state information is all-talk state information.

[0037] Furthermore, the echo suppression module is further configured to determine a first ratio of a second autopower spectrum of the estimated echo signal to a first autopower spectrum of the first audio signal; and determine the preset first similarity threshold based on the first ratio and a preset correction coefficient.

[0038] Furthermore, the device further comprises:

[0039] The correction module is configured to determine a third autopower spectrum of the residual signal and a second ratio of the third autopower spectrum to the first autopower spectrum; if the audio state information is determined to be single-speaking state information based on the voice activity detection result and the similarity, and the second ratio is greater than a preset second similarity threshold, the single-speaking state information is corrected to dual-speaking state information; if the audio state information is determined to be dual-speaking state information based on the voice activity detection result and the similarity, and the third autopower spectrum is less than a preset third similarity threshold, the dual-speaking state information is corrected to single-speaking state information.

[0040] Furthermore, the echo suppression module is specifically configured to obtain a first echo suppression aggressiveness factor of a previous audio frame in response to the audio state information being the single-talk state information; and determine the echo suppression aggressiveness factor based on the sum of the first echo suppression aggressiveness factor and a preset first step length factor.

[0041] Furthermore, the echo suppression module is specifically configured to obtain a first echo suppression aggressiveness factor of a previous audio frame in response to the audio state information being the double-talk state information or the all-talk state information; and determine the echo suppression aggressiveness factor based on a difference between the first echo suppression aggressiveness factor and a preset second step size factor.

[0042] On the other hand, the present application provides an electronic device, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus;

[0043] Memory for storing computer programs;

[0044] The processor is configured to implement any of the above method steps when executing a program stored in the memory.

[0045] On the other hand, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, any of the method steps described above is implemented.

[0046] The present application provides an echo suppression method, apparatus, electronic device, and storage medium. The method includes: obtaining a first audio signal collected by an audio capture device and a second audio signal output by an audio output device; determining an estimated echo signal corresponding to the second audio signal through linear filtering, and determining a residual signal based on the first audio signal and the estimated echo signal; performing voice activity detection on the first audio signal and the second audio signal, respectively, and determining the similarity between the first audio signal and the estimated echo signal; determining audio state information based on the voice activity detection result and the similarity, determining an echo suppression aggressiveness factor based on the audio state information, and performing echo suppression processing on the residual signal based on the echo suppression aggressiveness factor.

[0047] The above technical solution has the following advantages or beneficial effects:

[0048] In this application, a first audio signal captured by an audio capture device and a second audio signal output by an audio output device are obtained, an estimated echo signal corresponding to the second audio signal is determined, and a residual signal is determined based on the first audio signal and the estimated echo signal. Voice activity detection is performed on each of the first and second audio signals, and the similarity between the first and second audio signals is determined. The audio state information is determined by combining the voice activity detection results and the similarity between the first and second audio signals, thereby making the determination of the audio state information more accurate, and thus the determination of the echo suppression aggressiveness factor more accurate, resulting in better echo suppression results. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0050] Figure 1 Schematic diagram of the echo suppression process provided by this application;

[0051] Figure 2 Detailed flow chart of echo suppression provided for this application;

[0052] Figure 3 A schematic diagram of the structure of the echo suppression device provided in this application;

[0053] Figure 4 This is a schematic diagram of the electronic device structure provided in this application. DETAILED DESCRIPTION

[0054] The present application will be further described in detail below with reference to the accompanying drawings. It is apparent that the embodiments described are only a portion of the embodiments of the present application, not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present application without creative effort are intended to fall within the scope of protection of the present application.

[0055] Figure 1 This is a schematic diagram of the echo suppression process provided by this application. The process includes the following steps:

[0056] S101: Acquire a first audio signal collected by an audio acquisition device and a second audio signal output by an audio output device; determine an estimated echo signal corresponding to the second audio signal through linear filtering, and determine a residual signal based on the first audio signal and the estimated echo signal.

[0057] S102: Perform voice activity detection on the first audio signal and the second audio signal respectively, and determine the similarity between the first audio signal and the estimated echo signal.

[0058] S103: Determine audio state information according to the voice activity detection result and the similarity, determine an echo suppression aggressiveness factor according to the audio state information, and perform echo suppression processing on the residual signal according to the echo suppression aggressiveness factor.

[0059] The echo suppression method provided in the present application is applied to an electronic device, which may be a PC, a tablet computer, or a server.

[0060] An electronic device acquires a first audio signal captured by an audio capture device and a second audio signal output by an audio output device. The audio capture device includes a microphone, and the audio output device includes a speaker. In this application, the audio signal captured by the audio capture device is referred to as the first audio signal, and the audio signal output by the audio output device is referred to as the second audio signal. The first audio signal includes an audio signal input by a user into the audio capture device and an audio signal transmitted to the audio capture device after the second audio signal is reflected.

[0061] The first audio signal and the second audio signal are converted from the time domain to the frequency domain. Specifically, a frame overlapping segmentation method is employed, where the overlap between frames is a frame shift, and the frame shift value is half the frame length. The first audio signal and the second audio signal of the current frame, along with the audio signal captured by the audio capture device and the second audio signal output by the audio output device in the previous frame, are converted from time domain signals to frequency domain signals using a short-time Fourier transform.

[0062] Linear filtering is performed on the second audio signal in the frequency domain to obtain an estimated echo signal corresponding to the second audio signal. Linear filtering algorithms include adaptive filters such as LMS, NLMS, RLS, and Kalman. The filter coefficients in the filter are called echo paths. The product of the filter coefficients and the second audio signal in the frequency domain is used as the estimated echo signal corresponding to the second audio signal. The difference between the first audio signal and the estimated echo signal is used as the residual signal.

[0063] For example, the first audio signal captured by the audio acquisition device is d(n), and the second audio signal output by the audio output device is x(n). Using the short-time Fourier transform method, the time domain signal is converted into a frequency domain signal. The frequency domain audio signal corresponding to d(n) is D(k), and the frequency domain audio signal corresponding to x(n) is X(k). The frequency domain form of the filter coefficient is recorded as W(k), and the echo signal corresponding to the second audio signal is recorded as Y(k). Among them, Y(k) = W(k)X(k). The time domain estimated echo signal corresponding to Y(k) is recorded as y(n), and the residual signal is recorded as e(n). Among them, e(n) = d(n)-y(n).

[0064] It should be noted that, in this application, the single-talk and double-talk status information is used to guide whether the filter coefficients are updated, as follows:

[0065]

[0066] Taking the echo cancellation NLMS algorithm as an example, the update method is as follows:

[0067]

[0068] W(k+1)=W(k)+μ(k)X H (k)E(k);

[0069] α and Δ are preset fixed values.

[0070] The electronic device obtains a first audio signal and a second audio signal, determines an estimated echo signal corresponding to the second audio signal, determines a residual signal based on a difference between the first audio signal and the estimated echo signal, and then performs voice activity detection on the first audio signal and the second audio signal, respectively, to obtain a voice activity detection result for the first audio signal and a voice activity detection result for the second audio signal. The voice activity detection result includes whether voice activity is detected or not detected.

[0071] The electronic device determines the similarity between the first audio signal and the estimated echo signal, wherein the similarity between the first audio signal and the estimated echo signal can be determined by using calculation methods such as cosine similarity, cepstral distance, and KL divergence.

[0072] Audio state information is determined based on the voice activity detection result and the similarity, where the audio state information includes at least one of single-talk state information, double-talk state information, and all-talk state information. An echo suppression aggressiveness factor is then determined based on the audio state information, wherein the echo suppression aggressiveness factor corresponding to the single-talk state information is higher than the echo suppression aggressiveness factor corresponding to the double-talk state information, and the echo suppression aggressiveness factor corresponding to the single-talk state information is higher than the echo suppression aggressiveness factor corresponding to the all-talk state information. The echo suppression aggressiveness factor corresponding to the double-talk state information and the echo suppression aggressiveness factor corresponding to the all-talk state information may be the same or different.

[0073] The residual signal is subjected to echo suppression processing according to the echo suppression aggressiveness factor. Specifically, the gain G value is determined by combining the echo suppression aggressiveness factor and the echo post-processing module, and the gain G value is used to perform echo suppression processing on the residual signal to obtain an audio signal after echo cancellation.

[0074] In this application, a first audio signal captured by an audio capture device and a second audio signal output by an audio output device are obtained, an estimated echo signal corresponding to the second audio signal is determined, and a residual signal is determined based on the first audio signal and the estimated echo signal. Voice activity detection is performed on each of the first and second audio signals, and the similarity between the first and second audio signals is determined. The voice activity detection results and the similarity of the first and second audio signals are combined to determine whether the audio state information is single-speaker or dual-speaker, thereby making the determination of the audio state information more accurate, thereby improving the accuracy of the echo suppression aggressiveness factor and the echo suppression effect.

[0075] In the single-talk state, the residual signal contains residual echo and microphone noise. Echo suppression is necessary. Since there is no speaker at the microphone, a stronger level of echo suppression can be used. This also suppresses microphone noise, which is itself an undesirable component.

[0076] In the dual-talk state or the full-pass state, the residual signal contains residual echo, noise at the microphone end, and speech at the microphone end. It is necessary to retain the speech at the microphone end as much as possible to make it clear and coherent.

[0077] In the present application, in order to more accurately determine the similarity between the first audio signal and the echo signal, determining the similarity between the first audio signal and the estimated echo signal includes:

[0078] A first autopower spectrum of the first audio signal, a second autopower spectrum of the estimated echo signal, and a cross-power spectrum between the first audio signal and the estimated echo signal are determined respectively, and a similarity between the first audio signal and the estimated echo signal is determined based on the first autopower spectrum, the second autopower spectrum, and the cross-power spectrum.

[0079] In this application, the first autopower spectrum S of the first audio signal is determined dd , estimate the second autopower spectrum S of the echo signal yy and the cross power spectrum S of the first audio signal and the estimated echo signal yd The specific determination process is as follows:

[0080] S dd =αS dd +(1-α)D(k)D * (k);

[0081] S yy =αS yy +(1-α)Y(k)Y * (k);

[0082] S yd =αS yd +(1-α)Y(k)D * (k).

[0083] Where α is a preset smoothing coefficient, such as 0.75, 0.8, 0.85, etc. * indicates transposition.

[0084] According to the first autopower spectrum S dd , the second autopower spectrum S yy and the cross power spectrum S yd Determine the similarity between the first audio signal and the estimated echo signal. The specific determination process is as follows:

[0085]

[0086] Where * represents transposition, k represents the kth frequency point, and similarity(k) represents the similarity between the first audio signal and the estimated echo signal at the kth frequency point.

[0087] Considering that audio signal energy is generally concentrated within a certain frequency band, in order to further obtain more reliable audio state information, in this application, determining the similarity between the first audio signal and the estimated echo signal based on the first autopower spectrum, the second autopower spectrum, and the cross-power spectrum includes:

[0088] Obtaining and determining the sub-similarity corresponding to each frequency point in a preset frequency band respectively based on the first auto-power spectrum, the second auto-power spectrum, and the cross-power spectrum corresponding to each frequency point;

[0089] The similarity between the first audio signal and the estimated echo signal is determined according to the sub-similarity corresponding to each frequency point.

[0090] For example, each frequency point in the preset frequency band includes frequency points K1 to K2, and the first autopower spectrum, second autopower spectrum and cross-power spectrum corresponding to each frequency point in frequency points K1 to K2 are obtained respectively, and then according to the formula Determine the sub-similarity corresponding to each frequency point. Then, determine the similarity between the first audio signal and the estimated echo signal based on the sub-similarity corresponding to each frequency point. For example, the sum of the sub-similarity corresponding to each frequency point is used as the similarity between the first audio signal and the estimated echo signal.

[0091] It should be noted that the present application may also adopt the following method to determine the similarity between the first audio signal and the estimated echo signal. Specifically, the similarity between the first audio signal and the estimated echo signal is determined as follows: yd , and determine the similarity between the first audio signal and the second audio signal according to the above method xd , and then combined with similarity yd and similarity xd Determine the similarity between the final first audio signal and the estimated echo signal sum The specific expressions are as follows:

[0092]

[0093] α represents the similarity weight value between the first audio signal and the estimated echo signal, and β represents the similarity weight value between the first audio signal and the second audio signal.

[0094] In this application, in order to make the determination of the audio state information more accurate, the determining of the audio state information according to the voice activity detection result and the similarity includes:

[0095] If it is determined that the voice activity detection result of the first audio signal is no voice activity detected, and the voice activity detection result of the second audio signal is detected voice activity, and the similarity is less than a preset first similarity threshold, determining that the audio state information is single-speaking state information;

[0096] If it is determined that the voice activity detection results of the first audio signal and the second audio signal are both voice activity detected, and the similarity is greater than the preset first similarity threshold, determining that the audio state information is dual-talk state information;

[0097] If the voice activity detection result of the second audio signal is determined to be no voice activity detected, the audio state information is determined to be all-through state information.

[0098] In the present application, if it is determined that the voice activity detection result of the first audio signal is that no voice activity is detected, and the voice activity detection result of the second audio signal is that voice activity is detected, and the similarity is less than the preset first similarity threshold, the audio state information is determined to be single-speaking state information. The preset first similarity threshold can be an empirical value, such as 0.8, 0.9, etc. If it is determined that the voice activity detection results of the first audio signal and the second audio signal are both that voice activity is detected, and the similarity is greater than the preset first similarity threshold, the audio state information is determined to be dual-speaking state information. In addition, if it is determined that the voice activity detection result of the second audio signal is that no voice activity is detected, there is no need to judge the size between the similarity and the first similarity threshold, and the audio state information is directly determined to be full-pass state information. The details are as follows:

[0099]

[0100] In the formula, status represents the audio status information, far vad represents the voice activity detection result of the second audio signal, near vad represents the voice activity detection result of the first audio signal, T1 represents the preset first similarity threshold, similarity yd represents the similarity between the first audio signal and the echo signal, pass represents the pass-through status information, far vad=1 indicates that the voice activity detection result of the second audio signal detects voice activity, far vad=0 indicates that the voice activity detection result of the second audio signal does not detect voice activity, near vad=1 indicates that the voice activity detection result of the first audio signal detects voice activity, and near vad=0 indicates that the voice activity detection result of the first audio signal does not detect voice activity.

[0101] In order to make the preset first similarity threshold more accurate, in this application, the process of determining the preset first similarity threshold includes:

[0102] determining a first ratio of a second autopower spectrum of the estimated echo signal to a first autopower spectrum of the first audio signal;

[0103] The preset first similarity threshold is determined according to the first ratio and a preset correction coefficient.

[0104] In this application, a first ratio of the second autopower spectrum of the estimated echo signal to the first autopower spectrum of the first audio signal is first determined. The product of the first ratio and a preset correction coefficient is used as a preset first similarity threshold. Specifically, it is expressed as: v represents the preset correction coefficient, and the value of v is between 0 and 1.

[0105] In this application, considering that in actual usage scenarios, for a smaller first audio signal, to ensure more complete echo signal suppression, it is more inclined to determine it as single-speaking state information; for a larger first audio signal, to ensure audio signal continuity, it is more inclined to determine it as double-speaking state information. To more accurately perform echo suppression, the audio state information includes at least single-speaking state information and double-speaking state information. After determining the audio state information based on the voice activity detection result and the similarity, and before determining the echo suppression aggressiveness factor based on the audio state information, the method further includes:

[0106] Determining a third autopower spectrum of the residual signal, and a second ratio of the third autopower spectrum to the first autopower spectrum;

[0107] If the audio state information is determined to be single-speaking state information based on the voice activity detection result and the similarity, and the second ratio is greater than a preset second similarity threshold, the single-speaking state information is corrected to dual-speaking state information;

[0108] If the audio state information is determined to be dual-talk state information based on the voice activity detection result and the similarity, and the third autopower spectrum is less than a preset third similarity threshold, the dual-talk state information is corrected to single-talk state information.

[0109] In this application, audio state information includes at least single-talk state information and double-talk state information, which includes the determination of single-talk state information and double-talk state information, as well as the determination of other state information in addition to single-talk state information and double-talk state information. In order to more accurately perform echo suppression, the embodiments of this application mainly involve the correction process of the determined single-talk state information and double-talk state information. The details are as follows:

[0110] This application determines the third autopower spectrum of the residual signal, which is expressed as: S ee =αS ee +(1-α)E(k)E * (k); where S ee is the third autopower spectrum, α is the smoothing coefficient, E(k) is the residual signal energy value, and * indicates transposition.

[0111] Determine the second ratio of the third auto power spectrum to the first auto power spectrum as S ee / S dd .

[0112] If the audio state information is determined to be single-speaking state information based on the voice activity detection result and the similarity, and the second ratio is greater than the preset second similarity threshold T2, the single-speaking state information is modified to dual-speaking state information; if the audio state information is determined to be dual-speaking state information based on the voice activity detection result and the similarity, and the third autopower spectrum is less than the preset third similarity threshold T3, the dual-speaking state information is modified to single-speaking state information. Specifically expressed as:

[0113] When status=single talk, if When the status is set to "double talk", it is corrected to: status = double talk;

[0114] When status=Dual talk, if S ee <T3, corrected to: status = single talk.

[0115] Among them, the preset first similarity threshold, the preset second similarity threshold and the preset third similarity threshold can be set according to the actual scene and needs, and the size relationship between the three is not limited. Wherein, γ represents a preset coefficient, which is generally selected as a value between 0 and 1, and △ represents a preset lower limit value of T3, which is generally selected as a smaller value close to 0.

[0116] The echo suppression aggressiveness factor in the post-processing module of echo cancellation is defined as gamma, and the gamma value range is defined as [0, 1]. In single-talk mode, the goal is to minimize echo, so a higher gamma value is used. In dual-talk mode or pass mode, the goal is to preserve the primary audio signal as much as possible, so a lower gamma value is used.

[0117] In a normal conference scenario, there is no strict distinction between single-speaking and double-speaking status information. In order to ensure that the audio signal after echo post-processing suppresses the echo as much as possible while ensuring the continuity of the first audio signal, the gamma value of the current frame can be determined in combination with the single-speaking and double-speaking status information of the current frame and the gamma value of the previous frame. Based on the above considerations, in this application, the audio status information at least includes the single-speaking status information, and the echo suppression aggressiveness factor determined according to the audio status information includes:

[0118] In response to the audio state information being the single-talk state information, obtaining a first echo suppression aggressiveness factor of a previous audio frame;

[0119] The echo suppression aggressiveness factor is determined according to the sum of the first echo suppression aggressiveness factor and a preset first step length factor.

[0120] In this application, audio state information includes at least single-talk state information, which includes the determination of single-talk state information and other state information in addition to single-talk state information. For example, the determination of single-talk state information and double-talk state information; or the determination of single-talk state information and full-talk state information; or the determination of single-talk state information, double-talk state information, and full-talk state information. It should be noted that for an audio frame, its audio state information is one of single-talk state information, double-talk state information, and full-talk state information.

[0121] In the present application, a first echo suppression aggressiveness factor gamma(n-1) of the previous audio frame is first obtained. If the audio state information is single-talk state information, the sum of the first echo suppression aggressiveness factor and the preset first step length factor is determined, and the sum can be directly determined as the echo suppression aggressiveness factor. Alternatively, the first echo suppression aggressiveness factor gamma(n-1) is increased by the preset first step length factor △1 to obtain a second echo suppression aggressiveness factor, and the echo suppression aggressiveness factor is determined based on the second echo suppression aggressiveness factor. The second echo suppression aggressiveness factor can be substituted into the preset functional relationship F1 to obtain the echo suppression aggressiveness factor.

[0122] In the present application, the audio state information includes at least the double-talk state information or the full-talk state information, and determining the echo suppression aggressiveness factor according to the audio state information includes:

[0123] In response to the audio state information being the double-talk state information or the all-talk state information, obtaining a first echo suppression aggressiveness factor of a previous audio frame;

[0124] The echo suppression aggressiveness factor is determined according to a difference between the first echo suppression aggressiveness factor and a preset second step size factor.

[0125] In this application, audio state information includes at least dual-talk state information or full-talk state information, which includes the following situations: determination of dual-talk state information, determination of full-talk state information, determination of single-talk state information and dual-talk state information, determination of full-talk state information and single-talk state information, determination of dual-talk state information and full-talk state information, and determination of single-talk state information, dual-talk state information, and full-talk state information. Similarly, for an audio frame, its audio state information is one of single-talk state information, dual-talk state information, and full-talk state information.

[0126] If the audio state information is dual-talk state information or all-talk state information, the difference between the first echo suppression aggressiveness factor and the preset second step size factor is determined, and the difference can be directly determined as the echo suppression aggressiveness factor. Alternatively, the first echo suppression aggressiveness factor gamma(n-1) is reduced by the preset second step size factor Δ2 to obtain a third echo suppression aggressiveness factor, and the echo suppression aggressiveness factor is determined based on the third echo suppression aggressiveness factor. The third echo suppression aggressiveness factor can be substituted into the preset functional relationship F2 to obtain the echo suppression aggressiveness factor.

[0127] The preset functional relationships F1 and F2 are, for example, linear functions of y=ax, where a is the coefficient of x and a can be 1, 1.5, 2, etc.

[0128] The echo suppression aggressiveness factor gamma is combined with the gain G value obtained by the echo post-processing module. This gain G value can be calculated using a Wiener-based or correlation-based method. The residual signal e(n) is post-processed based on the gain G value to obtain the final echo-cancelled audio signal.

[0129] Figure 2 The detailed flow chart of echo suppression provided by this application is as follows: Figure 2 Shown, including:

[0130] S201: Acquire a first audio signal collected by an audio acquisition device and a second audio signal output by an audio output device; determine an estimated echo signal corresponding to the second audio signal through linear filtering, and determine a residual signal based on the first audio signal and the estimated echo signal.

[0131] S202: Perform voice activity detection on the first audio signal and the second audio signal respectively.

[0132] S203: Obtain and determine the sub-similarity corresponding to each frequency point based on the first auto-power spectrum, the second auto-power spectrum, and the cross-power spectrum corresponding to each frequency point in the preset frequency band; and determine the similarity between the first audio signal and the estimated echo signal based on the sub-similarity corresponding to each frequency point.

[0133] S204: If it is determined that the voice activity detection result of the first audio signal is no voice activity detected, and the voice activity detection result of the second audio signal is detected voice activity, and the similarity is less than a preset first similarity threshold, the audio state information is determined to be single-talk state information; if it is determined that the voice activity detection results of the first audio signal and the second audio signal are both detected voice activity, and the similarity is greater than the preset first similarity threshold, the audio state information is determined to be dual-talk state information; if it is determined that the voice activity detection result of the second audio signal is no voice activity detected, the audio state information is determined to be all-talk state information.

[0134] S205: Determine a third autopower spectrum of the residual signal and a second ratio of the third autopower spectrum to the first autopower spectrum; if the audio state information is determined to be single-speaking state information based on the voice activity detection result and the similarity, and the second ratio is greater than a preset second similarity threshold, correct the single-speaking state information to dual-speaking state information; if the audio state information is determined to be dual-speaking state information based on the voice activity detection result and the similarity, and the third autopower spectrum is less than a preset third similarity threshold, correct the dual-speaking state information to single-speaking state information.

[0135] S206: Obtain a first echo suppression aggressiveness factor of the previous audio frame; in response to the audio state information being single-talk state information, determine the echo suppression aggressiveness factor according to the sum of the first echo suppression aggressiveness factor and a preset first step length factor; in response to the audio state information being double-talk state information or all-talk state information, determine the echo suppression aggressiveness factor according to the difference between the first echo suppression aggressiveness factor and a preset second step length factor.

[0136] S207: Perform echo suppression processing on the residual signal according to the echo suppression aggressiveness factor.

[0137] Figure 3 This is a schematic diagram of the structure of the echo suppression device provided in this application, which includes:

[0138] A determination module 31 is configured to obtain a first audio signal collected by an audio collection device and a second audio signal output by an audio output device; determine an estimated echo signal corresponding to the second audio signal through linear filtering, and determine a residual signal based on the first audio signal and the estimated echo signal;

[0139] The echo suppression module 32 is configured to perform voice activity detection on the first audio signal and the second audio signal, respectively, and determine the similarity between the first audio signal and the estimated echo signal; determine audio state information based on the voice activity detection result and the similarity, determine an echo suppression aggressiveness factor based on the audio state information, and perform echo suppression processing on the residual signal based on the echo suppression aggressiveness factor.

[0140] The echo suppression module 32 is specifically configured to respectively determine a first autopower spectrum of the first audio signal, a second autopower spectrum of the estimated echo signal, and a cross-power spectrum between the first audio signal and the estimated echo signal, and determine a similarity between the first audio signal and the estimated echo signal based on the first autopower spectrum, the second autopower spectrum, and the cross-power spectrum.

[0141] The echo suppression module 32 is specifically configured to respectively obtain and determine the sub-similarity corresponding to each frequency point in a preset frequency band based on the first auto-power spectrum, the second auto-power spectrum, and the cross-power spectrum corresponding to each frequency point; and determine the similarity between the first audio signal and the estimated echo signal based on the sub-similarity corresponding to each frequency point.

[0142] The echo suppression module 32 is specifically configured to: determine, if the voice activity detection result of the first audio signal is determined to be no voice activity detected, and the voice activity detection result of the second audio signal is determined to be detected voice activity, and the similarity is less than a preset first similarity threshold, determine the audio state information as single-talk state information; determine, if the voice activity detection results of the first audio signal and the second audio signal are both determined to be voice activity detected, and the similarity is greater than the preset first similarity threshold, determine the audio state information as dual-talk state information; and determine, if the voice activity detection result of the second audio signal is determined to be no voice activity detected, determine the audio state information as all-talk state information.

[0143] The echo suppression module 32 is further configured to determine a first ratio of the second autopower spectrum of the estimated echo signal to the first autopower spectrum of the first audio signal; and determine the preset first similarity threshold based on the first ratio and a preset correction coefficient.

[0144] The device further comprises:

[0145] The correction module 33 is configured to determine a third autopower spectrum of the residual signal and a second ratio of the third autopower spectrum to the first autopower spectrum; if the audio state information is determined to be single-speaking state information based on the voice activity detection result and the similarity, and the second ratio is greater than a preset second similarity threshold, the single-speaking state information is corrected to dual-speaking state information; if the audio state information is determined to be dual-speaking state information based on the voice activity detection result and the similarity, and the third autopower spectrum is less than a preset third similarity threshold, the dual-speaking state information is corrected to single-speaking state information.

[0146] The echo suppression module 32 is specifically configured to obtain a first echo suppression aggressiveness factor of a previous audio frame in response to the audio state information being the single-talk state information; and determine the echo suppression aggressiveness factor according to the sum of the first echo suppression aggressiveness factor and a preset first step length factor.

[0147] The echo suppression module 32 is specifically configured to obtain a first echo suppression aggressiveness factor of a previous audio frame in response to the audio state information being the double-talk state information or the all-talk state information; and determine the echo suppression aggressiveness factor based on a difference between the first echo suppression aggressiveness factor and a preset second step size factor.

[0148] The present application also provides an electronic device, such as Figure 4 As shown, it includes: a processor 301, a communication interface 302, a memory 303 and a communication bus 304, wherein the processor 301, the communication interface 302, and the memory 303 communicate with each other through the communication bus 304;

[0149] The memory 303 stores a computer program, and when the program is executed by the processor 301 , the processor 301 performs any of the above method steps.

[0150] The communication bus mentioned in the electronic device mentioned above may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, only one thick line is used in the figure, but this does not mean that there is only one bus or only one type of bus.

[0151] The communication interface 302 is used for communication between the electronic device and other devices.

[0152] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk memory. Alternatively, the memory may be at least one storage device located away from the processor.

[0153] The above-mentioned processor can be a general-purpose processor, including a central processing unit, a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit, a field programmable gate array or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component, etc.

[0154] The present application also provides a computer storage readable storage medium, which stores a computer program that can be executed by an electronic device. When the program runs on the electronic device, the electronic device implements any of the above method steps when executing.

[0155] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.

[0156] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.

Claims

1. An echo suppression method, characterized in that: The method comprises: Acquire a first audio signal collected by an audio acquisition device and a second audio signal output by an audio output device; determine an estimated echo signal corresponding to the second audio signal through linear filtering, and determine a residual signal based on the first audio signal and the estimated echo signal; performing voice activity detection on each of the first audio signal and the second audio signal, and determining a similarity between the first audio signal and the estimated echo signal; determining audio state information based on the voice activity detection result and the similarity, determining an echo suppression aggressiveness factor based on the audio state information, and performing echo suppression processing on the residual signal based on the echo suppression aggressiveness factor; Determining the audio state information according to the voice activity detection result and the similarity includes: If it is determined that the voice activity detection result of the first audio signal is no voice activity detected, and the voice activity detection result of the second audio signal is detected voice activity, and the similarity is less than a preset first similarity threshold, determining that the audio state information is single-speaking state information; If it is determined that the voice activity detection results of the first audio signal and the second audio signal are both voice activity detected, and the similarity is greater than the preset first similarity threshold, determining that the audio state information is dual-talk state information; If the voice activity detection result of the second audio signal is determined to be no voice activity detected, the audio state information is determined to be all-through state information.

2. The method according to claim 1, wherein Determining the similarity between the first audio signal and the estimated echo signal includes: A first autopower spectrum of the first audio signal, a second autopower spectrum of the estimated echo signal, and a cross-power spectrum between the first audio signal and the estimated echo signal are respectively determined, and a similarity between the first audio signal and the estimated echo signal is determined based on the first autopower spectrum, the second autopower spectrum, and the cross-power spectrum.

3. The method according to claim 2, wherein The determining the similarity between the first audio signal and the estimated echo signal according to the first auto power spectrum, the second auto power spectrum, and the cross power spectrum includes: Obtaining and determining the sub-similarity corresponding to each frequency point in a preset frequency band respectively based on the first auto-power spectrum, the second auto-power spectrum, and the cross-power spectrum corresponding to each frequency point; The similarity between the first audio signal and the estimated echo signal is determined according to the sub-similarity corresponding to each frequency point.

4. The method according to claim 1, wherein The process of determining the preset first similarity threshold includes: determining a first ratio of a second autopower spectrum of the estimated echo signal to a first autopower spectrum of the first audio signal; The preset first similarity threshold is determined according to the first ratio and a preset correction coefficient.

5. The method according to claim 2, wherein The audio state information includes at least single-talk state information and dual-talk state information. After determining the audio state information according to the voice activity detection result and the similarity, and before determining the echo suppression aggressiveness factor according to the audio state information, the method further includes: Determining a third autopower spectrum of the residual signal, and a second ratio of the third autopower spectrum to the first autopower spectrum; If the audio state information is determined to be single-speaking state information based on the voice activity detection result and the similarity, and the second ratio is greater than a preset second similarity threshold, the single-speaking state information is corrected to dual-speaking state information; If the audio state information is determined to be dual-talk state information based on the voice activity detection result and the similarity, and the third autopower spectrum is less than a preset third similarity threshold, the dual-talk state information is corrected to single-talk state information.

6. The method according to claim 1, wherein The audio state information includes at least the single-talk state information, and determining the echo suppression aggressiveness factor according to the audio state information includes: In response to the audio state information being the single-talk state information, obtaining a first echo suppression aggressiveness factor of a previous audio frame; The echo suppression aggressiveness factor is determined according to the sum of the first echo suppression aggressiveness factor and a preset first step length factor.

7. The method according to claim 1, wherein The audio state information includes at least the double-talk state information or the full-talk state information, and determining the echo suppression aggressiveness factor according to the audio state information includes: In response to the audio state information being the double-talk state information or the all-talk state information, obtaining a first echo suppression aggressiveness factor of a previous audio frame; The echo suppression aggressiveness factor is determined according to a difference between the first echo suppression aggressiveness factor and a preset second step size factor.

8. An echo suppression device, characterized in that: The device comprises: a determination module configured to obtain a first audio signal collected by an audio collection device and a second audio signal output by an audio output device; determine an estimated echo signal corresponding to the second audio signal through linear filtering, and determine a residual signal based on the first audio signal and the estimated echo signal; an echo suppression module, configured to perform voice activity detection on the first audio signal and the second audio signal, respectively, and determine a similarity between the first audio signal and the estimated echo signal; determine audio state information based on the voice activity detection result and the similarity, determine an echo suppression aggressiveness factor based on the audio state information, and perform echo suppression processing on the residual signal based on the echo suppression aggressiveness factor; If the voice activity detection result of the first audio signal is determined to be no voice activity detected, and the voice activity detection result of the second audio signal is determined to be detected, and the similarity is less than a preset first similarity threshold, the audio state information is determined to be single-speaking state information; If it is determined that the voice activity detection results of the first audio signal and the second audio signal are both voice activity detected, and the similarity is greater than the preset first similarity threshold, determining that the audio state information is dual-talk state information; If the voice activity detection result of the second audio signal is determined to be no voice activity detected, the audio state information is determined to be all-through state information.

9. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus; Memory for storing computer programs; A processor, configured to implement the method steps described in any one of claims 1 to 7 when executing a program stored in a memory.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method steps according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Method and device for eliminating echo

    CN102065190A

  • AU1907799A