A method and system for suppressing residual echo based on special acoustic structure

Through the echo residue suppression method based on a special audio structure, the beam space scanning and the echo signal residual probability function are used to solve the problem of distinguishing echo residue and proximal speech signals in the prior art, and more effective echo suppression and speech retention effects are achieved.

CN115527549BActive Publication Date: 2025-05-02BEIJING SOUND PLUS TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211220177.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-30
Publication Date
2025-05-02
Estimated Expiration
2042-09-30

AI Technical Summary

Technical Problem

The existing echo cancellation method based on LMS algorithm has noise residual problem in duplex situations, making it difficult to effectively distinguish echo residual from near-end voice signals, affecting the performance of subsequent processing.

Method used

A method of echo residual suppression based on a special audio structure is proposed. The speaker audio components in multi-channel frequency domain signals are removed through optimized echo cancellation technology, beam space scanning technology is used to count the beam output energy, construct an echo signal residual probability function based on spatial characteristics, and the echo residual suppression algorithm is optimized to distinguish echo residual from near-end speech signals.

Benefits of technology

Effectively distinguish and suppress echo residues, especially when the proximal voice is weak, weak voice signals can be detected, improving the echo residue suppression and automatic gain control performance of subsequent processing, ensuring effective precedence voice when there is proximal voice.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115527549B_ABST
    Figure CN115527549B_ABST
Patent Text Reader

Abstract

The present application provides a method for suppressing residual echo based on a special structure of an acoustic system, including: using a microphone array to obtain a multi-channel speech time domain signal and converting it into a multi-channel frequency domain signal; performing optimized echo cancellation on the multi-channel frequency domain signal to obtain multi-channel data with residual echo signals, and performing beam space scanning; constructing an echo signal residual probability function based on spatial characteristics according to the positioning result of the beam space scanning; using an echo residual suppression algorithm optimized based on the echo signal residual probability function, performing echo residual suppression on the multi-channel data with residual echo signals to obtain a target signal. The present application can effectively distinguish residual echo from near-end speech signals, and effectively detect weak speech signals when the near-end speech is weak, thereby improving the performance of subsequent single-channel speech enhancement and automatic gain control processing, and realizing the function of further suppressing echoes when large echo residues are present, and effectively retaining near-end speech when there is near-end speech.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of voice data processing, and in particular to a method and system for suppressing residual echo based on a special acoustic structure. Background Art

[0002] With the widespread use of voice interaction devices such as smart speakers and conference systems, these devices continue to pursue portability and miniaturization while ensuring high volume. Therefore, the sound unit will inevitably be close to the microphone. Therefore, when the near-end speaker is far away from the device, the signal-to-echo ratio of the microphone receiving signal is often very low, which brings a series of adverse effects to subsequent voice enhancement technology. In the current commonly used voice enhancement technology solutions, echo cancellation processing is given priority.

[0003] Traditional echo cancellation methods often use the correlation between the reference signal and the received signal to adaptively control the update of the filter, but this method has performance limitations, especially in the duplex case, relying solely on the correlation characteristics for control cannot achieve satisfactory results. In recent years, domestic and foreign researchers have also proposed some echo suppression methods based on machine learning. These methods require online or offline supervised learning of residual echo characteristics to ultimately achieve echo suppression, but such methods cannot be applied to all types of residual noise, and will destroy the phase difference characteristics between the signals of each channel, which is not conducive to subsequent wave direction estimation and beamforming processing. Therefore, the commonly used method in actual products is still to use the traditional echo cancellation method based on the minimum mean square error algorithm (LMS, Least Mean Mquare) for processing.

[0004] In practical applications, in simplex mode, only the speaker or near-end speech signal exists, so the existing adaptive echo cancellation method based on the LMS algorithm can more accurately control the update of the filter coefficients, thereby obtaining a better voice interaction effect. In duplex mode, in order not to eliminate the near-end speech, the filter update rate is often reduced, resulting in residual noise. The ratio of the echo residual and the near-end speech component in the output signal after filtering is usually impossible to estimate, which brings many uncertainties to subsequent post-filtering and automatic gain control and other subsequent processing methods.

[0005] Therefore, the existing echo cancellation method based on the LMS algorithm has certain limitations in practical applications. Summary of the invention

[0006] In order to solve the above problems, the present application proposes a residual echo suppression method based on a special acoustic structure, which can effectively distinguish between residual echo and near-end speech signals. In particular, when the near-end speech is weak, the method can effectively detect the weak speech signal, thereby improving the subsequent processing of residual echo suppression and automatic gain control performance, and further suppressing the echo when a large echo remains. At the same time, the function of effectively retaining the near-end speech when there is near-end speech has important application value.

[0007] The present application provides a method for suppressing residual echo based on a special structure of an acoustic system, comprising:

[0008] Removing the loudspeaker audio component in the multi-channel frequency domain signal according to the optimized echo cancellation technology to obtain multi-channel data with residual echo signal; the echo signal is the loudspeaker audio component in the multi-channel frequency domain signal;

[0009] Counting the beam output energy of the multi-channel data with residual echo signals in various directions according to the beam space scanning technology; the various directions are located in the plane formed by the array elements of the microphone array;

[0010] According to the periodic intensity variation characteristics of the peak value curve of the beam output energy of the echo signal in each direction with the angle, a residual probability function of the echo signal based on the spatial characteristics is constructed; the spatial characteristics include that the peak value curve of the beam output energy of the direct sound energy component of the echo signal in each direction with the angle presents periodicity due to the same acoustic path difference between the loudspeaker and each array element of the microphone array;

[0011] According to the echo residual suppression algorithm optimized based on the echo signal residual probability function as an auxiliary parameter, echo residual suppression is performed on the multi-channel data with echo signal residual to obtain a target signal.

[0012] In a possible implementation, the optimization of the optimized echo cancellation algorithm includes optimizing the echo cancellation algorithm of the next frame signal of the multi-channel frequency domain signal using the echo signal residual probability function calculated from the current frame signal of the multi-channel frequency domain signal.

[0013] In a possible implementation, the residual echo suppression algorithm includes an optimized single-channel speech enhancement algorithm.

[0014] The method of performing echo residual suppression on the multi-channel data having echo signal residuals includes:

[0015] According to the optimized single-channel speech enhancement algorithm, an echo residual suppression operation is performed on any single-channel data in the multi-channel data with the echo signal residual to obtain a first processed signal.

[0016] In another possible implementation, the residual echo suppression algorithm includes an optimized single-channel speech enhancement algorithm.

[0017] The method of performing echo residual suppression on the multi-channel data having echo signal residuals includes:

[0018] According to the direction of arrival estimation and beamforming algorithm, target direction estimation and beamforming operations are performed on the multi-channel data with residual echo signals to obtain single-channel data with residual echo signals;

[0019] According to the optimized single-channel speech enhancement algorithm, an echo residual suppression operation is performed on the single-channel data with the echo signal residual to obtain a first processed signal.

[0020] In a possible implementation manner, the residual echo suppression algorithm further includes an optimized automatic gain control algorithm.

[0021] The performing echo residual suppression on the multi-channel data having echo signal residuals also includes:

[0022] According to the optimized automatic gain control algorithm, residual echo suppression is performed on the first processed signal to obtain a target signal.

[0023] The present application provides a residual echo suppression system based on a special acoustic structure, comprising:

[0024] A loudspeaker is located above or below the center of the microphone array whose array elements form a uniform circular array; it is used to transmit the signal of the loudspeaker audio component;

[0025] The microphone array, whose array elements form a uniform circular array, is used to obtain a multi-channel speech signal composed of a near-end speaker's speech component and a loudspeaker's audio component, and convert the multi-channel speech signal into a multi-channel frequency domain signal.

[0026] The microphone array is further used to remove the loudspeaker audio component in the multi-channel frequency domain signal according to the optimized echo cancellation technology to obtain multi-channel data with residual echo signal; the echo signal is the loudspeaker audio component in the multi-channel frequency domain signal;

[0027] Counting the beam output energy of multi-channel data with residual echo signals in various directions according to the beam space scanning technology; the various directions are located in the plane formed by the array elements of the microphone array;

[0028] According to the periodic intensity variation characteristics of the peak value curve of the beam output energy of the echo signal in each direction with the angle, a residual probability function of the echo signal based on the spatial characteristics is constructed; the spatial characteristics include that the peak value curve of the beam output energy of the direct sound energy component of the echo signal in each direction with the angle presents periodicity due to the same acoustic path difference between the loudspeaker and each array element of the microphone array;

[0029] According to the echo residual suppression algorithm optimized based on the echo signal residual probability function as an auxiliary parameter, echo residual suppression is performed on the multi-channel data with echo signal residual to obtain a target signal.

[0030] In a possible implementation, the optimization of the optimized echo cancellation algorithm includes optimizing the echo cancellation algorithm of the next frame signal of the multi-channel frequency domain signal using the echo signal residual probability function calculated from the current frame signal of the multi-channel frequency domain signal.

[0031] In a possible implementation, the residual echo suppression algorithm includes an optimized single-channel speech enhancement algorithm.

[0032] The method of performing echo residual suppression on the multi-channel data having echo signal residuals includes:

[0033] According to the optimized single-channel speech enhancement algorithm, an echo residual suppression operation is performed on any single-channel data in the multi-channel data with the echo signal residual to obtain a first processed signal.

[0034] In another possible implementation, the residual echo suppression algorithm includes an optimized single-channel speech enhancement algorithm.

[0035] The method of performing echo residual suppression on the multi-channel data having echo signal residuals includes:

[0036] According to the direction of arrival estimation and beamforming algorithm, target direction estimation and beamforming operations are performed on the multi-channel data with residual echo signals to obtain single-channel data with residual echo signals;

[0037] According to the optimized single-channel speech enhancement algorithm, an echo residual suppression operation is performed on the single-channel data with the echo signal residual to obtain a first processed signal.

[0038] In a possible implementation manner, the residual echo suppression algorithm further includes an optimized automatic gain control algorithm.

[0039] The performing echo residual suppression on the multi-channel data having echo signal residuals also includes:

[0040] According to the optimized automatic gain control algorithm, residual echo suppression is performed on the first processed signal to obtain a target signal.

[0041] The present application utilizes the characteristics that the acoustic path differences from the speaker sound unit to each array element of the microphone array are the same, while the acoustic path differences from the external near-end signal to each array element of the microphone array are different, to construct an echo signal residual probability function based on the acoustic path difference, which is used to effectively distinguish the echo residual from the near-end signal. A larger value can be obtained when there is an echo residual and a smaller value can be obtained when the residual echo is small. The judgment function can also be used to assist in improving the echo cancellation performance, and subsequent processing such as subsequent speech enhancement and automatic gain control can be performed, which can further suppress the residual echo and retain the near-end signal. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 It is a schematic diagram of a sound pickup scene and a special sound structure provided in an embodiment of the present application;

[0043] Figure 2 is a flow chart of a method for suppressing residual echo based on a special acoustic structure provided by an embodiment of the present application;

[0044] Figure 3 is a schematic diagram of a testing device provided in an embodiment of the present application;

[0045] Figure 4 It is a spectrogram of a microphone receiving signal after echo cancellation based on a pure echo situation provided in an embodiment of the present application;

[0046] Figure 5 It is a spectrogram after echo cancellation of a microphone receiving signal based on pure near-end speech provided in an embodiment of the present application;

[0047] Figure 6 It is a spectrogram after echo cancellation of a microphone receiving signal based on coexistence of a small echo and near-end speech provided in an embodiment of the present application;

[0048] Figure 7 It is a spectrogram after echo cancellation of a microphone receiving signal based on coexistence of a large echo and near-end speech provided in an embodiment of the present application;

[0049] Figure 8 It is a flow chart of an algorithm for constructing a module of an echo residual judgment function based on a special structure of an acoustic instrument provided in an embodiment of the present application;

[0050] Fig. 9 It is a flowchart of a module algorithm of an echo residual suppression method based on a special structure of an acoustic provided in an embodiment of the present application;

[0051] Fig.10It is a comparison chart of the signal spectra of the traditional subsequent processing and the comprehensive processing based on the residual echo probability function provided by the embodiment of the present application;

[0052] Fig.11 It is a structural diagram of a residual echo suppression system based on a special acoustic structure provided in an embodiment of the present application. DETAILED DESCRIPTION

[0053] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described below in conjunction with the accompanying drawings.

[0054] In the description of the embodiments of the present application, words such as "exemplary", "for example" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary", "for example" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary", "for example" or "for example" is intended to present related concepts in a concrete way.

[0055] In the description of the embodiments of the present application, the term "and / or" is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, B exists alone, and A and B exist at the same time. In addition, unless otherwise specified, the term "multiple" means two or more. For example, multiple systems refers to two or more systems, and multiple screen terminals refers to two or more screen terminals.

[0056] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. The terms "include", "comprises", "has" and their variations all mean "including but not limited to", unless otherwise specifically emphasized.

[0057] The method proposed in this application is applicable to microphone arrays and speaker playback systems with special structures, such as desktop conference systems, smart speakers, etc. Without loss of generality, this application takes a small desktop conference speaker with a uniform circular array microphone as an example, where the speaker unit needs to be located above or below the center of the speaker. The principles of other arrays are similar and will not be described separately in this application.

[0058] Figure 1 This is a schematic diagram of a sound pickup scenario and a special sound structure provided by this application. A typical application scenario and a special sound structure such as Figure 1 As shown,

[0059] Figure 1 The left picture in the figure shows a sound pickup scene. The sound pickup device is a microphone array evenly distributed in a circular shape. The signals received by the microphone array include target signals, echo signals, and environmental noise signals. Specifically, when a near-end speaker is talking in a space, the microphone array can receive both the target signal and echo signal of the near-end speaker and the target signal and echo signal emitted by the speaker.

[0060] Figure 1 The right figure in the figure is a special sound structure. A microphone circular array with M elements evenly distributed is placed in a three-dimensional space. The center of the array coincides with the origin O, and the distance between the array element and the center of the circle is δ1. Without loss of generality, the M elements of the microphone array are numbered mic1, mic2, ..., micM in a counterclockwise direction, where the direction of mic1 is set to 0 degrees, the direction of mic2 is set to 360 / M degrees, and so on. The speaker unit S1 is located below the center of the microphone array. The value of M can be determined arbitrarily, for example, Figure 1 In the example, assuming that M=4, the microphone array includes 4 microphone array elements.

[0061] In one example, with the rise of smart office equipment, especially the popularity of conference systems, multi-channel speech enhancement technologies based on microphone arrays, such as echo cancellation, direction of arrival estimation and beamforming, have been widely used. Among them, echo cancellation technology uses a reference signal to eliminate the speaker audio component in the microphone receiving signal through an adaptive filtering method, thereby improving the performance of subsequent processing, such as direction of arrival, beamforming and other speech enhancement processing, and improving the quality of voice interaction.

[0062] In practical applications, the received signals of each microphone are usually first de-echoed, and the processed signals are subjected to direction of arrival (DOA) estimation and beamforming processing to form a directional beam in the near-end target direction, extract the target signal and suppress the reverberation sound components in other directions. Finally, subsequent processing operations such as single-channel speech enhancement, equalization, and automatic gain control are performed to obtain the final enhanced target speech signal. However, when the echo is strong, the echo cancellation method based on the LMS algorithm often has echo residues. The size of the echo residue is not only related to the speaker volume, but also to various factors such as the transfer function and speaker linearity. When the echo residue is large, it will seriously affect the quality of voice interaction. Therefore, how to effectively judge the size of the echo residue after echo cancellation processing and how to distinguish it from the near-end speech signal have very important application value.

[0063] Based on this, the embodiment of the present application utilizes the characteristics that the acoustic path difference between the speaker sound unit and each array element of the microphone array is the same, while the acoustic path difference between the external near-end signal and each array element of the microphone array is different, and proposes a residual echo suppression method based on a special acoustic structure, which is applicable to voice and audio signals, and can be applied to both real-time voice communication systems and non-real-time voice signal enhancement technologies. It includes the following steps:

[0064] (1) The multi-channel data after echo cancellation is subjected to beam space scanning, and a residual probability function of the echo signal based on spatial characteristics is constructed according to the positioning results. This function is used to effectively distinguish the echo residual from the near-end signal. A larger value can be obtained when there is echo residual, and a smaller value can be obtained when the residual echo is small.

[0065] (2) Based on the probability function, multiple processing modules are optimized, including echo cancellation, single-channel speech enhancement, and automatic gain control.

[0066] Experimental results show that the technical solution proposed in this application can effectively distinguish between echo residues and near-end speech signals, especially in the case of long-distance weak near-end speech, and can effectively detect weak speech signals, thereby improving the performance of subsequent processing echo residue suppression and automatic gain control, and further suppressing echoes when large echo residues exist, while effectively retaining the function of near-end speech when there is near-end speech.

[0067] Figure 2 This is a flow chart of a method for suppressing residual echo based on a special acoustic structure provided by the present application.

[0068] like Figure 2 The method comprises the following steps S201 to S205.

[0069] In step S201, a multi-channel speech signal consisting of a near-end speaker speech component and a loudspeaker audio component is acquired by using a microphone array; the multi-channel speech signal is converted into a multi-channel frequency domain signal; the array elements of the microphone array form a uniform circular array; and the loudspeaker is located above or below the center of the microphone array.

[0070] based on Figure 1 In the special acoustic structure shown in the figure, the microphone array elements in the array are numbered mic1, mic2, ..., micM in a counterclockwise rotation order, where the direction of mic1 is set to 0 degrees, the direction of mic2 is set to 360 / M degrees, and so on. Assume that the signal x received by the i-th microphone array element is i (n) is:

[0071] x i (n) = s i (n)+e i (n)+vi (n) (1)

[0072] Where i = 1, 2, LM, s i (n), e i (n) and v i (n) are the target speech signal, echo signal and ambient noise signal received by the i-th microphone array element. The near-end speaker’s speech is the target speech signal, and the signal output by the loudspeaker is the echo signal after various reflections. The time domain signal x(t) received by the array is [x1(t), x2(t), ..., x M (t)] After short-time Fourier Transform (STFT), we get the lth frame, N FFT The kth spectral component of the point FFT:

[0073]

[0074] Where m = 1, 2, LM, X m (k,l) represents the received signal of the mth microphone array element. The first part It includes the direct sound component and reverberation sound component of the near-end speaker's voice. Represents Ω d (k) The corresponding steering vector, S d (k, l) is the signal energy corresponding to that direction, where d = 0 corresponds to the direct sound component, and the others are reverberation sound components. Similarly, the second part It includes the direct sound component of the loudspeaker audio signal and the reverberation sound component composed of other reflection surfaces. Represents Ω g (k) The corresponding steering vector, E g (k, l) is the echo signal energy corresponding to that direction, where g = 0 corresponds to the direct sound component, and the others are reverberation sound components. Similarly, the third part V(k, l) represents the noise signal received by all elements in the microphone array, including V1(k, l)LV m (k,l)LV M (k,l), where m = 1, 2, LM, V m (k, l) represents the noise signal received by the mth microphone array element. In the application of desktop conference audio, since the speaker is very close to the microphone array, the second part The direct sound component is the main one, accounting for the vast majority of energy components.

[0075] based on Figure 1 The special sound structure shown, assuming that the microphone receives the signal X i(k, l), i = 1, 2, ..., M, after being processed by the classic LMS-based echo cancellation algorithm, the output signal is recorded as Z i (k,l),i=1,2,...,M. Taking the traditional SRP-PHAT method as an example for beam scanning, the specific technical solution is as follows:

[0076] The cross-correlation between the output signal Z1(k, l) after the first microphone echo cancellation and the output signal Z2(k, l) after the second microphone echo cancellation is:

[0077]

[0078] where ω=2πf k , f k is the frequency corresponding to the kth frequency point, τ 12 Represents the delay between two output signals, which is mainly determined by the distance between the sound source and each microphone array element pair. Considering the case of multiple pairs of output signals, all output signals are combined and paired, and the results of the combined pairing calculations are cumulatively added to obtain the output power of the controllable beamformer:

[0079]

[0080] where τ nm represents the delay between the output signal of the nth microphone after echo cancellation and the output signal of the mth microphone after echo cancellation. Considering the SRP-PHAT algorithm, the amplitude influence of each frequency point is removed and only the phase information is retained, let:

[0081]

[0082] From equations (4) and (5), we can get:

[0083]

[0084] In practical applications, it is usually assumed that the sound source (or the projection of the incoming wave direction on the horizontal plane) and the microphone array are located in the same plane, and the near-end sound source signal is θ i If incident, By θ i Determine. Scan the plane composed of the elements of the microphone array 360 degrees, and equation (6) can be rewritten as:

[0085]

[0086] In order to reduce the jitter of the beam scanning output results, Smoothing is performed in the time domain, for example, P s (θ i ,l)=α s Ps (θ i ,l-1)+(1-α s )·P(θ i ,l), where α s is the smoothing factor, and this application uses α s =0.8 as an example for simulation. For the beam output of each frame, the normalization method is usually used for processing:

[0087]

[0088] based on Figure 1 In the special acoustic structure shown in the figure, the acoustic path difference from the speaker generating unit to each microphone array element is the same, and beam scanning is performed in the form of formula (7), and the beam output result is repeated periodically at an interval of (360 / M) degrees. However, the acoustic path difference from the near-end speaker's voice to each microphone array element cannot be the same, so it is difficult to have a periodically repeated beam output phenomenon in practical applications.

[0089] like Figure 3 The test device of this application is shown. The conference audio system includes a uniform circular array with a radius of 6 cm composed of 4 microphone arrays, and the speaker is located directly below the center of the device. Taking the actual test data as an example, Figure 4 (a)~(c) to Figure 7 (a) to (c) respectively give the pure echo case, pure near-end speech, the simultaneous existence of both (small echo) and the simultaneous existence of both (large echo), (a) the spectrogram of the original microphone array received signal, (b) the spectrogram of the single-channel output signal after echo cancellation, and the corresponding data calculated by equations (7) and (8), (c) the normalized beam output signal spectrogram, where the scanning interval angle is 10 degrees and the frequency range is 300 Hz to 4000 Hz.

[0090] From the analysis of the above results, we can find the following conclusions:

[0091] (1) Figure 4 The results (a) to (c) show that when only the loudspeaker is emitting sound, the normalized beam output results show obvious periodicity with angle, that is, a large peak appears every 90 degrees starting from 0 degrees, and a peak also appears every 90 degrees starting from 45 degrees. This is because the symmetry of the microphone placement causes the scanning beam output results between adjacent microphones to be symmetrical about the middle angle of the two microphone array elements.

[0092] (2) Figure 5 The results (a) to (c) show that when there is only a near-end signal, the beam output results show obvious single sound source peak characteristics, and even for several other weaker pseudo peaks, there is no obvious periodic repetition phenomenon.

[0093] (3) When the echo residue and the near-end target exist at the same time (i.e., the "double talk" situation), Figure 6 The results (a) to (c) show that when the echo residue is small, the beam output still shows a strong single sound source peak characteristic. Figure 7 The results (a) to (c) show that when the echo residual is strong, the beam output still shows obvious periodic characteristics. When the echo residual signal energy is close to the near-end energy, the beam output also has certain periodic characteristics, but the peak amplitude and intensity will be interfered by the near-end speech.

[0094] In practical applications, in order to judge the amount of echo residue, the correlation between the output signal and the reference signal after echo cancellation processing is often used, or the convergence of the filter is judged to judge the degree of echo residue. These methods all use the correlation characteristics between different signals for processing, and do not use the spatial characteristics of the speaker and the near-end sound source for judgment.

[0095] In one example, based on Figure 1 The special sound structure shown in the figure uses the fact that the acoustic path difference between the loudspeaker and each microphone array element is basically the same, while the acoustic path difference between the external near-end signal and each microphone array element is relatively different. Figures 4 to 7 In order to improve the performance of the adaptive beamforming algorithm, the present application proposes to construct an echo residual function based on the spatial characteristic judgment result, which obtains a larger value when the echo residual is larger and a smaller value when the echo residual is smaller. It can effectively determine whether there is an echo residual and determine the energy relationship between the echo residual and the near-end speech signal.

[0096] like Figure 8 As shown, it is a flowchart of an echo residual judgment function construction module algorithm based on a special structure of an acoustic provided by an embodiment of the present application, which is mainly divided into the following two steps:

[0097] (1) Use the Steered Response Power with phase Transform (SRP-PHAT) beamforming method to scan the beam and count the beam output energy in each direction;

[0098] (2) According to the beam output results, an echo residual judgment function is constructed. When the echo residual is large, a higher probability function value can be obtained, and when the echo residual is small, a lower probability function value can be obtained. This provides auxiliary parameters for subsequent post-processing, thereby improving the echo residual suppression capability.

[0099] The above two steps of constructing the echo residual judgment function are embodied in step S202 to step S204.

[0100] In step S202, the speaker audio component in the multi-channel frequency domain signal is removed according to the optimized echo cancellation technology to obtain multi-channel data with residual echo signal; the echo signal is the speaker audio component in the multi-channel frequency domain signal.

[0101] The optimization of the optimized echo cancellation algorithm includes optimizing the echo cancellation algorithm of the next frame signal of the multi-channel frequency domain signal by using the echo signal residual probability function calculated by the current frame signal of the multi-channel frequency domain signal.

[0102] In step S203, the beam output energy of the multi-channel data with residual echo signals in each direction is counted according to the beam space scanning technology; each direction is located in the plane formed by the array elements of the microphone array.

[0103] In step S204, according to the periodic intensity variation characteristics of the beam output energy peak curve of the echo signal in each direction with the angle, an echo signal residual probability function based on the spatial characteristics is constructed; the spatial characteristics include that the direct sound energy component of the echo signal in each direction has a periodicity in the beam output energy peak curve with the angle because the loudspeaker has the same acoustic path difference to each array element of the microphone array.

[0104] based on Figure 1 The special acoustic structure shown, assuming that M=4, this application takes the case where the microphone array includes 4 microphone array elements as an example for the following specific description.

[0105] The specific calculation method of the echo signal residual probability function is as follows:

[0106] (1) The beam scanning normalization result P)(θ of the first frame in formula (8) is i ,l), the possible first peak area is recorded as ψ=[ψ1,ψ2,...ψ M ], that is, the angle index corresponding to 0 degrees, 90 degrees, 180 degrees and 270 degrees. The possible second peak area is recorded as φ=[φ1,φ2,...φ M ], that is, the angle index corresponding to 45 degrees, 135 degrees, 225 degrees and 315 degrees;

[0107] (2) For the four peak indexes in the first peak region, calculate the angle of each peak index and the corresponding angles of the two adjacent angles P)(θ i ,l), the average value is calculated as follows:

[0108]

[0109] Similarly, for the four peak indexes in the second peak region, calculate the angle of each peak index and the corresponding angles of the two adjacent angles. The average value of is:

[0110]

[0111] (3) When near-end speech and echo residual exist at the same time, as the near-end signal speech energy increases, the echo residual corresponds to The periodic peak structure will be destroyed and become less and less obvious. Using the above characteristics, the threshold value th0 is set. In this application, th0 = 0.5 is set, and M1 (ψ i ,l),i=1,2,3,4, the number of μ1 and M2(φ i ,l),i=1,2,3,4, the number of μ2 greater than the threshold. When μ1 and μ2 are larger, it means that the characteristics of large echo residues in these areas are more obvious, so a state judgment method is designed:

[0112]

[0113] (4) According to the result obtained in the previous step (3), when Flag(l) = 1, it means that at this time There is a very obvious periodic peak structure, that is, there is a very strong echo residue. When Flag(l) = -1, it means that at this time There is almost no periodic peak structure at the potential angle, and it can be considered that there is almost no echo residual. When Flag(l) = 0, it means that at this time Although there are a few peaks in the potential angle, the periodicity is not strong enough. At this time, it can be considered that there are both residual echo and near-end speech, which requires further judgment and processing.

[0114] (5) When Flag(l) = 0, for M1(ψ i ,l),i=1,2,3,4, take the average of the two smallest values ​​and get U1(l); similarly, for M2(φ i ,l),i=1,2,3,4, take the average of the two smallest values ​​to get U2(l). The larger U1(l) or U2(l) is, the larger the peak value of the potential angle is. The two smallest values ​​are selected to avoid the influence of the near-end signal on the peak judgment of the potential angle as much as possible.

[0115] (6) According to Flag(l), U1(l) and U2(l), design the echo residual judgment function P based on DOA echo (l), P echo The larger the (l) is, the more residual echo there is. The specific calculation formula is:

[0116]

[0117] Using the above steps (1) to (6), the beam scanning normalization result P)(θ i ,l) Calculate the echo residual judgment function P echo (l). Figure 4 (a)~(c) to Figure 7 In the case of data corresponding to (a) to (c), Figure 4 (d)~ Figure 7 (d) shows the residual echo judgment function calculated by formula (12). From the results, it can be seen that the method of the present application can obtain a higher P when there is only echo. echo (l), and a smaller P is obtained when there is only a near-end signal echo (l), when the echo residual and the near-end speech signal exist at the same time, as the echo residual increases, P echo Therefore, the method proposed in this application can effectively judge the severity of the residual echo after echo cancellation processing, and this information will be helpful for the optimization design of various subsequent processing modules.

[0118] In practical applications, in order to ensure the near-end voice quality in duplex conditions, the filter update rate is often reduced, resulting in more noise residue. Although an excessively large update rate has a better echo removal effect, it may cause near-end voice distortion. Therefore, how to reasonably set the filter update step size in practical applications is of great significance. This application proposes an echo cancellation filter update optimization method based on an echo residual judgment function.

[0119] The specific analysis is as follows:

[0120] In the duplex case, in order not to eliminate the near-end speech, the filter update rate is often reduced to cause residual noise. The ratio of the echo residual amount and the near-end speech component in the output signal after filtering is usually impossible to estimate. Therefore, how to reasonably set the filter update step size in practical applications is of great significance. This application proposes an echo cancellation filter update optimization method based on the echo residual judgment function. The specific method is as follows:

[0121] For the mth microphone, the lth frame, the received signal X at the kth frequency point m (k, l), the loudspeaker reference signal is recorded as R(k, l), then the filter output Z m (k,l) and the residual signal E m (k,l) are:

[0122]

[0123] E m (k,l)=Xm (k,l)-Z m (k,l-1) (14)

[0124] Among them, W m (k, l) are the echo cancellation filter coefficients. The classic LMS method uses the normalized least mean square algorithm to update the echo cancellation filter:

[0125] W m (k,l)=W m (k,l-1)+2μ(k,l)E m (k,l)X m (k,l) (15)

[0126] In order to speed up the convergence speed, it is necessary to select a suitable step size μ(k,l). The common method is to solve the step size by minimizing the mean square error. The specific formula is:

[0127] μ(k,l)=1 / |X m (k,l)| 2 (16)

[0128] This application proposes to use a step size control method based on an echo residual judgment function to automatically increase the step size when the echo residual is large, thereby speeding up the filter update and achieving a better echo suppression effect, and automatically reduce the step size when the echo residual is small, thereby slowing down the filter update. The optimized echo cancellation filter is:

[0129]

[0130] Among them, P echo (l-1) is the echo residual judgment function calculated in the previous frame. Since the echo cancellation step is before calculating SRP-PHAT, P echo (l) There is generally no drastic change between frames, so the P of the previous frame can be used echo (l-1) assists in step size judgment, γ is a parameter added to prevent the denominator from being too small, and in this application, γ=0.001.

[0131] In one example, after the echo cancellation process, the output result usually needs to be post-processed to further extract and enhance the near-end speech signal. Common post-processing modules include speech enhancement and automatic gain control. Among them, speech enhancement technology includes traditional single-channel speech enhancement, multi-channel speech enhancement, and speech enhancement technology based on machine learning. In order to solve the above problems, the present application proposes a single-channel noise reduction optimization method based on a residual echo probability function, and an automatic gain control optimization method based on a residual echo probability function.

[0132] The above process of optimizing the single-channel speech enhancement algorithm and the automatic gain control algorithm based on the echo signal residual probability function is specifically embodied in step S205 described below.

[0133] In step S205, according to the echo residual suppression algorithm optimized based on the echo signal residual probability function as an auxiliary parameter, the echo residual suppression is performed on the multi-channel data with the echo signal residual to obtain the target signal.

[0134] In one example, the echo residual suppression algorithm includes an optimized single-channel speech enhancement algorithm.

[0135] Perform echo suppression on multi-channel data with residual echo signals, including:

[0136] According to the optimized single-channel speech enhancement algorithm, an echo residual suppression operation is performed on any single-channel data in the multi-channel data having the echo signal residual to obtain a first processed signal.

[0137] In another example, the echo residual suppression algorithm includes an optimized single-channel speech enhancement algorithm.

[0138] Perform echo suppression on multi-channel data with residual echo signals, including:

[0139] According to the direction of arrival estimation and beamforming algorithm, target direction estimation and beamforming operations are performed on the multi-channel data with residual echo signals to obtain single-channel data with residual echo signals;

[0140] According to the optimized single-channel speech enhancement algorithm, an echo residual suppression operation is performed on the single-channel data with the echo signal residual to obtain a first processed signal.

[0141] In practical applications, each channel after echo cancellation may be subjected to subsequent processing respectively, or beamforming may be performed on multi-channel signals according to DOA information first, and subsequent processing may be performed on the single-channel signal after beamforming.

[0142] In this application, only a channel output signal after echo cancellation is used as an example to perform single-channel speech enhancement optimization processing. The single-channel speech enhancement optimization processing method based on the beamforming output signal is similar and will not be repeated in this application. The specific analysis is as follows:

[0143] Traditional single-channel speech enhancement algorithms, such as the classic spectral subtraction, MinimaControl led Recursive Averaging (MCRA) and Optimal ly-modified log-spectral amplitude estimation (OM-LSA) algorithms, all process in the frequency domain. By calculating the probability of speech at the frequency point, the corresponding gain G(k,l) is calculated, and the final target speech estimation signal is obtained:

[0144]

[0145] In the output signal after echo cancellation processing, since the residual echo is also non-steady in many cases and even has certain speech structure characteristics, it is very easy to be judged as the target speech signal and retained, thus affecting the processing effect.

[0146] In order to solve this problem, this application takes the OM-LSA algorithm as an example and provides an optimization method based on the residual echo probability function. This function can be used to distinguish the characteristics of the residual echo and the near-end speech signal, thereby improving the speech enhancement result. Here is a brief description of the specific steps of the OM-LSA algorithm:

[0147] Assume that H0 indicates that the target speech does not appear, and H1 indicates that the speech appears. Then the signal after echo processing is Z(k,l)=S(k,l)+D(k,l), where S(k,l) is the target speech component, and D(k,l) is the noise residue and background noise. According to the OM-LSA method, the estimated value of the target speech can be expressed as:

[0148]

[0149] Among them, q(k,l)=P(H0(k,l)) represents the probability of the absence of prior speech, which is usually calculated as follows:

[0150] q(k,t)=1-P local (k,l)P global (k,l)P frame (l)(20)

[0151] Among them, P local (k, l) represents the probability of local speech existence determined by the prior signal-to-noise ratio ε(k, l). The solution of ε(k, l) refers to the subsequent formula (26), P local (k,l) is calculated as follows:

[0152]

[0153] Among them, ε minRepresents the upper and lower limits of the prior signal-to-noise ratio. In order to reduce the variance, P local (k,l) generally uses the prior signal-to-noise ratio of two adjacent frequency points for averaging, while P global (k, l) represents the global probability of speech existence determined by the prior signal-to-noise ratio, which is generally averaged by the prior signal-to-noise ratios of multiple adjacent frequency points, such as the average of 20 to 30 adjacent frequency points. frame (l) represents the probability of speech existence in a frame determined by the prior signal-to-noise ratio, which is obtained by averaging the prior signal-to-noise ratios of all frequency points. On the other hand, in formula (19) Represents the probability of conditional speech occurrence. The specific calculation method is as follows:

[0154]

[0155] in, Indicates the gain function when speech is not present. Generally, it can be set to a lower fixed value. H1 (k, t) represents the gain function when speech occurs, which is calculated as:

[0156]

[0157] in, represents the posterior signal-to-noise ratio, is the estimated noise power spectrum, and the specific calculation method is:

[0158]

[0159] in, is the smoothing parameter, p'(k,l)=α p p'(k,l-1)+(1-α p )I(k,l) is the conditional speech occurrence probability.

[0160]

[0161] in, Z min (k,l)=min{Z min (k,l-1),Z s (k,l)}, used to determine the ratio of the current frame frequency to the steady-state noise. s (k,l)=α s Z s (k,l-1)+(1-α s )Z(k,l) is smoothed using the first-order smoothing method. In formula (11) represents the a priori signal-to-noise ratio, which is calculated as:

[0162]

[0163] Substituting the above calculation results into formula (19), we can get the final speech estimation value

[0164] The analysis of the above process shows that the key factors affecting the accuracy of the results include effective noise power spectrum estimation, estimation of the probability of the absence of prior speech, etc. The existing method mainly makes judgments based on the energy characteristics of the spectrum structure, and does not use the spatial characteristics of the echo signal for processing. Here, the residual echo probability function P proposed in step S204 is used to echo (l) is optimized. The function is as described in formula (12), which mainly includes the following steps:

[0165] (1) Using P echo (l) Optimize the conditional speech occurrence probability formula. First, optimize formula (25). echo (l) The frame with smaller echo residual is optimized as follows:

[0166]

[0167] When the echo residue is large, especially when the residual signal has certain speech non-stationary characteristics, the original formula may be misjudged as H1, but after being modified by the present application, it can still be judged as H0.

[0168] (2) There is no probability formula for the prior speech q(k,t) = 1-P local (k,l)P global (k,l)P frame (l) Optimize the probability of speech existence in the frame P frame (l) The calculation formula is optimized as follows:

[0169]

[0170] Among them, k f It is related to the number of smoothed frequency points, here we take k f In this way, a lower probability of frame speech existence can be obtained when the echo residual is large, thereby effectively improving the estimation performance of the residual echo power spectrum.

[0171] (3) Gain function G when speech is not present H0 (k,l) is usually set to a low fixed value G min , this value is generally not too small, otherwise it is easy to cause distortion of the target speech in the case of misjudgment, but it will also lead to a decrease in the suppression performance of the residual echo. Therefore, we propose a gradual lower limit setting method, namely Where G(P echo (l)) is a piecewise function. In the case of a clear large echo residue, G(P echo (l)) can be set to a smaller value G m2 , in the case of clear absence of residual echo G(P echo (l)) value is set to G m1 , in other cases G(P echo (l))With P echo (l) increases and becomes smaller. The specific expression is:

[0172]

[0173] Among them, this application sets G m1 =0.1,σ1=0.2,G m1 =0.025,σ2=0.8.

[0174] By optimizing the OM-LSA algorithm in the above manner, the residual noise can be further suppressed in the case of large residual echo, thereby improving the suppression performance of the residual echo.

[0175] In summary, the above single-channel speech enhancement operation is performed on the multi-channel data with residual echo signal to obtain the estimated value of the target speech as the first processed signal.

[0176] In one example, the echo residual suppression algorithm further includes an optimized automatic gain control algorithm.

[0177] Performing echo residual suppression on multi-channel data with residual echo signals, including:

[0178] According to the optimized automatic gain control algorithm, the first processed signal is subjected to echo residual suppression to obtain the target signal. The specific analysis is as follows:

[0179] After post-filtering, it is often necessary to use automatic gain control to keep the gain of the speaker output signal within a preset reasonable range to avoid sudden changes. However, the traditional automatic gain method determines the gain by judging whether the current frame signal is a speech signal or a noise signal. However, when the output signal contains echo residue, especially when the echo is strong, it is difficult to distinguish between echo residue and near-end weak speech signal by simply using the signal spectrum characteristics to judge whether it is a speech signal, so it is possible to misjudge the echo residue as a weak speech signal and amplify it.

[0180] In order to solve the above problems, the present application proposes an automatic gain control optimization method based on a residual echo probability function, and performs automatic gain control processing on the output signal after echo cancellation and subsequent processing enhancement. The specific steps are as follows:

[0181] (1) The estimated value of the target speech Perform ISTFT processing to obtain the time domain signal For time domain signals Perform frame processing, and take every 512 points as one frame in the case of 16kHz sampling. The time domain signal of each frame is recorded as sfrm, and the average energy averEnergy is calculated;

[0182] (2) Calculate the prior gain frontGain. The specific calculation method is:

[0183] if(frontGain*max(abs(sfrm))>=ref)

[0184] frontGain=ref*frontGain / (frontGain*max(abs(sfrm)));

[0185] else

[0186] frontGain=frontGain;

[0187] end

[0188] Where, ref is a fixed reference value, which is 0.3 here. abs() is an absolute value operation for each receiving point of the time domain signal.

[0189] (3) Calculate the posterior gain postGain. The specific calculation method is:

[0190] if(P echo (l)<0.3)

[0191] postGain=0.975*frontGain;

[0192] else

[0193] if(P frame (l)<0.3)

[0194] postGain = frontGain;

[0195] else

[0196] if(frontGain*averEnergy>0.5)

[0197] postGain=0.975*frontGain;

[0198] else

[0199] postGain=1.025*frontGain;

[0200] end

[0201] end

[0202] end

[0203] Among them, P echo (l) is the current frame noise residual judgment function estimated by formula (12), P frame (l) is the probability of speech existence in the current frame estimated using formula (28).

[0204] (4) Calculate the output signal yfrm = postGain*sfrm;

[0205] (5) Repeat steps (1) to (4) for a new input signal.

[0206] In summary, the automatic gain control processing operation is performed on the first processed signal to obtain each frame signal yfrm, and each frame signal is integrated to obtain y(t) as the target signal.

[0207] Fig. 9 is a flowchart of a module algorithm of an echo residual suppression method based on a special structure of an acoustic provided in an embodiment of the present application. Figure 2 The technical solutions of steps S201 to S205 in the embodiment are modularly expressed, and the specific steps are as follows:

[0208] (1) For a voice interaction system with a special acoustic structure, the received signal is framed and converted into a frequency domain signal using STFT, and the LMS algorithm based on residual echo probability function control proposed in step 202 is used to optimize the echo cancellation processing of the received signal of each microphone;

[0209] (2) Using the SRP-PHAT algorithm to perform beam scanning on the two-dimensional plane, and using formulas (7) to (8) to process, a normalized beam output result is obtained;

[0210] (3) Based on the normalized beam output results, the echo residual judgment function P is calculated using equations (9) to (12): echo (l);

[0211] (4) Using echo residual judgment function P echo (1) Optimizing the single-channel noise reduction method in the subsequent processing module, using formulas (27), (28), and (29) to optimize the probability of conditional speech occurrence, the probability of prior speech non-existence, and the gain function when speech does not occur, respectively;

[0212] (5) Using echo residual judgment function P echo(l) Optimizing the automatic gain control method in the subsequent processing module, processing the signal after single-channel noise reduction in step (4) to obtain the final output signal.

[0213] Fig.10 The conventional subsequent processing and the comprehensive processing signal spectrum comparison diagram based on the residual echo probability function provided by the embodiment of the present application, wherein the conventional subsequent processing adopts OM-LSA and the common automatic gain control method (different from the use of P echo (l) control strategy), and the comprehensive method of the present application includes adopting multiple residual echo probability function-based auxiliary control methods proposed in steps S201 to S205.

[0214] Use Figure 1 The actual desktop conferencing system shown in the figure, which consists of a special structure of a microphone array including 4 array elements, was tested and the signal-to-response ratio was -40dB.

[0215] Fig.10 From top to bottom are: (a) original received signal, (b) loudspeaker reference signal, (c) single-channel output after echo cancellation, (d) output result of traditional post-processing method, (e) output result of the comprehensive method of the present application. It can be seen from the results that in the case of only large echo (0s-2s, 5.8s-6.2s, 9.0s-9.3s), there are more residual noises in the output results of the traditional post-processing method, while the method proposed in the present application can effectively suppress the residual noise in the case of large echo residues, and the noise reduction effect has obvious advantages over the traditional post-processing method. At the same time, when near-end speech and echo residues exist at the same time (2.8s-5.8s, 6.2s-9.0s), there are still obvious echo residues in the low- and medium-frequency parts of the output signal of the traditional post-processing method, while the echo residue of the method proposed in the present application is significantly smaller than that of the traditional method. In summary, the method proposed in the present application can significantly improve the residual echo suppression performance.

[0216] Fig.11 1 is a schematic diagram of a residual echo suppression system 1100 based on a special acoustic structure provided in an embodiment of the present application, including:

[0217] The loudspeaker 1101 is located above or below the center of the microphone array whose array elements form a uniform circular array, and is used to transmit the signal of the loudspeaker audio component.

[0218] The microphone array 1102, whose array elements form a uniform circular array, is used to obtain a multi-channel speech signal composed of a near-end speaker's speech component and a loudspeaker's audio component, and convert the multi-channel speech signal into a multi-channel frequency domain signal.

[0219] The microphone array is also used to remove the loudspeaker audio component in the multi-channel frequency domain signal according to the optimized echo cancellation technology to obtain multi-channel data with residual echo signal; the echo signal is the loudspeaker audio component in the multi-channel frequency domain signal;

[0220] Counting the beam output energy of multi-channel data with residual echo signals in each direction according to the beam space scanning technology; each direction is located in the plane composed of the array elements of the microphone array;

[0221] According to the periodic intensity variation characteristics of the peak curve of the beam output energy of the echo signal in each direction with the angle, the residual probability function of the echo signal based on the spatial characteristics is constructed; the spatial characteristics include that the peak curve of the beam output energy of the direct sound energy component of the echo signal in each direction with the angle presents periodicity due to the same acoustic path difference between the loudspeaker and each array element of the microphone array;

[0222] According to an echo residual suppression algorithm optimized based on an echo signal residual probability function as an auxiliary parameter, echo residual suppression is performed on multi-channel data with echo signal residuals to obtain a target signal.

[0223] In one example, the echo residual suppression algorithm includes an optimized single-channel speech enhancement algorithm.

[0224] Perform echo suppression on multi-channel data with residual echo signals, including:

[0225] According to the optimized single-channel speech enhancement algorithm, an echo residual suppression operation is performed on any single-channel data in the multi-channel data having the echo signal residual to obtain a first processed signal.

[0226] In another example, the echo residual suppression algorithm includes an optimized single-channel speech enhancement algorithm.

[0227] Perform echo suppression on multi-channel data with residual echo signals, including:

[0228] According to the direction of arrival estimation and beamforming algorithm, target direction estimation and beamforming operations are performed on the multi-channel data with residual echo signals to obtain single-channel data with residual echo signals;

[0229] According to the optimized single-channel speech enhancement algorithm, an echo residual suppression operation is performed on the single-channel data with the echo signal residual to obtain a first processed signal.

[0230] In one example, the echo residual suppression algorithm further includes an optimized automatic gain control algorithm.

[0231] Performing echo residual suppression on multi-channel data with residual echo signals, including:

[0232] According to the optimized automatic gain control algorithm, residual echo suppression is performed on the first processed signal to obtain a target signal.

[0233] The present application proposes a method for suppressing echo residues and distinguishing echo residues from near-end signals based on a special structure of an acoustic system. By designing a special positional relationship between a microphone array and a speaker sounding unit, the acoustic path difference between the sounding unit and each microphone array element is the same, while the acoustic path difference between other external sound sources and each microphone array element is different. Utilizing this feature, the present application proposes a method for distinguishing echo residues and near-end signals based on a special structure of an acoustic system. Utilizing the feature that the acoustic path difference between the speaker sounding unit and each array element of the microphone array is the same, while the acoustic path difference between the external near-end signal and each array element of the microphone array is different, a spatial characteristic-based echo signal residual probability function is constructed to effectively distinguish echo residues from near-end signals. A larger value can be obtained when there is an echo residue and a smaller value can be obtained when the residual echo is small. The judgment function can also assist in improving the echo cancellation performance, and subsequent processing such as subsequent speech enhancement and automatic gain control can be performed, so that the residual echo can be further suppressed and the near-end signal can be retained.

[0234] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented by software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website site, computer, server or data center to another website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid state disk (SSD)), etc.

[0235] It is understood that the various numerical numbers involved in the embodiments of the present application are only for the convenience of description and are not used to limit the scope of the embodiments of the present application. It should be understood that in the embodiments of the present application, the size of the sequence number of the above-mentioned processes does not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0236] The specific implementation methods described above further illustrate the purpose, technical solutions and beneficial effects of the present application in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present application. Any modifications, equivalent substitutions, improvements, etc. made on the basis of the technical solutions of the present application should be included in the scope of protection of the present application.

Claims

1. A method for suppressing residual echo based on a special acoustic structure, comprising: A multi-channel speech signal consisting of near-end speaker speech and loudspeaker audio components is acquired using a microphone array; Convert a multi-channel speech signal into a multi-channel frequency domain signal; The array elements of the microphone array form a uniform circular array; the loudspeaker is located above or below the center of the microphone array; Removing the loudspeaker audio component in the multi-channel frequency domain signal according to the optimized echo cancellation technology to obtain multi-channel data with residual echo signal; the echo signal is the loudspeaker audio component in the multi-channel frequency domain signal; Counting the beam output energy of the multi-channel data with residual echo signals in various directions according to the beam space scanning technology; the various directions are located in the plane formed by the array elements of the microphone array; According to the periodic intensity variation characteristics of the peak value curve of the beam output energy of the echo signal in each direction with the angle, a residual probability function of the echo signal based on the spatial characteristics is constructed; the spatial characteristics include that the peak value curve of the beam output energy of the direct sound energy component of the echo signal in each direction with the angle presents periodicity due to the same acoustic path difference between the loudspeaker and each array element of the microphone array; According to the echo residual suppression algorithm optimized based on the echo signal residual probability function as an auxiliary parameter, echo residual suppression is performed on the multi-channel data with echo signal residual to obtain a target signal.

2. The method according to claim 1, characterized in that The optimization of the optimized echo cancellation algorithm includes optimizing the echo cancellation algorithm of the next frame signal of the multi-channel frequency domain signal by using the echo signal residual probability function calculated from the current frame signal of the multi-channel frequency domain signal.

3. The method according to claim 1, characterized in that The residual echo suppression algorithm includes an optimized single-channel speech enhancement algorithm. The method of performing echo residual suppression on the multi-channel data having echo signal residuals includes: According to the optimized single-channel speech enhancement algorithm, an echo residual suppression operation is performed on any single-channel data in the multi-channel data with the echo signal residual to obtain a first processed signal.

4. The method according to claim 1, characterized in that The residual echo suppression algorithm includes an optimized single-channel speech enhancement algorithm. The method of performing echo residual suppression on the multi-channel data having echo signal residuals includes: According to the direction of arrival estimation and beamforming algorithm, target direction estimation and beamforming operations are performed on the multi-channel data with residual echo signals to obtain single-channel data with residual echo signals; According to the optimized single-channel speech enhancement algorithm, an echo residual suppression operation is performed on the single-channel data with the echo signal residual to obtain a first processed signal.

5. The method according to claim 3 or 4, characterized in that: The residual echo suppression algorithm also includes an optimized automatic gain control algorithm. The performing echo residual suppression on the multi-channel data having echo signal residuals also includes: According to the optimized automatic gain control algorithm, residual echo suppression is performed on the first processed signal to obtain a target signal.

6. A residual echo suppression system based on a special acoustic structure, comprising: A loudspeaker located above or below the center of the microphone array whose elements form a uniform circular array; Used to transmit the signal of the audio component of the speaker; A microphone array, whose array elements form a uniform circular array, is used to obtain a multi-channel speech signal composed of a near-end speaker's speech component and a loudspeaker's audio component, and convert the multi-channel speech signal into a multi-channel frequency domain signal; The microphone array is further used to remove the loudspeaker audio component in the multi-channel frequency domain signal according to the optimized echo cancellation technology to obtain multi-channel data with residual echo signal; the echo signal is the loudspeaker audio component in the multi-channel frequency domain signal; Counting the beam output energy of multi-channel data with residual echo signals in various directions according to the beam space scanning technology; the various directions are located in the plane formed by the array elements of the microphone array; According to the periodic intensity variation characteristics of the peak value curve of the beam output energy of the echo signal in each direction with the angle, a residual probability function of the echo signal based on the spatial characteristics is constructed; the spatial characteristics include that the peak value curve of the beam output energy of the direct sound energy component of the echo signal in each direction with the angle presents periodicity due to the same acoustic path difference between the loudspeaker and each array element of the microphone array; According to the echo residual suppression algorithm optimized based on the echo signal residual probability function as an auxiliary parameter, echo residual suppression is performed on the multi-channel data with echo signal residual to obtain a target signal.

7. The system according to claim 6, characterized in that The optimization of the optimized echo cancellation algorithm includes optimizing the echo cancellation algorithm of the next frame signal of the multi-channel frequency domain signal by using the echo signal residual probability function calculated from the current frame signal of the multi-channel frequency domain signal.

8. The system according to claim 6, characterized in that The residual echo suppression algorithm includes an optimized single-channel speech enhancement algorithm. The method of performing echo residual suppression on the multi-channel data having echo signal residuals includes: According to the optimized single-channel speech enhancement algorithm, an echo residual suppression operation is performed on any single-channel data in the multi-channel data with the echo signal residual to obtain a first processed signal.

9. The system according to claim 6, wherein the residual echo suppression algorithm comprises an optimized single-channel speech enhancement algorithm. The method of performing echo residual suppression on the multi-channel data having echo signal residuals includes: According to the direction of arrival estimation and beamforming algorithm, target direction estimation and beamforming operations are performed on the multi-channel data with residual echo signals to obtain single-channel data with residual echo signals; According to the optimized single-channel speech enhancement algorithm, an echo residual suppression operation is performed on the single-channel data with the echo signal residual to obtain a first processed signal.

10. The system according to claim 8 or 9, characterized in that The residual echo suppression algorithm also includes an optimized automatic gain control algorithm. The performing echo residual suppression on the multi-channel data having echo signal residuals also includes: According to the optimized automatic gain control algorithm, residual echo suppression is performed on the first processed signal to obtain a target signal.

Citation Information

Patent Citations

  • Residual echo suppression method based on multi-feature flow structure deep neural network

    CN112037809A

  • Echo suppression method and device

    CN112837697A