Vehicle-mounted Bluetooth call echo cancellation method and device, storage medium and electronic terminal

Through the combination algorithm of adaptive filtering and nonlinear filtering, combined with voice energy detection, the echo cancellation problem in complex acoustic environments in vehicle Bluetooth calls is solved, efficient and stable echo cancellation is achieved, and call quality is improved.

CN120390050APending Publication Date: 2025-07-29AUTOMOBILE RES INST OF TSINGHUA UNIV IN SUZHOU XIANGCHENG
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510513052.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

In vehicle Bluetooth calls, the prior art is difficult to effectively eliminate echoes in complex acoustic environments, especially in high volume and dual speaking states, resulting in a degradation in call quality and voice interference.

Method used

Adaptive filtering and nonlinear filtering combination algorithms are used, combined with voice energy detection, echoes are eliminated in stages, and filter parameters and suppression coefficients are dynamically adjusted through dual-talk detection to ensure the integrity of the target signal.

Benefits of technology

It effectively eliminates echoes in complex acoustic environments, avoids the "voice-eating" phenomenon, improves call clarity and coherence, and ensures a high-quality call experience for remote users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120390050A_ABST
    Figure CN120390050A_ABST
Patent Text Reader

Abstract

The invention discloses a vehicle-mounted Bluetooth call echo cancellation method and device, a storage medium and an electronic terminal, and the method comprises the steps: obtaining a first call voice signal, and carrying out the preprocessing of the first call voice signal, so as to obtain a to-be-detected signal, and the to-be-detected signal comprises a reference signal and a target signal; judging whether the call state of the to-be-detected signal is double-talk based on a double-talk detection method; if the call state is a double-talk state, performing first echo cancellation processing on the to-be-detected signal through a combination of an adaptive filtering algorithm and a nonlinear filtering algorithm to obtain a preliminary residual signal, then detecting the preliminary residual signal through a voice energy detection method, and judging whether residual echo exists in the preliminary residual signal; and if the residual echo exists in the preliminary residual signal, performing second echo cancellation processing on the residual signal to obtain a second call voice signal. According to the method, the limitation of a traditional single filtering method in a complex acoustic environment is overcome, the phenomenon of'sound absorption 'is effectively avoided, and the call definition is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of echo cancellation, and in particular to a method, device, storage medium and electronic terminal for canceling vehicle-mounted Bluetooth call echo. Background Art

[0002] With the continuous advancement of automotive intelligence, in-vehicle communication systems have become a core feature of the smart cockpit. In an in-vehicle Bluetooth hands-free call scenario, the user plays a reference signal through the vehicle's speakers, while the vehicle's microphones simultaneously capture the target signal, such as the driver or passenger's voice. However, due to the complexity of the acoustic environment, such as acoustic coupling between the speakers and microphones and reflections within the vehicle, the reference signal played by the speakers is easily picked up by the microphones again, resulting in an acoustic echo. If this echo is not effectively eliminated, the remote user will hear a delayed feedback of their own voice, seriously affecting call quality and user experience.

[0003] Existing echo cancellation technologies primarily rely on linear adaptive filtering algorithms, which estimate the impulse response of the echo path, generate an echo estimate signal, and then cancel it from the near-end signal. However, in-vehicle speakers are prone to nonlinear distortion at high volume, making it difficult for linear filters to fully cancel the echo, and residual echo can still interfere with calls. When both near-end and far-end users speak simultaneously, existing algorithms may over-suppress the target signal, causing speech interruptions or loss of clarity. Factors such as the interior layout and passenger count can cause the echo path to vary over time, making it difficult for traditional linear algorithms to track these changes, thus affecting cancellation effectiveness. While some improved solutions have attempted to incorporate nonlinear filtering or double-talk detection algorithms, they still suffer from high computational power consumption, complex parameter adjustments, and insufficient residual echo detection accuracy. Therefore, achieving efficient and stable echo cancellation in complex acoustic environments while avoiding loss of voice quality in double-talk situations has become a pressing technical challenge in the field of in-vehicle communications. Summary of the Invention

[0004] The purpose of the present invention is to solve the above problems and provide a method, device, storage medium and electronic terminal for canceling vehicle-mounted Bluetooth call echo.

[0005] The technical solution of the present invention is: a method for canceling echo of a vehicle-mounted Bluetooth call, comprising the following steps:

[0006] Obtain a first call voice signal and preprocess the first call voice signal to obtain a signal to be detected, where the signal to be detected includes a reference signal and a target signal; based on a double-talk detection method, determine whether the call state of the signal to be detected is double-talk; if the call state is double-talk, perform a first echo cancellation process on the signal to be detected through a combination of an adaptive filtering algorithm and a non-linear filtering algorithm to obtain a preliminary residual signal, and then, detect the preliminary residual signal through a voice energy detection method and determine whether there is residual echo in the preliminary residual signal; if there is residual echo in the preliminary residual signal, perform a second echo cancellation process on the residual signal to obtain a second call voice signal.

[0007] As an improvement of an embodiment of the present invention, the "obtain a first call voice signal and preprocess the first call voice signal to obtain a signal to be detected" specifically includes: performing frame segmentation on both the reference signal and the target signal, and performing block processing on the segmented reference signal and the segmented target signal according to a preset length to obtain the signal to be detected.

[0008] As an improvement of an embodiment of the present invention, the "based on a double-talk detection method, determine whether the call state of the signal to be detected is double-talk" specifically includes: calculating the cross-correlation value between the target signal and the reference signal, the first autocorrelation value of the target signal, and the second autocorrelation value of the reference signal; generating a double-talk detection decision parameter according to the cross-correlation value and / or the first autocorrelation value and / or the second autocorrelation value of the reference signal; comparing the double-talk detection decision parameter with a preset threshold, and determining whether the current call state is double-talk according to the comparison result.

[0009] As an improvement of an embodiment of the present invention, the "perform a first echo cancellation process on the signal to be detected through a combination of an adaptive filtering algorithm and a non-linear filtering algorithm to obtain a preliminary residual signal" specifically includes: performing a Fourier transform on the signal to be detected and processing it through an adaptive filtering algorithm to obtain a first residual signal; processing the first residual signal through non-linear filtering to obtain a preliminary residual signal.

[0010] As an improvement of an embodiment of the present invention, the "perform a Fourier transform on the signal to be detected and process it through an adaptive filtering algorithm to obtain a first residual signal" specifically includes: the adaptive filtering algorithm is a normalized least mean square adaptive filtering algorithm; performing a Fourier transform on the signal to be detected; performing a zero-padding operation on the filter weight coefficient of the adaptive filtering algorithm, and then, performing block filtering on the reference signal to obtain a filtering result; performing an inverse Fourier transform on the filtering result to obtain the first residual signal; the block filtering process is: where x(n) is the reference signal, d(n) is the reference signal, ω T is the filter weight coefficient at time period T; error(n) is the first residual signal.

[0011] As an improvement of an embodiment of the present invention, the "processing the first residual signal through nonlinear filtering to obtain a preliminary residual signal" specifically includes: calculating a coherence coefficient based on the power spectral densities of the first residual signal, the reference signal, and the target signal, and estimating the echo residual state; determining sub-band nonlinear gain parameters according to a preset suppression coefficient and the signal coherence coefficient; performing sub-band processing on the first residual signal in the frequency domain, and performing gain suppression on the signals of each band according to the nonlinear gain parameters; superimposing the suppressed sub-band signals to generate a second residual signal in the frequency domain; performing an inverse Fourier transform on the second residual signal in the frequency domain to obtain a preliminary residual signal.

[0012] As an improvement of an embodiment of the present invention, the "detecting the preliminary residual signal through a voice energy detection method and determining whether there is residual echo in the preliminary residual signal, and if there is residual echo in the preliminary residual signal, performing echo cancellation on the residual signal to obtain a second call voice signal" specifically includes: setting preset parameters, the preset parameters including a preset window length L, a preset power threshold P th , a preset effective speech percentage R th and a preset suppression coefficient α; calculating the power and point-by-point power of the preliminary residual signal; calculating the effective speech percentage by comparing the point-by-point power with the preset power threshold; if the effective speech percentage is less than the preset percentage, multiplying the preliminary residual signal by the preset suppression coefficient and adding comfort noise to obtain the second call voice signal.

[0013] To achieve one of the above-mentioned invention purposes, an embodiment of the present invention provides a vehicle-mounted Bluetooth call echo cancellation device, including the following modules: an information acquisition module, configured to acquire a first call voice signal and preprocess the first call voice signal to obtain a signal to be detected, the signal to be detected including a reference signal and a target signal; a state judgment module, configured to judge whether the call state of the signal to be detected is a double talk based on a double talk detection method; an echo cancellation module, if the call state is a double talk, performing a first echo cancellation process on the signal to be detected through a combination of an adaptive filtering algorithm and a nonlinear filtering algorithm to obtain a preliminary residual signal, and then, detecting the preliminary residual signal through a voice energy detection method and judging whether there is residual echo in the preliminary residual signal; if there is residual echo in the preliminary residual signal, performing a second echo cancellation process on the residual signal to obtain a second call voice signal.

[0014] To achieve one of the above-mentioned invention purposes, an embodiment of the present invention provides a storage medium storing program instructions, and when the program instructions are executed, the echo cancellation method described in any one of the above is implemented.

[0015] To achieve one of the above-mentioned invention purposes, an embodiment of the present invention provides an electronic terminal, including a processor and a memory, where the memory stores program instructions, and the processor runs the program instructions to implement the echo cancellation method described in any one of the above.

[0016] The vehicle-mounted Bluetooth call echo cancellation method, device, storage medium and electronic terminal provided by the embodiments of the present invention have the following advantages: Through the collaborative processing of linear adaptive filtering and non-linear filtering, the linear and non-linear echo residues are eliminated in stages. First, the main echo components are preliminarily suppressed by linear filtering, and then the residual echo is suppressed in frequency bands in the frequency domain by non-linear filtering, overcoming the limitations of traditional single filtering methods in complex acoustic environments. Based on the improved cross-correlation double-talk detection algorithm, the near-end and far-end call states are accurately judged. In the double-talk state, the suppression coefficient and filter update strategy are dynamically adjusted, the filter coefficient update is paused, and the non-linear gain parameter is optimized to effectively avoid the "sound eating" phenomenon and ensure the integrity of the target signal and the coherence of the call. At the same time, an energy detection-based residual echo cancellation mechanism is introduced, and comfort noise is superimposed to make up for the distortion of the speech signal or the "silent hole", and the progressive processing flow ensures that the residual echo of the far-end received speech signal is lower than the perception threshold, improving the call clarity. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is a flowchart of the vehicle-mounted Bluetooth call echo cancellation method of the present invention;

[0018] Figure 2 is a structural diagram of the vehicle-mounted Bluetooth call echo cancellation device of the present invention;

[0019] Figure 3 is a schematic diagram of the application scenario of the vehicle-mounted Bluetooth call echo cancellation method of the present invention;

[0020] Figure 4 is a comparison diagram of the effects of the vehicle-mounted Bluetooth call echo cancellation method of the present invention and the prior art;

[0021] Figure 5 is a structural diagram of the electronic terminal of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] The present invention will be described in detail below in conjunction with the specific embodiments shown in the drawings. However, these embodiments do not limit the present invention, and any structural, method, or functional transformation made by those of ordinary skill in the art based on these embodiments is included in the protection scope of the present invention.

[0023] If the present invention involves orientations (such as up, down, left, right, front, back, outside, inside, etc.) when being described, the involved orientations need to be defined.

[0024] The scope of the embodiments herein includes the entire scope of the claims and all available equivalents of the claims. Herein, the terms "first", "second", etc. are only used to distinguish one element from another, and do not require or imply any actual relationship or order between these elements. In fact, the first element can also be called the second element, and vice versa. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a structure, device or equipment including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such structure, device or equipment. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the structure, device or equipment including the said element. The various embodiments herein are described in a progressive manner, with each embodiment highlighting the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other.

[0025] The orientations or positional relationships indicated by the terms "longitudinal", "transverse", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", etc. herein are based on the orientations or positional relationships shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be construed as a limitation to the present invention. In the description herein, unless otherwise specified and defined, the terms "mounted", "connected", "coupled" shall be understood in a broad sense. For example, it may be a mechanical connection or an electrical connection, or may also be the communication inside two elements. It may be directly connected or indirectly connected through an intermediate medium. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.

[0026] Embodiment 1 of the present invention provides a method for eliminating in-vehicle Bluetooth call echo, as Figure 1 shown, which includes the following steps:

[0027] Step 101: Obtain a first call voice signal and preprocess the first call voice signal to obtain a signal to be detected, where the signal to be detected includes a reference signal and a target signal;

[0028] As Figure 2As shown, in the in-vehicle Bluetooth hands-free call scenario, first, the remote microphone 302 of the remote user device 301 collects the remote voice, processes it into a remote voice signal, and then sends it to the proximal user device 201. The proximal user device 201 amplifies the power of the remote voice signal and plays it through the proximal speaker 202. At the same time, the proximal microphone 203 collects the in-vehicle proximal voice signal and the remote voice signal played by the proximal speaker 202 to form a first call voice signal. Then, the first call voice signal is subjected to frame and block processing, and finally, a signal to be detected containing a reference signal block and a target signal block is obtained, where the reference signal is the remote voice signal received and amplified proximally, and the target signal is the sum of the proximal voice signal and the echo signal.

[0029] Step 102: Based on the two-talk detection method, determine whether the call state of the signal to be detected is two-talk.

[0030] In practice, first, calculate the target signal and the reference signal in the signal to be detected obtained through the processing of Step 101 to respectively obtain the cross-correlation value between the target signal and the reference signal, the first autocorrelation value of the target signal, and the second autocorrelation value of the reference signal. The cross-correlation value reflects the similarity and correlation between the target signal and the reference signal, and the autocorrelation value reflects the internal correlation of each signal. Then, generate a two-talk detection decision parameter based on these calculated values. According to the actual situation, one can choose to use only the cross-correlation value, only the autocorrelation value, or combine them to generate this decision parameter. Finally, compare the generated two-talk detection decision parameter with a pre-set threshold and determine whether the call state is two-talk. Two-talk detection can greatly improve the echo cancellation effect. In the two-talk state, adjust the echo cancellation algorithm parameters or adopt special strategies to avoid over-suppressing the proximal useful voice signal, make the echo cancellation more accurate, improve the call quality, promptly detect the two-talk state and take countermeasures to reduce problems such as voice distortion and stuttering, and make the call smoother and more comfortable.

[0031] Step 103: If the call state is two-talk, perform a first echo cancellation process on the signal to be detected through a combination of an adaptive filtering algorithm and a non-linear filtering algorithm to obtain a preliminary residual signal. Then, detect the preliminary residual signal through a voice energy detection method and determine whether there is residual echo in the preliminary residual signal. If there is residual echo in the preliminary residual signal, perform a second echo cancellation process on the residual signal to obtain a second call voice signal.

[0032] Here, the adaptive filtering algorithm is a type of filtering algorithm that can automatically adjust its own parameters according to the statistical characteristics changes of the input signal. It can effectively track the changes of the signal in an unknown environment, process the signal in the best way, dynamically adapt to the changes of the echo path in echo cancellation, and better estimate and eliminate the echo. The non-linear filtering algorithm is a filtering method for processing signals with non-linear characteristics. Different from traditional linear filtering, it can process complex signals that do not conform to linear laws, and can handle some non-linear echo components in echo cancellation to improve the echo cancellation effect. The speech energy detection method is to judge the presence, characteristics, etc. of the speech signal by calculating the energy of the speech signal, and can be used to detect whether there is residual echo in the signal in echo cancellation. In practice, when it is judged in step 102 that the call state is two-way talking, the adaptive filtering algorithm and the non-linear filtering algorithm are used to perform the first echo cancellation processing on the signal to be detected obtained in step 101. The adaptive filtering algorithm adjusts the parameters according to the dynamic changes of the echo path, makes a preliminary estimate and elimination of the echo, and the non-linear filtering algorithm processes the non-linear echo components therein. After the two are combined, a preliminary residual signal is obtained. Then, the speech energy detection method is used to detect the preliminary residual signal, calculate the energy of the preliminary residual signal and compare it with a preset energy threshold to judge whether there is residual echo in the preliminary residual signal. If it is detected that there is residual echo in the preliminary residual signal, the preliminary residual signal is subjected to a second echo cancellation processing to further eliminate the residual echo, and finally a second call speech signal is obtained. It can be understood that in this way, the adaptive filtering can be used to dynamically adapt to the echo path, the non-linear filtering can process complex non-linear echo components, and the advantages of the two are complementary, greatly improving the accuracy and comprehensiveness of echo cancellation, and effectively coping with the complex echo situation during two-way talking. Using the speech energy detection to judge whether there is residual echo in the preliminary residual signal ensures targeted elimination of the residual echo, thus significantly improving the call speech quality and minimizing the impact of the echo on the call clarity and fluency, bringing a better call experience to users.

[0033] In this embodiment, the "obtaining the first call speech signal and preprocessing the first call speech signal to obtain the signal to be detected" specifically includes: performing frame division processing on both the reference signal and the target signal, and performing block division processing on the frame-divided reference signal and the frame-divided target signal according to a preset length to obtain the signal to be detected.

[0034] Here, in the frame segmentation process, a fixed frame length and frame shift are set. The frame length is generally between 10 and 30 milliseconds, and 20 milliseconds can be selected. The frame shift is usually less than the frame length. For example, it is set to 10 milliseconds, indicating that there is an overlap of 10 milliseconds between two adjacent frames. According to such rules, the reference signal and the target signal are intercepted in sequence, and the continuous signal is divided into independent frames. Then, block processing is performed. A block length is preset in advance, such as 10 frames. According to the preset length, the frames in the segmented reference signal and target signal are combined into blocks in sequence. If the remaining number of frames is less than the block length, zero-padding or discarding processing can be performed according to the specific situation, and the signal to be detected composed of the reference signal block and the target signal block is obtained. It can be understood that the frame segmentation and block processing of the reference signal and the target signal provide a more effective data basis for subsequent processing steps such as double-talk detection and echo cancellation, thereby improving the performance of the entire echo cancellation system and the call quality.

[0035] In this embodiment, "judging whether the call state of the signal to be detected is double-talk based on the double-talk detection method" specifically includes: calculating the cross-correlation value between the target signal and the reference signal, the first autocorrelation value of the target signal, and the second autocorrelation value of the reference signal; generating a double-talk detection decision parameter according to the cross-correlation value and / or the first autocorrelation value and / or the second autocorrelation value of the reference signal; comparing the double-talk detection decision parameter with a preset threshold, and judging whether the current call state is double-talk according to the comparison result.

[0036] In practice, the double-talk detection method can be the Geigel algorithm, and the expression of the first decision parameter of the Geigel algorithm is: If p1(n)>δ1, then the call state is double-talk. The double-talk detection method can be the cross-correlation algorithm, and the expression of the second decision parameter of the cross-correlation algorithm is:

[0037] If p2(n)>δ2, then the call state is double-talk. The double-talk detection method can be the improved cross-correlation DTD algorithm, and the expression of the third decision parameter of the improved cross-correlation DTD algorithm is: If p3(n)<δ3, then the call state is double-talk. xcorr(d(n),x(n)) = E[d 2 (n)].E[x 2 (n)], where d(n) is the target signal, x(n) is the reference signal, δ1, δ2, and δ3 are all preset thresholds, Rdd is the first autocorrelation value, Rxx is the second autocorrelation value, and xcorr() is the cross-correlation function. The target signal d(n)=h T x(n)+v(n), where h Tis the echo path for time period T, and v(n) is the near-end speech signal.

[0038] In this embodiment, the step of "processing the signal to be detected through a combination of an adaptive filtering algorithm and a non-linear filtering algorithm to obtain a preliminary residual signal" specifically includes: performing a Fourier transform on the signal to be detected and processing it through an adaptive filtering algorithm to obtain a first residual signal; and processing the first residual signal through non-linear filtering to obtain a preliminary residual signal.

[0039] Here, the Fourier transform can convert a signal from the time domain to the frequency domain. In the time domain, the signal is represented as a function that varies with time, while in the frequency domain, the signal is represented as a combination of different frequency components. Through the Fourier transform, various frequency components contained in the signal and their corresponding amplitude and phase information can be analyzed. Specifically, to perform a fast Fourier transform (FFT) on the signal to be detected, it is first necessary to ensure that the signal length is an integer power of 2. If this is not satisfied, it can be achieved by padding with zeros. Reverse the binary representation of the sequence elements of the signal to be detected to obtain a new sequence number and rearrange them accordingly. Then, perform multi-level butterfly operations. Each level contains multiple butterfly units, and each butterfly unit processes two input samples to generate two output samples. According to the butterfly operation formula where is the rotation factor. Starting from the first level, gradually complete log2N levels of butterfly operations to finally obtain the frequency domain representation of the signal to be detected. It can be understood that applying the Fourier transform to the collected signal to be detected and converting it from the time domain to the frequency domain can clearly show the distribution of each frequency component in the signal, facilitating subsequent filtering processing.

[0040] In this embodiment, the step of "performing a Fourier transform on the signal to be detected and processing it through an adaptive filtering algorithm to obtain a first residual signal" specifically includes: the adaptive filtering algorithm is the normalized least mean square adaptive filtering algorithm; performing a Fourier transform on the signal to be detected; performing a zero-padding operation on the filter weight coefficients of the adaptive filtering algorithm, and then performing block filtering on the reference signal to obtain a filtering result; performing an inverse Fourier transform on the filtering result to obtain the first residual signal; the block filtering is: where x(n) is the reference signal, d(n) is the target signal, ω T is the filter weight coefficient at time period T; error(n) is the first residual signal.

[0041] Here, the filter weight coefficient ω of the adaptive filtering algorithm is initialized. To meet the requirements of block filtering processing, zero-padding operation needs to be performed on the filter weight coefficient. Let the length of the reference signal block be L, and the original length of the filter weight coefficient be Q. Then the length of the filter weight coefficient ω after zero-padding is L + Q - 1. The update of the filter weight coefficient satisfies the expression: It can be understood that the zero-padding operation can avoid aliasing phenomena during the block filtering process and ensure the accuracy of the filtering result.

[0042] In this embodiment, the step of "processing the first residual signal through non-linear filtering to obtain a preliminary residual signal" specifically includes: calculating a coherence coefficient based on the power spectral densities of the first residual signal, the reference signal, and the target signal, and estimating the echo residual state; determining a sub-band non-linear gain parameter according to a preset suppression coefficient and the signal coherence coefficient; performing sub-band processing on the first residual signal in the frequency domain, and suppressing the gain of each band signal according to the non-linear gain parameter; and superimposing the suppressed sub-band signals to obtain a preliminary residual signal.

[0043] In practice, first, windowing and Fourier transform processing are performed on the target signal d(n) and the first residual signal error(n). The windowing processing can select a Hanning window, and the windowed signals are respectively d w (n) = d(n)w(n) and

[0044] error w (n) = error(n)w(n), where w(n) is the Hanning window function. Then Fourier transform is performed to obtain d(n)' = FFT[d w (n)] and error(n)' = FFT[error w (n)]. Then, the power spectra of the target signal and the first residual signal are calculated. The power spectrum of the target signal is: Pd(n) = α * Pd(n - 1) + (1 - α) * real(d(n) * d(n)'), and the power spectrum of the first residual signal is: Pe(n) = α * Pe(n - 1) + (1 - α) * real(error(n) * error(n)'). Based on the power spectral densities of the first residual signal, the reference signal, and the target signal, the coherence coefficient γ(k) is calculated. The coherence coefficient reflects the correlation of different signals in the frequency domain, and its calculation formula can be determined according to specific signal processing theories. For example where P sx (k) is the cross-power spectrum of the target signal and the reference signal. Then, according to the filter error e f (n) and the target signal d(n), the divergence and convergence states of the current frame filter are analyzed. Through the time-domain response h(n) of the filter, the filter response peak A is calculatedmax = max|h(n)| and the maximum gain value G of the echo path max Calculate the path delay τ with respect to the filter d Update the active states of the far - end and target signals. According to the echo reverberation model, calculate the power P of the reference signal plus the reverberation signal x+r (k). Through the maximum echo gain G max and the maximum value x of the reference signal max , determine whether the echo signal is oversaturated. If the echo is oversaturated, there is residual echo in the first residual signal. Estimate the residual echo power spectrum P re (k).

[0045] According to the preset suppression coefficient α and the signal coherence coefficient γ(k), determine the sub - band non - linear gain parameter g(k). When calculating the non - linear gain value, there are preset thresholds T1 and T2. Calculate the ratios of the residual echo power P re (k) and the error power P e1 (k) respectively and the ratio of the residual echo power P re (k) and the noise power P n (k) respectively When both of these ratios are greater than the preset threshold, perform gain calculation, such as g(k)=α×(1 - γ(k)); otherwise, by default, no processing is done, g(k)=1. This is to reduce the suppression effect of this part, avoid loss of the target signal, and prevent the phenomenon of intermittent "voice loss" of the proximal voice during the call. Then, perform sub - band processing in the frequency domain on the first residual signal, divide it into multiple frequency bands E 1i (k), i = 1, 2, …, N, where N is the number of frequency bands. Suppress the gain of each frequency - band signal according to the non - linear gain parameter g(k) to obtain the suppressed sub - band signal E 2i (k)=g(k)E 1i (k). Superimpose the suppressed sub - band signals E 2i (k) to generate the second residual signal in the frequency domain

[0046] In this embodiment, the step of "detecting the preliminary residual signal by the voice energy detection method and determining whether there is residual echo in the preliminary residual signal, and if there is residual echo in the preliminary residual signal, performing echo cancellation on the residual signal to obtain the second call voice signal" specifically includes: setting preset parameters, and the preset parameters include a preset window length L, a preset power threshold P th , a preset effective speech percentage R thand a preset suppression coefficient α; calculate the power and the point-by-point power of the preliminary residual signal; calculate the effective speech percentage by comparing the point-by-point power with a preset power threshold; if the effective speech percentage is less than a preset percentage, multiply the preliminary residual signal by the preset suppression coefficient and add comfort noise to obtain the second call voice signal.

[0047] In practice, the setting of preset parameters is the basis of the entire speech energy detection method, including a preset window length L, a preset power threshold P th , a preset effective speech percentage R th and a preset suppression coefficient α. The preset power threshold P th and the preset effective speech percentage R th can be obtained in the following ways: estimated according to the real vehicle test environment. Since the ratio of the effective speech and the residual echo energy in the preliminary residual signal obtained through the previous steps is relatively large, appropriate preset parameters can effectively distinguish and identify, while reducing the computing power consumption and improving the call fluency; calculated according to the coherence between the target signal, the reference signal and the preliminary residual signal after the current frame; it can be understood that dynamically adjusting the preset parameters in the energy detection algorithm according to the actual in-vehicle acoustic environment can ensure the adaptability of the algorithm and improve the accuracy of the algorithm in complex and changing acoustic environments. The power P e of the preliminary residual signal = ∑error(i) 2 , where i is the number of sampling points within the preset window length L; the point-by-point power K(j) of the preliminary residual signal = |e(j)| 2 , and e(j) is the amplitude of the j-th point of the preliminary residual signal.

[0048] By comparing the point-by-point power K(j) of the preliminary residual signal with the preset power threshold P th , calculate the effective speech percentage R of the preliminary residual signal in the current frame. The specific operation is: first calculate the ratio of the point-by-point power K(j) to the power P e of the preliminary residual signal and combine it with the preset power threshold P th to judge whether each point is a valid signal and count. If K(j)>P th , then this point is a valid signal, and the effective signal percentage If R>R th , it means that there is no residual echo in the current frame; otherwise, there is a residual echo in the current frame. When it is judged that there is a residual echo, multiply the preliminary residual signal e(n) of the current frame by the preset suppression coefficient α to obtain αe(n), and add comfort noise n c (n) to obtain the second call voice signal y(n) = αe(n) + n c (n), where M is the number of valid signals and N is the total number of points of the preliminary residual signal.

[0049] The comfort noise is used to mitigate the uncomfortable effects such as pauses, voicelessness, and distortion that may occur during the cancellation step. Its basic principle is to mix the noise signal with the signal after echo cancellation processing, making the mixed signal more natural and comfortable audibly. First, calculate the smoothed target signal energy P ssin , and take the minimum value of P ssin and the initial noise power Pnoise0; calculate the noise power parameter Ψ, and its expression is: Ψ = Pnoise0 - min(P ssin , Pnoise0); thus calculate the noise power Pnoise, and its expression is: Pnoise = (min(P ssin , Pnoise0) + ε1 * Ψ) * ε2, where both ε1 and ε2 are initially given step sizes. After that, use the linear multiplicative congruential method to generate the initial random number X1 value, and then generate X2, X3... X n through a preset formula. The preset formula is: It can be understood that by generating comfort noise, white noise signals with specific frequencies can be added to the residual signal to simulate the auditory characteristics of the human ear, thereby improving the naturalness of the sound.

[0050] Figure 3 FIG. 18 shows the echo cancellation results of the same data by the existing echo cancellation method and the method of the present invention. Figure 3 From top to bottom in FIG. 18 are the reference signal, the target signal, and the processing result curve. In the target signal curve, the call status in the intervals of abscissa 10 - 28 and 40 - 55 is single talk, and only the echo projected by the reference signal exists; the call status in the interval of 60 - 72 is double talk, and the target signal is mixed with the reference signal echo. The processing result curve of the comparative document method shows that in the single talk mode, although the amplitude of the echo decreases but is not completely eliminated and is still within the recognizable range; in the double talk mode, the echo is hardly eliminated; the processing result curve of the method of the present invention shows that in the single talk mode, the echo is almost completely eliminated and presents a straight line; in the double talk mode, the echo signal is effectively eliminated. It can be understood that after processing the target signal through linear adaptive filtering and non - linear filtering, and then through the energy detection method, further residual echo judgment and elimination are performed on each frame of voice signal, reducing the residual echo in the signal transmitted to the far - end while avoiding any adverse factors such as voicelessness that affect the smoothness of the call, thereby improving the quality of the in - vehicle Bluetooth hands - free call process.

[0051] Embodiment 2 of the present invention provides an in - vehicle Bluetooth call echo cancellation device, as shown in Figure 4 , which includes the following modules:

[0052] An information acquisition module 40 is configured to acquire a first call voice signal and preprocess the first call voice signal to obtain a signal to be detected, where the signal to be detected includes a reference signal and a target signal.

[0053] A status determination module 41 is configured to determine whether the call status of the signal to be detected is a two-way call based on a two-way call detection method.

[0054] An echo cancellation module 42, if the call status is a two-way call, performs a first echo cancellation process on the signal to be detected through a combination of an adaptive filtering algorithm and a non-linear filtering algorithm to obtain a preliminary residual signal. Then, the preliminary residual signal is detected through a voice energy detection method to determine whether there is residual echo in the preliminary residual signal. If there is residual echo in the preliminary residual signal, a second echo cancellation process is performed on the residual signal to obtain a second call voice signal.

[0055] Embodiment 3 of the present invention provides a storage medium storing program instructions, and when the program instructions are executed, the echo cancellation method described in any one of the above is implemented.

[0056] Embodiment 4 of the present invention provides an electronic terminal, as Figure 5 shown, including a processor and a memory, where the memory stores program instructions, and the processor runs the program instructions to implement the echo cancellation method described in any one of the above.

[0057] The present invention may be an apparatus, a method, and / or a computer program product. The computer program product may include a readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present invention.

[0058] The storage medium may be a tangible device that holds and stores instructions for use by an instruction execution device. The storage medium may include, for example, but is not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punched card or raised structures in a groove having instructions stored thereon, and any suitable combination of the foregoing.

[0059] It should be understood that although this specification is described in terms of embodiments, not every embodiment contains only an independent technical solution. This narrative style of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

[0060] The series of detailed descriptions listed above are only specific descriptions of the feasible embodiments of the present invention, and they are not intended to limit the protection scope of the present invention. Any equivalent embodiments or modifications made without departing from the technical spirit of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for eliminating echo in in-vehicle Bluetooth calls, characterized in that, Including the following steps: Obtain a first call voice signal and preprocess the first call voice signal to obtain a signal to be detected, where the signal to be detected includes a reference signal and a target signal; Based on a two-talk detection method, determine whether the call state of the signal to be detected is two-talk; If the call state is two-talk, perform a first echo cancellation process on the signal to be detected through a combination of an adaptive filtering algorithm and a non-linear filtering algorithm to obtain a preliminary residual signal. Then, detect the preliminary residual signal through a voice energy detection method and determine whether there is residual echo in the preliminary residual signal; If there is residual echo in the preliminary residual signal, perform a second echo cancellation process on the residual signal to obtain a second call voice signal.

2. The echo cancellation method according to claim 1, characterized in that, The "obtain a first call voice signal and preprocess the first call voice signal to obtain a signal to be detected" specifically includes: performing frame division processing on both the reference signal and the target signal, and performing block division processing on the frame-divided reference signal and the frame-divided target signal according to a preset length to obtain the signal to be detected.

3. The echo cancellation method according to claim 1, characterized in that, The "based on a two-talk detection method, determine whether the call state of the signal to be detected is two-talk" specifically includes: Calculate the cross-correlation value between the target signal and the reference signal, the first autocorrelation value of the target signal, and the second autocorrelation value of the reference signal; generate a two-talk detection decision parameter according to the cross-correlation value and / or the first autocorrelation value and / or the second autocorrelation value of the reference signal; compare the two-talk detection decision parameter with a preset threshold, and determine whether the current call state is two-talk according to the comparison result.

4. The echo cancellation method according to claim 1, wherein The "perform a first echo cancellation process on the signal to be detected through a combination of an adaptive filtering algorithm and a non-linear filtering algorithm to obtain a preliminary residual signal" specifically includes: performing Fourier transform on the signal to be detected and processing it through an adaptive filtering algorithm to obtain a first residual signal; processing the first residual signal through non-linear filtering to obtain a preliminary residual signal.

5. The echo cancellation method according to claim 4, wherein The "perform Fourier transform on the signal to be detected and process it through an adaptive filtering algorithm to obtain a first residual signal" specifically includes: The adaptive filtering algorithm is a normalized least mean square adaptive filtering algorithm; perform Fourier transform on the signal to be detected; perform zero-padding operation on the filter weight coefficient of the adaptive filtering algorithm, and then perform block filtering processing on the reference signal to obtain a filtering result; perform inverse Fourier transform on the filtering result to obtain the first residual signal; The block filtering process is as follows: where x(n) is the reference signal, d(n) is the reference signal, ω T is the filter weight coefficient at time period T; error(n) is the first residual signal.

6. The echo cancellation method according to claim 4, characterized in that, The "process the first residual signal through non-linear filtering to obtain a preliminary residual signal" specifically includes: Calculate a coherence coefficient based on the power spectral densities of the first residual signal, the reference signal, and the target signal, and estimate the echo residual state; determine a sub-band non-linear gain parameter according to a preset suppression coefficient and the signal coherence coefficient; perform sub-band processing in the frequency domain on the first residual signal, and perform gain suppression on the signals in each band according to the non-linear gain parameter; superimpose the suppressed sub-band signals to generate a second residual signal in the frequency domain; perform inverse Fourier transform on the second residual signal in the frequency domain to obtain a preliminary residual signal.

7. The echo cancellation method according to claim 1, wherein The step of "detecting the preliminary residual signal through a voice energy detection method and determining whether there is residual echo in the preliminary residual signal, and if there is residual echo in the preliminary residual signal, performing echo cancellation on the residual signal to obtain a second call voice signal" specifically includes: Set preset parameters, where the preset parameters include a preset window length L, a preset power threshold P th , a preset effective speech percentage R th and a preset suppression coefficient α; calculate the power and the point-by-point power of the preliminary residual signal; calculate the effective speech percentage by comparing the point-by-point power with the preset power threshold; if the effective speech percentage is less than the preset percentage, multiply the preliminary residual signal by the preset suppression coefficient and add comfort noise to obtain the second call voice signal.

8. An in-vehicle Bluetooth call echo cancellation device, characterized in that, It includes the following modules: An information acquisition module, configured to acquire a first call voice signal and perform preprocessing on the first call voice signal to obtain a signal to be detected, where the signal to be detected includes a reference signal and a target signal; A state judgment module, configured to judge whether the call state of the signal to be detected is a double talk based on a double talk detection method; An echo cancellation module, if the call state is a double talk, performing a first echo cancellation process on the signal to be detected through a combination of an adaptive filtering algorithm and a non-linear filtering algorithm to obtain a preliminary residual signal, and then, detecting the preliminary residual signal through a voice energy detection method and determining whether there is residual echo in the preliminary residual signal; If there is residual echo in the preliminary residual signal, performing a second echo cancellation process on the residual signal to obtain a second call voice signal.

9. A storage medium stores program instructions, characterized in that, When the program instructions are executed, the echo cancellation method described in any one of claims 1 to 7 is implemented.

10. An electronic terminal, characterized in that, It includes a processor and a memory, the memory stores program instructions, and the processor runs the program instructions to implement the echo cancellation method described in any one of claims 1 to 7.

Citation Information

Cited By

  • In-vehicle call method and device, storage medium and electronic terminal

    CN120748424A