Call echo cancellation method for vehicle-mounted equipment

By obtaining pre-processing, correlation calculation and adaptive filtering of microphone and remote voice signals, the problem of poor echo cancellation effect of vehicle-mounted equipment is solved, efficient echo cancellation and stable communication are achieved, and it adapts to complex vehicle-mounted environments.

CN120748427APending Publication Date: 2025-10-03DALIAN HAITIAN IND TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510920720.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing echo cancellation technology for in-vehicle equipment is ineffective in complex acoustic environments, with residual echo, excessive latency, or system instability, making it difficult to adapt to resource-constrained in-vehicle embedded systems.

Method used

By acquiring microphone signals and far-end voice signals, performing preprocessing and buffer settings, calculating correlation, generating echo estimation signals and performing adaptive filtering, combining frequency domain analysis to suppress howling and achieve efficient echo cancellation.

Benefits of technology

It significantly improves the quality of in-vehicle voice communications, ensures call clarity, enhances adaptability and stability in complex in-vehicle environments, controls processing delays at a low level, and meets real-time communication needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120748427A_ABST
    Figure CN120748427A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of vehicle-mounted equipment echo cancellation, and discloses a vehicle-mounted equipment call echo cancellation method, which comprises the following steps: acquiring a microphone signal and a far-end voice signal; wherein the microphone signal comprises a near-end voice signal and an echo signal; preprocessing the near-end voice signal and the far-end voice signal, and setting a buffer region interval; calculating the correlation between the microphone signal and the remote voice signal according to the microphone signal and the remote voice signal; generating an echo estimation signal according to the far-end voice signal and the estimation of the echo path, and removing the echo estimation signal from the microphone signal; and updating the estimation of the echo path, and performing post-processing on the microphone signal from which the echo estimation signal is removed. According to the method, the vehicle-mounted voice communication quality is remarkably improved, the echo interference is effectively eliminated, the adaptability and the stability in a complex vehicle-mounted environment are enhanced, and the real-time communication requirement is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of echo cancellation of vehicle-mounted equipment, and in particular to a method for canceling call echo of vehicle-mounted equipment. Background Art

[0002] Against the backdrop of the rapid development of intelligent transportation and in-vehicle communication technologies, real-time voice communication between on-board devices (such as the DACU in the driver's cab and the PECU in the passenger compartment) has become a key link in ensuring driving safety and improving service quality.

[0003] However, due to the complex acoustic characteristics of the in-vehicle environment (such as sound reflections in enclosed spaces and engine noise interference) and equipment layout limitations, the voice signals collected by the microphone often contain a large amount of echo and noise, seriously affecting the clarity and intelligibility of calls. Traditional echo cancellation technologies (such as fixed filter methods and simple spectral subtraction) lack environmental adaptability and are prone to problems such as residual echo, excessive latency, or system instability in in-vehicle scenarios. Although deep learning-based methods have excellent theoretical effects, they require high hardware computing power and are difficult to adapt to resource-constrained in-vehicle embedded systems (such as FPGA chips). In addition, existing solutions have shortcomings in processing buffer size and balancing signal latency and stability, which often lead to voice freezes or incomplete echo suppression.

[0004] Therefore, it is necessary to provide a method for canceling echo of a call in a vehicle-mounted device to solve the problem of poor echo suppression effect of the echo cancellation technology in the prior art. Summary of the Invention

[0005] In view of this, the present invention proposes a method for canceling call echo of an in-vehicle device, aiming to solve the problem of poor echo suppression effect of the echo cancellation technology in the prior art.

[0006] The present invention proposes a method for canceling echo of a vehicle-mounted device call, comprising:

[0007] Acquire a microphone signal and a far-end voice signal; wherein the microphone signal includes a near-end voice signal and an echo signal;

[0008] Preprocessing the near-end voice signal and the far-end voice signal, and setting a buffer interval;

[0009] Calculating a correlation between the microphone signal and the far-end voice signal;

[0010] generating an echo estimation signal based on the far-end speech signal and an estimation of the echo path, and removing the echo estimation signal from the microphone signal;

[0011] The estimate of the echo path is updated and the microphone signal is post-processed with the echo estimate signal removed.

[0012] Furthermore, the preprocessing of the near-end voice signal and the far-end voice signal includes:

[0013] Removing DC components from the near-end speech signal and the far-end speech signal by a high-pass filter;

[0014] Dynamic gain control based on short-time energy detection adjusts the amplitudes of the near-end voice signal and the far-end voice signal to a target range;

[0015] Performing frame peak normalization processing on the near-end speech signal and the far-end speech signal.

[0016] Furthermore, the setting of the buffer zone includes:

[0017] Obtain historical data and set the buffer lower limit according to the maximum value τmax of the historical echo delay in the historical data:

[0018] Nmin=fs×(τmax+Δτ);

[0019] In the above formula, Nmin represents the lower limit of the buffer, fs represents the signal sampling rate, and Δτ is the reserved safety margin;

[0020] Set the buffer limit based on the delay threshold Tmax:

[0021] Nmax=fs×Tmax;

[0022] In the above formula, Nmax represents the upper limit of the buffer, and fs represents the signal sampling rate;

[0023] The buffer interval is obtained as [Nmin, Nmax].

[0024] Furthermore, the calculating of the correlation between the microphone signal and the far-end voice signal includes:

[0025] The correlation between the microphone signal and the far-end voice signal is calculated using the following formula:

[0026]

[0027] In the above formula, R xy (k) represents the correlation between the microphone signal and the far-end speech signal, N represents the number of samples of the audio frame selected for processing, y(n) represents the microphone signal, x(n) represents the far-end speech signal, k represents the sampling delay length of the original sound and the echo, and n represents the nth audio sample signal.

[0028] Furthermore, the generating of the echo estimation signal according to the far-end speech signal and the estimation of the echo path includes:

[0029] The echo estimation signal is obtained by the following formula:

[0030] d^(n)=h(n)*x(n);

[0031] In the above formula, d^(n) represents the echo estimation signal, h(n) represents the echo estimation path, and x(n) represents the far-end speech signal.

[0032] Furthermore, the step of removing the echo estimation signal from the microphone signal includes:

[0033] The echo estimation signal is removed by the following formula:

[0034] e(n)=y(n)-d^(n);

[0035] In the above formula, e(n) represents the signal after echo cancellation, y(n) represents the microphone signal, and d^(n) represents the echo estimation signal.

[0036] Furthermore, the updating of the echo path estimation includes:

[0037] Update the echo path estimate using an adaptive filter:

[0038] h(n+1)=h(n)+μ·e(n)·x(n);

[0039] In the above formula, h(n+1) represents the updated echo path, h(n) represents the echo estimation path, μ represents the step size factor, e(n) represents the signal after echo cancellation, and x(n) represents the far-end speech signal.

[0040] Furthermore, the post-processing of the microphone signal from which the echo estimation signal is removed includes:

[0041] Perform fast Fourier transform on the echo-cancelled signal, calculate the power spectrum, traverse all frequencies within 5kHz, and calculate the energy change rate;

[0042] Marking frequency points according to the energy change rate, and determining whether howling occurs;

[0043] An IIR notch filter with Q=8 is designed for each howling frequency to perform filtering.

[0044] Furthermore, the marking of frequency points according to the energy change rate to determine whether howling occurs includes:

[0045] The frequency points where the energy change rate is greater than 3.0 and the bandwidth is less than 100 Hz are marked. The marked points are counted continuously. If the frequency points appear for more than or equal to 3 consecutive frames, howling is confirmed.

[0046] Furthermore, the IIR notch filter with Q=8 is designed for each howling frequency, and after filtering, the filtering process includes:

[0047] When the energy of five consecutive frames decreases, the IIR notch filter corresponding to the marked point is removed.

[0048] Compared with the prior art, the beneficial effects of the present invention are as follows: the present invention first obtains the near-end voice signal containing echo and the far-end voice signal respectively through the microphone and the communication bus, laying the data foundation. The pre-processing step improves the signal quality, and at the same time reasonably sets the buffer interval to balance the delay and stability; the correlation calculation step uses the cross-correlation function and peak detection to determine the echo delay, providing key parameters for subsequent processing; the echo estimation and elimination step generates an echo estimation signal based on the far-end voice signal and delay and subtracts it from the microphone signal to achieve preliminary echo elimination; the updated echo path estimation can adapt to environmental changes; the post-processing further eliminates residual echo and noise, and improves voice clarity. Furthermore, the present invention significantly improves the quality of in-vehicle voice communication, effectively eliminates echo interference, makes the voices of both parties in the call clear and distinguishable, and improves user experience; through pre-processing and adaptive mechanisms, it enhances adaptability and stability in complex in-vehicle environments (such as tunnels, downtown, etc.); reasonable buffer settings and efficient algorithm design control the processing delay at a low level while ensuring the echo cancellation effect, meeting real-time communication needs. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present invention. The same reference symbols are used throughout the drawings to represent the same components. In the drawings:

[0050] Figure 1 This is a flow chart of a method for canceling call echo in an in-vehicle device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0051] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art. It should be noted that, unless there is a conflict, the embodiments of the present disclosure and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with the embodiments.

[0052] In some embodiments of the present application, see Figure 1 As shown, this embodiment provides a method for canceling a call echo of an in-vehicle device, comprising the following steps:

[0053] S100, obtaining a microphone signal and a far-end voice signal; wherein the microphone signal includes a near-end voice signal and an echo signal;

[0054] S200, preprocessing the near-end voice signal and the far-end voice signal, and setting a buffer interval;

[0055] S300, calculating a correlation between the microphone signal and the far-end voice signal;

[0056] S400, generating an echo estimation signal according to the far-end speech signal and an estimation of the echo path, and removing the echo estimation signal from the microphone signal;

[0057] S500: Update the estimation of the echo path, and perform post-processing on the microphone signal from which the echo estimation signal is removed.

[0058] It can be understood that the method first obtains the near-end voice signal containing echo and the far-end voice signal through the microphone and the communication bus respectively, laying the data foundation. The preprocessing step improves the signal quality, and at the same time reasonably sets the buffer interval to balance the delay and stability; the correlation calculation step uses the cross-correlation function and peak detection to determine the echo delay, providing key parameters for subsequent processing; the echo estimation and elimination step generates an echo estimation signal based on the far-end voice signal and delay and subtracts it from the microphone signal to achieve preliminary echo elimination; the updated echo path estimation can adapt to environmental changes; the post-processing further eliminates residual echo and noise to improve voice clarity. Furthermore, the present invention significantly improves the quality of in-vehicle voice communication, effectively eliminates echo interference, makes the voices of both parties in the call clear and distinguishable, and improves the user experience; through preprocessing and adaptive mechanisms, it enhances adaptability and stability in complex in-vehicle environments (such as tunnels, downtown areas, etc.); reasonable buffer settings and efficient algorithm design, while ensuring the echo cancellation effect, control the processing delay at a low level to meet real-time communication needs.

[0059] In some embodiments of the present application, the preprocessing of the near-end voice signal and the far-end voice signal includes:

[0060] Removing DC components from the near-end speech signal and the far-end speech signal by a high-pass filter;

[0061] Dynamic gain control based on short-time energy detection adjusts the amplitudes of the near-end voice signal and the far-end voice signal to a target range;

[0062] Performing frame peak normalization processing on the near-end speech signal and the far-end speech signal.

[0063] As can be understood, the preprocessing scheme of the present invention significantly improves the effectiveness and stability of subsequent echo cancellation through a triple signal optimization mechanism. First, DC component suppression enhances signal quality. In an in-vehicle environment, microphone circuit bias voltage or electromagnetic interference often introduces a DC component, causing signal baseline drift. The application of a high-pass filter (e.g., with a cutoff frequency of 20Hz) effectively filters out this DC component, preventing its interference with subsequent adaptive filtering. Second, dynamic gain control balances signal energy. Dynamic gain control (AGC) based on short-term energy detection can adaptively adjust the signal amplitude. In an in-vehicle scenario, near-end speech may fluctuate significantly due to changes in user distance or volume, while far-end speech is affected by communication link attenuation. AGC calculates the short-term signal energy (e.g., a 30ms window) and dynamically adjusts the gain (e.g., a target energy of -15dBFS) to maintain the signal within an appropriate dynamic range. This operation not only avoids clipping distortion but also compresses the signal energy fluctuation range from ±12dB to ±3dB, significantly improving the stability of the adaptive filtering. In addition, frame-based peak normalization improves algorithm accuracy. Frame-based peak normalization (e.g., one frame every 10ms) normalizes the signal amplitude to a specific range, eliminating amplitude variations introduced by different devices or environments. This process is crucial for echo cancellation algorithms: on the one hand, it stabilizes the adaptive filter step size parameters (e.g., the μ value in the NLMS algorithm) across the entire frequency band, preventing localized divergence caused by uneven signal amplitude. On the other hand, the normalized signal produces clearer peaks in correlation calculations, reducing the echo delay estimation error from approximately ±2ms to ±0.5ms, significantly improving the accuracy of echo path modeling. The synergistic effect of preprocessing makes the preprocessed signal more suitable for adaptive filtering.

[0064] In some embodiments of the present application, the setting of the buffer zone includes:

[0065] Obtain historical data and set the buffer lower limit according to the maximum value τmax of the historical echo delay in the historical data:

[0066] Nmin=fs×(τmax+Δτ);

[0067] In the above formula, Nmin represents the lower limit of the buffer, fs represents the signal sampling rate, and Δτ is the reserved safety margin;

[0068] Set the buffer limit based on the delay threshold Tmax:

[0069] Nmax=fs×Tmax;

[0070] In the above formula, Nmax represents the upper limit of the buffer, and fs represents the signal sampling rate;

[0071] The buffer interval is obtained as [Nmin, Nmax].

[0072] It is understandable that the method of dynamically setting the buffer interval based on historical data significantly improves the robustness and real-time performance of echo cancellation by accurately quantifying the delay boundary. On the one hand, the lower limit Nmin is set based on the historical maximum echo delay τmax, and a safety margin Δτ (such as 5-10ms) is reserved to ensure that the echo path changes in extreme cases can be captured and to avoid echo residuals caused by insufficient delay estimation; on the other hand, the buffer upper limit Nmax is limited by the delay threshold Tmax to prevent excessive buffering from introducing unacceptable communication delays (such as the industry standard requirement that the one-way delay is less than 200ms). This dual-boundary constraint mechanism achieves the best balance between echo cancellation effect and system real-time performance. Specifically, in a certain vehicle-mounted system, the sampling rate fs = 16kHz, and through historical data statistics, it is found that the maximum echo delay τmax in tunnel scenarios is 40ms, and the delay in normal scenarios is about 20ms. The lower limit of the buffer, Nmin, is calculated as 16000 × (0.04 + 0.005) = 720 samples. If the delay threshold, Tmax, is set to 100ms, then the upper limit, Nmax, is calculated as 16000 × 0.1 = 1600 samples. During runtime, the buffer size is dynamically adjusted based on the current scenario: in complex scenarios like tunnels, a buffer close to Nmin (e.g., 800 samples) is used to ensure effective echo cancellation; in standard scenarios, a smaller buffer (e.g., 400 samples) is used to reduce latency.

[0073] In some embodiments of the present application, the calculating of the correlation between the microphone signal and the far-end voice signal includes:

[0074] The correlation between the microphone signal and the far-end voice signal is calculated using the following formula:

[0075]

[0076] In the above formula, R xy (k) represents the correlation between the microphone signal and the far-end speech signal, N represents the number of samples of the audio frame selected for processing, y(n) represents the microphone signal, x(n) represents the far-end speech signal, k represents the sampling delay length of the original sound and the echo, and n represents the nth audio sample signal.

[0077] It's easy to understand that the delay estimation method based on the cross-correlation function provides highly accurate delay parameters for echo cancellation by mathematically quantifying the temporal correlation between the microphone signal and the far-end speech. First, the cross-correlation peak position directly corresponds to the propagation delay of the echo path, intuitively reflecting the physical characteristics of the acoustic system. Second, the time-domain integral operation (the summation term in the formula) effectively smooths random noise, maintaining high estimation accuracy, especially in low signal-to-noise ratio environments. Furthermore, compared to frequency-domain methods (such as generalized cross-correlation), direct time-domain calculation avoids the overhead of FFT transforms, making it more suitable for resource-constrained in-vehicle embedded systems.

[0078] In some embodiments of the present application, generating an echo estimation signal based on a far-end speech signal and an estimation of an echo path includes:

[0079] The echo estimation signal is obtained by the following formula:

[0080] d^(n)=h(n)*x(n);

[0081] In the above formula, d^(n) represents the echo estimation signal, h(n) represents the echo estimation path, and x(n) represents the far-end speech signal.

[0082] It's understandable that, first, combining the echo path characteristics with the far-end voice signal accurately simulates the actual echo generation process, providing a reliable reference for subsequent echo cancellation. Secondly, using this formula to generate an estimated signal allows for close integration of the echo estimation process with the adaptive filtering algorithm, facilitating dynamic adjustments based on environmental changes, thereby improving the real-time performance and accuracy of echo cancellation. Furthermore, this approach offers clear computational logic and manageable computational complexity, making it suitable for efficient implementation on the FPGA chip of an onboard device. Specifically, on a moving bus, the onboard device acquires a far-end voice signal with a sampling rate of 16kHz. Preliminary correlation calculations indicate an echo path delay of approximately 10ms, or 80 sampling points. An adaptive filtering algorithm is used to preliminarily determine the echo path estimate, which is then convolved with the far-end voice signal to produce the echo estimate signal.

[0083] In some embodiments of the present application, removing the echo estimation signal from the microphone signal includes:

[0084] The echo estimation signal is removed by the following formula:

[0085] e(n)=y(n)-d^(n);

[0086] In the above formula, e(n) represents the signal after echo cancellation, y(n) represents the microphone signal, and d^(n) represents the echo estimation signal.

[0087] As you can understand, this formula is based on a simple time-domain subtraction, subtracting the original microphone signal containing near-end speech and echo from the simulated echo estimation signal sample by sample. This method intuitively and quickly cancels the echo component, directly outputting a relatively pure near-end speech signal. This subtraction operation has low computational complexity and adapts to the limited hardware resources of in-vehicle equipment. It also works closely with the adaptive filtering algorithm, allowing real-time updates based on environmental changes to continuously optimize the echo cancellation effect. For example, in an in-vehicle communication system, the driver and passengers communicate through the in-vehicle device. Due to the enclosed space structure of the vehicle, the voice captured by the microphone signal is mixed with a significant echo. By first generating an echo estimation signal based on the far-end speech signal and then processing it using the formula, the clarity of the conversation between the driver and passenger can be significantly improved, echo interference is essentially eliminated, and the quality of in-vehicle voice communication can be effectively guaranteed.

[0088] In some embodiments of the present application, updating the echo path estimate includes:

[0089] Update the echo path estimate using an adaptive filter:

[0090] h(n+1)=h(n)+μ·e(n)·x(n);

[0091] In the above formula, h(n+1) represents the updated echo path, h(n) represents the echo estimation path, μ represents the step size factor, e(n) represents the signal after echo cancellation, and x(n) represents the far-end speech signal.

[0092] It's no secret that the adaptive filter-based echo path update mechanism significantly improves the robustness of echo cancellation by learning about changes in the acoustic environment in real time. First, the residual signal after echo cancellation is used as error feedback to continuously fine-tune the echo path estimate, tracking the time-varying characteristics of the echo path caused by factors such as passenger movement and window openings in the vehicle environment. Second, the introduction of a step-size factor balances the algorithm's convergence speed with steady-state error, maintaining estimation accuracy in complex acoustic scenarios. For example, when a vehicle processes speech signals at a 16kHz sampling rate and a vehicle enters a tunnel from an open road, the echo path delay suddenly increases from 20ms to 40ms. Upon detecting an increase in residual signal energy, the adaptive filter formula adjusts the step-size factor. If the current estimate corresponds to a 20ms delay, then after receiving 100 frames of data (approximately 6.25ms), the estimated delay gradually approaches 40ms, restoring the echo suppression ratio from an initial 15dB to over 25dB, ensuring clear voice quality for both callers.

[0093] In some embodiments of the present application, the post-processing of the microphone signal from which the echo estimation signal is removed includes:

[0094] Perform fast Fourier transform on the echo-cancelled signal, calculate the power spectrum, traverse all frequencies within 5kHz, and calculate the energy change rate;

[0095] Marking frequency points according to the energy change rate, and determining whether howling occurs;

[0096] An IIR notch filter with Q=8 is designed for each howling frequency to perform filtering.

[0097] It is understandable that this post-processing solution based on frequency domain analysis and adaptive filtering significantly improves the stability and clarity of in-vehicle voice communications by accurately identifying and suppressing howling frequencies. By converting the time domain signal into a frequency domain representation through Fast Fourier Transform (FFT), it can capture energy distribution changes within 5kHz in real time and accurately lock the howling frequency (usually manifested as a sudden increase in narrowband energy). It dynamically generates an IIR notch filter with Q=8 (the higher the Q value, the narrower the bandwidth) based on the energy change rate, suppressing howling while preserving the voice energy to the greatest extent. The filtered signal is continuously monitored to ensure that howling is effectively suppressed and no new howling frequencies are generated.

[0098] In some embodiments of the present application, marking the frequency points according to the energy change rate to determine whether howling occurs includes:

[0099] The frequency points where the energy change rate is greater than 3.0 and the bandwidth is less than 100 Hz are marked. The marked points are counted continuously. If the frequency points appear for more than or equal to 3 consecutive frames, howling is confirmed.

[0100] As can be understood, this howling detection rule based on energy change rate, bandwidth, and duration of frames achieves high accuracy and low false positive rate in howling detection through multi-dimensional feature constraints. First, using an energy change rate greater than 3.0 as a screening criterion effectively identifies frequency points that experience significant energy spikes compared to the historical average, distinguishing normal speech fluctuations from howling. Limiting the bandwidth to less than 100Hz leverages the physical law of howling frequencies exhibiting narrowband characteristics to eliminate interference from broadband noise or speech harmonics. Furthermore, counting the markers for three consecutive frames further mitigates false positives caused by transient signal fluctuations. Suppression measures are only triggered when a frequency point consistently exhibits howling characteristics. This combined judgment method ensures rapid response to sudden howling events while significantly reducing the probability of false triggers. In the complex acoustic environment of vehicles, it can accurately capture true howling events, ensuring the stability and reliability of voice communications and avoiding unnecessary damage to normal voice signals caused by false positives.

[0101] In some embodiments of the present application, the design of an IIR notch filter with Q=8 for each howling frequency, after filtering processing, includes:

[0102] When the energy of five consecutive frames decreases, the IIR notch filter corresponding to the marked point is removed.

[0103] It is understandable that the dynamic removal mechanism of the filter based on the continuous decline in energy significantly improves the flexibility of howling suppression and voice protection capabilities through closed-loop feedback. First, when the howling frequency energy is detected to drop to a normal level for 5 consecutive frames (about 50-100ms), it is determined that the howling has been effectively suppressed, and the filter is removed in time to prevent the voice components in this frequency band (such as high-frequency harmonics) from being excessively attenuated due to long-term filtering, thereby ensuring the naturalness of the voice. Secondly, in the vehicle-mounted scenario, howling may occur briefly due to factors such as passenger movement and device position adjustment. The dynamic removal mechanism can quickly release filtering resources, avoid interference of fixed filters on subsequent normal signals, and improve the system's adaptability to dynamic environments. Through the complete closed loop of "detection-suppression-release", a dynamic balance between acoustic performance and voice quality is achieved, which is especially suitable for frequent short-term howling scenarios in the vehicle environment.

[0104] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or a combination of software and hardware embodiments. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0105] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0106] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0107] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0108] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the scope of protection of the claims of the present invention.

Claims

1. A method for canceling echo of a vehicle-mounted device, characterized in that: include: Acquire a microphone signal and a far-end voice signal; wherein the microphone signal includes a near-end voice signal and an echo signal; Preprocessing the near-end voice signal and the far-end voice signal, and setting a buffer interval; Calculating a correlation between the microphone signal and the far-end voice signal; generating an echo estimation signal based on the far-end speech signal and an estimation of the echo path, and removing the echo estimation signal from the microphone signal; The estimate of the echo path is updated and the microphone signal is post-processed with the echo estimate signal removed.

2. The method for canceling echo of vehicle-mounted equipment according to claim 1, characterized in that: The preprocessing of the near-end voice signal and the far-end voice signal includes: Removing DC components from the near-end speech signal and the far-end speech signal by a high-pass filter; Dynamic gain control based on short-time energy detection adjusts the amplitudes of the near-end voice signal and the far-end voice signal to a target range; Performing frame peak normalization processing on the near-end speech signal and the far-end speech signal.

3. The method for canceling echo of vehicle-mounted equipment according to claim 1, characterized in that: The setting of the buffer zone includes: Obtain historical data and set the buffer lower limit according to the maximum value τmax of the historical echo delay in the historical data: Nmin=fs×(τmax+Δτ); In the above formula, Nmin represents the lower limit of the buffer, fs represents the signal sampling rate, and Δτ is the reserved safety margin; Set the buffer limit based on the delay threshold Tmax: Nmax=fs×Tmax; In the above formula, Nmax represents the upper limit of the buffer, and fs represents the signal sampling rate; The buffer interval is obtained as [Nmin, Nmax].

4. The method for canceling echo of vehicle-mounted equipment according to claim 1, characterized in that: The step of calculating the correlation between the microphone signal and the far-end voice signal includes: The correlation between the microphone signal and the far-end voice signal is calculated using the following formula: In the above formula, R xy (k) represents the correlation between the microphone signal and the far-end speech signal, N represents the number of samples of the audio frame selected for processing, y(n) represents the microphone signal, x(n) represents the far-end speech signal, k represents the sampling delay length of the original sound and the echo, and n represents the nth audio sample signal.

5. The method for canceling echo of vehicle-mounted equipment according to claim 4, characterized in that: The step of generating an echo estimation signal according to the far-end speech signal and the estimation of the echo path includes: The echo estimation signal is obtained by the following formula: d^(n)=h(n)*x(n); In the above formula, d^(n) represents the echo estimation signal, h(n) represents the echo estimation path, and x(n) represents the far-end speech signal.

6. The method for canceling echo of vehicle-mounted equipment according to claim 5, characterized in that: The step of removing the echo estimation signal from the microphone signal includes: The echo estimation signal is removed by the following formula: e(n)=y(n)-d^(n); In the above formula, e(n) represents the signal after echo cancellation, y(n) represents the microphone signal, and d^(n) represents the echo estimation signal.

7. The method for canceling echo of vehicle-mounted equipment according to claim 6, characterized in that: The updating of the echo path estimation includes: Update the echo path estimate using an adaptive filter: h(n+1)=h(n)+μ·e(n)·x(n); In the above formula, h(n+1) represents the updated echo path, h(n) represents the echo estimation path, μ represents the step size factor, e(n) represents the signal after echo cancellation, and x(n) represents the far-end speech signal.

8. The method for canceling echo of vehicle-mounted equipment according to claim 6, characterized in that: The post-processing of the microphone signal from which the echo estimation signal is removed comprises: Perform fast Fourier transform on the echo-cancelled signal, calculate the power spectrum, traverse all frequencies within 5kHz, and calculate the energy change rate; Marking frequency points according to the energy change rate, and determining whether howling occurs; An IIR notch filter with Q=8 is designed for each howling frequency to perform filtering.

9. The method for canceling echo of vehicle-mounted equipment according to claim 8, characterized in that: The marking of frequency points according to the energy change rate to determine whether howling occurs includes: The frequency points where the energy change rate is greater than 3.0 and the bandwidth is less than 100 Hz are marked. The marked points are counted continuously. If the frequency points appear for more than or equal to 3 consecutive frames, howling is confirmed.

10. The vehicle-mounted device call echo cancellation method according to claim 8, characterized in that: The IIR notch filter with Q=8 is designed for each howling frequency, and after filtering, it includes: When the energy of five consecutive frames decreases, the IIR notch filter corresponding to the marked point is removed.