Echo cancellation method and server for vehicle cabin

By determining the reference signal and target signal in the vehicle cockpit and using specific signal processing technology to echo cancellation of the speaker signal, the problem of echo interference in the vehicle is solved and the quality of voice interaction and user experience is improved.

CN120126495APending Publication Date: 2025-06-10GUANGZHOU XIAOPENG MOTORS TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510287171.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

Echo interference in the vehicle cockpit affects the voice interaction function, resulting in poor user voice interaction experience.

Method used

By determining the reference signal and the first target signal, the speaker signal is echo cancelled using a weighted recursive least squares filter and a packet timing convolution recursive network.

Benefits of technology

Effectively eliminate in-car echoes, reduce noise interference, improve the quality and accuracy of voice interactions, and improve user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126495A_ABST
    Figure CN120126495A_ABST
Patent Text Reader

Abstract

The invention discloses an echo cancellation method for a vehicle cabin, a server and a computer readable storage medium. The method comprises the following steps: determining a reference signal according to a loudspeaker signal; then, determining a first target signal according to the microphone signal and the reference signal; and finally, performing echo cancellation processing on the loudspeaker signal according to the microphone signal, the first target signal and the reference signal. Thus, echo cancellation processing is performed on the loudspeaker signal through the obtained microphone signal, the determined first target signal and the reference signal, echoes formed in the vehicle are eliminated, noise interference is reduced, the quality of interactive voice is improved, and the voice interaction experience of a user is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of voice interaction, and particularly to an echo cancellation method, a server, and a computer-readable storage medium for a vehicle cockpit. Background Art

[0002] In the related art, a user can control a vehicle through voice interaction. However, since the interior space of the vehicle reflects the sound propagated inside the vehicle and is picked up by the vehicle microphone, an echo is formed, which affects the voice interaction function of the vehicle and results in a poor voice interaction experience for the user. Summary of the Invention

[0003] The present application provides an echo cancellation method, a server, and a computer-readable storage medium for a vehicle cockpit.

[0004] An embodiment of the present application provides an echo cancellation method for a vehicle cockpit, the method comprising:

[0005] Determining a reference signal according to a speaker signal;

[0006] Determining a first target signal according to a microphone signal and the reference signal;

[0007] Performing echo cancellation processing on the speaker signal according to the microphone signal, the first target signal, and the reference signal.

[0008] In this way, the server determines a reference signal according to the speaker signal. Then, the server determines a first target signal according to the microphone signal and the reference signal. Finally, the server performs echo cancellation processing on the speaker signal according to the microphone signal, the first target signal, and the reference signal. In this way, by using the obtained microphone signal, the determined first target signal, and the reference signal, echo cancellation processing is performed on the speaker signal, the echo formed inside the vehicle is eliminated, the noise interference is reduced, the quality of the interactive voice is improved, and the user's voice interaction experience is enhanced.

[0009] In some embodiments, the determining a reference signal according to a speaker signal includes:

[0010] Performing splitting processing on the speaker signal;

[0011] Determining a first type of reference signal based on the spatial position of the speaker in the vehicle cockpit;

[0012] Determining a second type of reference signal based on the function priority of the vehicle-mounted system;

[0013] Determining the reference signal according to the first type of reference signal and the second type of reference signal.

[0014] In this way, the server splits and processes the speaker signals. Then, based on the spatial positions of the speakers in the vehicle cockpit, the server determines the first type of reference signals. Next, based on the functional priorities of the vehicle system, the server determines the second type of reference signals. Finally, the server determines the reference signals according to the first type of reference signals and the second type of reference signals. In this way, by splitting and combining the speaker signals to generate multiple types of reference signals, the propagation path of sound waves in the vehicle interior space can be comprehensively simulated, thereby improving the accuracy and elimination effect of echo cancellation. Moreover, by grouping multiple speakers in the vehicle audio channel system based on the spatial positions of the speakers in the vehicle cockpit and the functional priorities of the vehicle system, there is no need to process each channel separately, greatly reducing the computational complexity.

[0015] In some embodiments, determining the first target signal according to the microphone signal and the reference signal includes:

[0016] Determining a microphone frequency-domain signal according to the microphone signal;

[0017] Determining a reference frequency-domain signal according to the reference signal;

[0018] Determining a residual error signal according to the microphone frequency-domain signal and the reference frequency-domain signal;

[0019] Determining the first target signal according to the residual error signal.

[0020] In this way, the server determines a microphone frequency-domain signal according to the microphone signal. Then, the server determines a reference frequency-domain signal according to the reference signal. Next, the server determines a residual error signal according to the microphone frequency-domain signal and the reference frequency-domain signal. Finally, the server determines the first target signal according to the residual error signal. In this way, the first target signal can be determined, the echo component in the acquired signal can be reduced, the accuracy of speech recognition can be improved, and thus the user voice interaction experience can be enhanced.

[0021] In some embodiments, determining the residual error signal according to the microphone frequency-domain signal and the reference frequency-domain signal includes:

[0022] Determining a current error signal according to the microphone signal, the filter coefficient vector at the previous moment, and the reference frequency-domain signal;

[0023] Determining a gain vector according to a pre-determined forgetting factor, the inverse correlation matrix at the previous moment, and the reference frequency-domain signal;

[0024] Determining the current filter coefficient vector according to the filter coefficient vector at the previous moment, the current error signal, and the gain vector;

[0025] Determine the residual error signal based on the current filter coefficient vector, the microphone frequency domain signal, and the reference frequency domain signal.

[0026] In this way, the server determines the current error signal based on the microphone signal, the filter coefficient vector at the previous moment, and the reference frequency domain signal. Then, the server determines the gain vector based on the pre-determined forgetting factor, the inverse correlation matrix at the previous moment, and the reference frequency domain signal. Next, the server determines the current filter coefficient vector based on the filter coefficient vector at the previous moment, the current error signal, and the gain vector. Finally, the server determines the residual error signal based on the current filter coefficient vector, the microphone frequency domain signal, and the reference frequency domain signal. In this way, it is possible to dynamically adjust the filter coefficients according to the dynamic changes in the in-vehicle acoustic environment, determine the residual error signal, thereby reducing the linear echo component in the acquired signal, adapting to different acoustic environments, and ensuring the echo cancellation effect.

[0027] In some embodiments, the method further includes:

[0028] Update the filter coefficient vector at the previous moment according to the current filter coefficient vector.

[0029] In this way, the server updates the filter coefficient vector at the previous moment according to the current filter coefficient vector. In this way, by continuously updating the filter coefficients, it is possible to adapt to the changes in the in-vehicle acoustic environment, thereby accurately eliminating the echo in the microphone signal, improving the quality of the voice signal, and maintaining a stable echo cancellation effect.

[0030] In some embodiments, the method further includes:

[0031] Determine the current inverse correlation matrix according to the inverse correlation matrix at the previous moment, the gain vector, the forgetting factor, and the reference frequency domain signal;

[0032] Update the inverse correlation matrix at the previous moment according to the current inverse correlation matrix.

[0033] In this way, the server determines the current inverse correlation matrix according to the inverse correlation matrix at the previous moment, the gain vector, the forgetting factor, and the reference frequency domain signal. Then, the server updates the inverse correlation matrix at the previous moment according to the current inverse correlation matrix. In this way, by continuously updating the inverse correlation matrix, it is possible to accurately reflect the correlation between the reference signals, thereby effectively eliminating the echo in the microphone signal, improving the quality of the voice signal, and enhancing the accuracy of speech recognition.

[0034] In some embodiments, the reference signal includes a first type of reference signal and a second type of reference signal. The first type of reference signal includes a plurality of first-type reference sub-signals, and each of the first-type reference sub-signals is determined based on a spatial position of a speaker in the vehicle cockpit. The echo cancellation processing of the speaker signal according to the microphone signal, the first target signal, and the reference signal includes:

[0035] Performing a superposition process on the plurality of first-type reference sub-signals and the second type of reference signal to determine a target reference signal;

[0036] Performing echo cancellation processing on the speaker signal according to the target reference signal, the microphone signal, and the first target signal.

[0037] In this way, the server performs a superposition process on the plurality of first-type reference sub-signals and the second type of reference signal to determine a target reference signal. Then, the server performs echo cancellation processing on the speaker signal according to the target reference signal, the microphone signal, and the first target signal. In this way, by classifying the reference signal into two types and performing a superposition process, the echo in the in-vehicle signal can be accurately removed, thereby effectively eliminating the echo.

[0038] In some embodiments, the echo cancellation processing of the speaker signal according to the target reference signal, the microphone signal, and the first target signal includes:

[0039] Performing feature extraction processing on the target reference signal, the microphone signal, and the first target signal to determine a target joint feature;

[0040] Performing encoding processing on the target joint feature to determine a high-dimensional embedding feature;

[0041] Performing a temporal linking process on the high-dimensional embedding feature to determine a temporal enhancement feature;

[0042] Performing decoding processing on the high-dimensional embedding feature and the temporal enhancement feature to determine a target reconstructed spectrum;

[0043] Performing mapping processing on the target reconstructed spectrum to implement echo cancellation processing of the speaker signal.

[0044] In this way, the server performs feature extraction processing on the target reference signal, the microphone signal, and the first target signal to determine the target joint features. Then, the server performs encoding processing on the target joint features to determine the high-dimensional embedded features. Next, the server performs temporal linking processing on the high-dimensional embedded features to determine the temporal enhanced features. Subsequently, the server performs decoding processing on the high-dimensional embedded features and the temporal enhanced features to determine the target reconstructed spectrum. Finally, mapping processing is performed on the target reconstructed spectrum to achieve echo cancellation processing of the speaker signal. In this way, by capturing short-term features and modeling long-term dependencies, non-linear echo modeling and noise reduction are performed on in-vehicle signals, thereby balancing echo cancellation and speech retention, reducing speech distortion, increasing speech quality, and improving the speech interaction experience.

[0045] An embodiment of the present application provides a server, which includes a processor and a memory. A computer program is stored on the memory. When the computer program is executed by the processor, the above method is implemented.

[0046] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above method are implemented.

[0047] Additional aspects and advantages of the embodiments of the present application will be given in part in the following description, will become apparent in part from the following description, or will be understood through the practice of the embodiments of the present application. Description of the Drawings

[0048] The above and / or additional aspects and advantages of the present application will become apparent and easy to understand from the description of the embodiments in conjunction with the following drawings, where:

[0049] Figure 1 is one of the flow diagrams of the echo cancellation method of some embodiments of the present application;

[0050] Figure 2 is another flow diagram of the echo cancellation method of some embodiments of the present application;

[0051] Figure 3 is yet another flow diagram of the echo cancellation method of some embodiments of the present application;

[0052] Figure 4 is still another flow diagram of the echo cancellation method of some embodiments of the present application;

[0053] Figure 5 is yet still another flow diagram of the echo cancellation method of some embodiments of the present application;

[0054] Figure 6 is another flow diagram of the echo cancellation method of some embodiments of the present application;

[0055] Figure 7 It is the seventh flow schematic diagram of the echo cancellation method of some embodiments of the present application;

[0056] Figure 8 It is the eighth flow schematic diagram of the echo cancellation method of some embodiments of the present application;

[0057] Figure 9 It is the processing flow schematic diagram of the input signal of the grouped temporal convolutional recurrent network. Detailed implementation manners

[0058] The following details the implementation manners of the present application. Examples of the implementation manners are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The implementation manners described below with reference to the accompanying drawings are exemplary and are only used to explain the implementation manners of the present application, and should not be construed as a limitation to the implementation manners of the present application.

[0059] In the related art of in-vehicle dialogue systems, users can control the vehicle through voice interaction, which can reduce the driver's operation burden and improve driving safety during driving. However, since the interior space of the vehicle reflects the sound propagated inside the vehicle and is picked up by the vehicle microphone, an echo is formed, which will affect the voice interaction function of the vehicle.

[0060] Specifically, the interior space of the vehicle is relatively enclosed, with multiple sound sources and reflecting surfaces, resulting in a complex sound wave propagation path, forming multiple reflections and interferences, and making the interior acoustic environment complex. When the speaker plays a sound, part of the sound will be picked up by the microphone after being reflected by the interior space of the vehicle, forming an echo. The echo is mixed with the original voice signal, resulting in the distortion of the voice signal and affecting the accuracy of voice recognition. In addition, there are various noise sources in the vehicle, such as engine noise, wind noise, and tire noise. These noises are superimposed on the echo, further reducing the signal-to-noise ratio of the voice signal and making voice recognition more difficult.

[0061] Therefore, the interference of in-vehicle echo and noise will increase the voice recognition error rate, affect the fluency and accuracy of voice interaction, and thus reduce the user's voice interaction experience.

[0062] Based on the above problems, please refer to Figure 1 , the implementation manners of the present application provide an echo cancellation method for a vehicle cockpit, and the method includes:

[0063] 01: Determine a reference signal according to the speaker signal;

[0064] 02: Determine a first target signal according to the microphone signal and the reference signal;

[0065] 03: Perform echo cancellation processing on the speaker signal according to the microphone signal, the first target signal, and the reference signal.

[0066] The embodiments of the present application further provide a server, including a memory and a processor. The training method of the speech recognition model in the embodiments of the present application can be implemented by the server in the embodiments of the present application. Specifically, a computer program is stored in the memory, and the processor is configured to determine a reference signal according to the speaker signal, and determine a first target signal according to the microphone signal and the reference signal. The processor is further configured to perform echo cancellation processing on the speaker signal according to the microphone signal, the first target signal, and the reference signal.

[0067] The embodiments of the present application further provide an echo cancellation device. The training method of the speech recognition model in the embodiments of the present application can be implemented by the echo cancellation device in the embodiments of the present application. Specifically, the echo cancellation device includes a determination module and an echo cancellation module. The determination module is configured to determine a reference signal according to the speaker signal, and is further configured to determine a first target signal according to the microphone signal and the reference signal. The echo cancellation module is configured to perform echo cancellation processing on the speaker signal according to the microphone signal, the first target signal, and the reference signal.

[0068] Specifically, the echo cancellation method provided in the embodiments of the present application is directed to the acoustic environment of the vehicle cockpit. By eliminating the echo formed by the sound played by the in-vehicle speaker being picked up by the microphone after reflection in the vehicle space, the clarity and accuracy of the voice signal are improved, thereby improving the accuracy of speech recognition and enhancing the user's voice interaction experience.

[0069] The speaker signal refers to the sound signal that drives the speakers in the in-vehicle audio system, including navigation announcements, music, voice responses, etc. In the embodiments of the present application, taking the in-vehicle 7.1.4 channel system as an example, the echo cancellation method provided in the embodiments of the present application is described. The in-vehicle 7.1.4 channel system refers to a complete audio playback system composed of 7 main audio channels, 1 subwoofer channel, and 4 overhead channels in the in-vehicle audio system. The 7 main audio channels include the front left channel, the front right channel, the rear left channel, the rear right channel, the left center channel, the right center channel, and the center channel, which are used to play the main audio signals, such as human voices, musical instrument sounds, etc. The subwoofer channel is also called the heavy bass channel and is responsible for playing low-frequency audio signals, such as drum sounds, double bass sounds, etc. The overhead channels are located on the roof and are responsible for playing audio signals from the overhead direction, such as the sound of an airplane flying by, raindrops falling on the roof, etc.

[0070] The reference signal is obtained by processing the sound signal played by the speaker and is used to predict and cancel the echo collected by the microphone.

[0071] The microphone signal refers to the sound signals collected by the microphone from multiple sound sources, including the original speech signal, echo, and ambient noise, which is the basis for echo cancellation and noise reduction processing. The original speech signal refers to the speech request issued by the user. The echo refers to the sound that is picked up by the microphone again after the sound played by the speaker is reflected. The ambient noise refers to various noises inside and outside the vehicle, such as engine noise, road noise, and wind noise, etc.

[0072] The first target signal refers to the signal that after linear echo cancellation processing, the echo component in the microphone signal is suppressed and is closer to the original speech signal, which can be used as the input signal for subsequent processing to improve the clarity and quality of voice interaction.

[0073] The server determines the reference signal according to the collected speaker signal.

[0074] Next, the server determines the first target signal according to the microphone signal and the determined reference signal above. It should be noted that here, a weighted recursive least squares (WRLS) filter is used to perform linear echo cancellation on the microphone signal and the reference signal to obtain the signal after linear echo cancellation processing, that is, the first target signal.

[0075] Finally, the server performs echo cancellation processing on the speaker signal according to the microphone signal, the first target signal, and the reference signal. Specifically, a grouped temporal convolutional recurrent network (GTCRN) is used to jointly process the first target signal, the original microphone signal, and the reference signal, and simultaneously complete the tasks of non-linear echo cancellation and noise reduction, and finally obtain a high-quality voice signal.

[0076] In summary, in the echo cancellation method and server for a vehicle cockpit provided by the embodiments of the present application, the server determines the target timestamp information associated with the target training data according to the obtained target training data, where the target training data includes target audio data and target text data corresponding to the target audio data. Next, the server determines the target masked training data according to the target training data and the target timestamp information. Finally, the server trains the speech recognition model according to the target masked training data. In this way, through the timestamp-driven collaborative masking technology, the target training data is actively processed during the training stage of the speech recognition model to determine the target masked training data, which can enable the speech recognition model to learn deep context semantics and anti-interference features, reduce the mis-touch rate, and improve the user experience. And while reducing the mis-touch rate, the data diversity is retained and the model generalization ability is improved.

[0077] Please refer toFigure 2 , in some embodiments, step 01 (determining a reference signal according to the speaker signal) includes:

[0078] 011: performing splitting processing on the speaker signal;

[0079] 012: determining a first type of reference signal based on the spatial position of the speaker in the vehicle cockpit;

[0080] 013: determining a second type of reference signal based on the function priority of the vehicle-mounted system;

[0081] 014: determining the reference signal according to the first type of reference signal and the second type of reference signal.

[0082] In some embodiments, the echo cancellation device further includes a splitting module, and the splitting module is used to perform splitting processing on the speaker signal. The determining module is further used to determine a first type of reference signal based on the spatial position of the speaker in the vehicle cockpit. And determine a second type of reference signal based on the function priority of the vehicle-mounted system. And determine the reference signal according to the first type of reference signal and the second type of reference signal.

[0083] In some embodiments, the processor is further used to perform splitting processing on the speaker signal. And determine a first type of reference signal based on the spatial position of the speaker in the vehicle cockpit. The processor is further used to determine a second type of reference signal based on the function priority of the vehicle-mounted system. And determine the reference signal according to the first type of reference signal and the second type of reference signal.

[0084] Specifically, the splitting processing refers to decomposing the integrated-channel speaker signal into multiple independent channel signals according to specific rules and methods.

[0085] First, the server performs splitting processing on the speaker signal to determine multiple independent channel signals.

[0086] Next, based on the spatial position of the speaker in the vehicle cockpit, the server determines a first type of reference signal. That is, according to the position of the speaker in the vehicle cockpit, the above-determined independent channel signals are combined. For example, the signals of the left speakers are mixed into one path, and the signals of the right speakers are mixed into another path. Then the left speaker signal and the right speaker signal are collectively referred to as the first type of reference signal. It should be noted that in the embodiments of the present application, for the in-vehicle 7.1.4 channel system, all the speaker signals on the left in the vehicle cockpit are mixed into reference signal 1, all the speaker signals on the right in the vehicle cockpit are mixed into reference signal 2, and the sky channel signal is separately extracted as reference signal 3.

[0087] Then, based on the functional priorities of the in-vehicle system, the server determines the second type of reference signal. That is, according to the functional priorities of the in-vehicle system, the speaker signals of some specific functions are selected as the second type of reference signal, which can accurately eliminate the echo generated by specific functions and avoid interfering with voice interaction. In some embodiments, the voice reply signal or the navigation broadcast signal during voice interaction is separately extracted as the reference signal 4.

[0088] Finally, the server determines the reference signal for echo cancellation and noise reduction based on the first type of reference signal and the second type of reference signal.

[0089] In this way, based on the spatial positions of the speakers in the vehicle cockpit and the functional priorities of the in-vehicle system, multiple speakers in the in-vehicle 7.1.4 channel system are grouped, and there is no need to process each channel separately, which greatly reduces the computational complexity. The in-vehicle 7.1.4 channel system includes more than 12 speakers. If each channel is processed separately, the computational complexity is 12×12 = 144 groups of paths. However, when multiple speakers in the in-vehicle 7.1.4 channel system are grouped based on the spatial positions of the speakers in the vehicle cockpit and the functional priorities of the in-vehicle system, the reference signal is compressed to 4 paths, and the computational complexity is reduced to 4×4 = 16 groups of paths.

[0090] In this way, the server splits and processes the speaker signals. Then, based on the spatial positions of the speakers in the vehicle cockpit, the server determines the first type of reference signal. Then, based on the functional priorities of the in-vehicle system, the server determines the second type of reference signal. Finally, the server determines the reference signal based on the first type of reference signal and the second type of reference signal. In this way, by splitting and combining the speaker signals, multiple types of reference signals are generated, which can comprehensively simulate the propagation path of sound waves in the vehicle interior space, thereby improving the accuracy and elimination effect of echo cancellation. And, based on the spatial positions of the speakers in the vehicle cockpit and the functional priorities of the in-vehicle system, multiple speakers in the in-vehicle channel system are grouped, and there is no need to process each channel separately, which greatly reduces the computational complexity.

[0091] Please refer to Figure 3 , in some embodiments, step 02 (determining the first target signal according to the microphone signal and the reference signal) includes:

[0092] 021: Determine the microphone frequency domain signal according to the microphone signal;

[0093] 022: Determine the reference frequency domain signal according to the reference signal;

[0094] 023: Determine the residual error signal according to the microphone frequency domain signal and the reference frequency domain signal;

[0095] 024: Determine the first target signal according to the residual error signal.

[0096] In some embodiments, the determination module is further configured to determine a microphone frequency-domain signal based on the microphone signal, and determine a reference frequency-domain signal based on the reference signal. The determination module is further configured to determine a residual error signal based on the microphone frequency-domain signal and the reference frequency-domain signal, and determine a first target signal based on the residual error signal.

[0097] In some embodiments, the processor is further configured to determine a microphone frequency-domain signal based on the microphone signal, and determine a reference frequency-domain signal based on the reference signal. The processor is further configured to determine a residual error signal based on the microphone frequency-domain signal and the reference frequency-domain signal, and determine a first target signal based on the residual error signal.

[0098] Specifically, the microphone frequency-domain signal refers to the sound signal picked up by the microphone, which is obtained after being converted from the time domain to the frequency domain through the fast Fourier transform (FFT).

[0099] The fast Fourier transform algorithm (FFT) is an algorithm for calculating an approximation of the discrete Fourier transform (DFT). Based on the properties of the DFT and the idea of recursive decomposition, it can decompose the time-domain signal into a combination of sine waves and cosine waves of different frequencies, thereby converting the signal from the time domain to the frequency domain, revealing the frequency characteristics of the signal, and performing various frequency-domain operations. It decomposes the DFT into multiple smaller DFTs and then gradually combines these results to obtain the final DFT. The DFT can convert a signal from the time domain to the frequency domain and reveal the frequency composition of the signal.

[0100] The reference frequency-domain signal refers to the signal obtained by splitting and combining the in-vehicle speaker signal and then converting it to the frequency domain through the fast Fourier transform. In some embodiments, if all the speaker signals on the left side in the vehicle cockpit are mixed as reference signal 1, all the speaker signals on the right side in the vehicle cockpit are mixed as reference signal 2, the sky channel signal is separately extracted as reference signal 3, and the voice reply signal or the navigation broadcast signal during voice interaction is separately extracted as reference signal 4. Then, the reference frequency-domain signal will form a four-dimensional matrix from these four groups of reference signals through the fast Fourier transform.

[0101] The residual error signal refers to the difference signal between the microphone frequency-domain signal and the reference frequency-domain signal after being processed by a weighted recursive least squares filter, including the uneliminated echo components, as well as possible noise and errors in the signal processing process. Thus, the residual error signal needs to be further processed by a non-linear processing module to achieve more thorough echo cancellation and noise reduction.

[0102] The server performs fast Fourier transform processing on the microphone signal, converts the time-domain signal to the frequency domain, and determines the microphone frequency-domain signal Y(n). Next, the server performs fast Fourier transform processing on the reference signal, converts the time-domain signal to the frequency domain, and determines the reference frequency-domain signal x(n). Then, the server determines the residual error signal based on the microphone frequency-domain signal Y(n) and the reference frequency-domain signal x(n). Finally, the server performs inverse fast Fourier transform processing on the residual error signal to determine the first target signal.

[0103] Inverse fast Fourier transform processing refers to the inverse operation of the fast Fourier transform (FFT), which converts a frequency-domain signal back to a time-domain signal. The process of inverse fast Fourier transform processing is similar to that of fast Fourier transform processing, but the steps are reversed. Inverse fast Fourier transform processing uses the fast algorithm of the discrete Fourier transform (DFT), while inverse fast Fourier transform processing uses the inverse operation of the fast algorithm of the discrete Fourier transform.

[0104] Thus, the server determines the microphone frequency-domain signal based on the microphone signal. Next, the server determines the reference frequency-domain signal based on the reference signal. Then, the server determines the residual error signal based on the microphone frequency-domain signal and the reference frequency-domain signal. Finally, the server determines the first target signal based on the residual error signal. In this way, the first target signal can be determined, the echo component in the acquired signal can be reduced, the accuracy of speech recognition can be improved, and thus the user's speech interaction experience can be enhanced.

[0105] Please refer to Figure 4 , in some embodiments, step 023 (determining the residual error signal based on the microphone frequency-domain signal and the reference frequency-domain signal) includes:

[0106] 0231: Determine the current error signal based on the microphone signal, the filter coefficient vector at the previous moment, and the reference frequency-domain signal;

[0107] 0232: Determine the gain vector based on the pre-determined forgetting factor, the inverse correlation matrix at the previous moment, and the reference frequency-domain signal;

[0108] 0233: Determine the current filter coefficient vector based on the filter coefficient vector at the previous moment, the current error signal, and the gain vector;

[0109] 0234: Determine the residual error signal based on the current filter coefficient vector, the microphone frequency-domain signal, and the reference frequency-domain signal.

[0110] In some embodiments, the determination module is further configured to determine a current error signal according to the microphone signal, the filter coefficient vector at the previous moment, and the reference frequency-domain signal. And determine a gain vector according to a predetermined forgetting factor, the inverse correlation matrix at the previous moment, and the reference frequency-domain signal. The determination module is further configured to determine the current filter coefficient vector according to the filter coefficient vector at the previous moment, the current error signal, and the gain vector. And determine a residual error signal according to the current filter coefficient vector, the microphone frequency-domain signal, and the reference frequency-domain signal.

[0111] In some embodiments, the processor is further configured to determine a current error signal according to the microphone signal, the filter coefficient vector at the previous moment, and the reference frequency-domain signal. And determine a gain vector according to a predetermined forgetting factor, the inverse correlation matrix at the previous moment, and the reference frequency-domain signal. The processor is further configured to determine the current filter coefficient vector according to the filter coefficient vector at the previous moment, the current error signal, and the gain vector. And determine a residual error signal according to the current filter coefficient vector, the microphone frequency-domain signal, and the reference frequency-domain signal.

[0112] Specifically, in the embodiments of the present application, based on a weighted recursive least squares filter, a residual error signal is determined according to the microphone frequency-domain signal and the reference frequency-domain signal. The weighted recursive least squares algorithm is a signal processing algorithm used to estimate unknown parameters or models in a signal, belonging to the category of adaptive filters, which can dynamically adjust its own parameters according to the statistical characteristics of the input signal, so as to achieve a better filtering effect.

[0113] The goal of the weighted recursive least squares algorithm is to find a set of filter coefficient vectors that minimize the error between the output of the filter and the desired signal. The implementation method is to recursively update the filter coefficient vector, that is, using the weighted least squares criterion, assigning different weights to data at different times to update the filter coefficient vector, so as to better adapt to the dynamic changes of the signal. Moreover, the weighted recursive least squares filter also introduces a forgetting factor to control the weight of historical data. A smaller forgetting factor means more emphasis on recent data, while a larger forgetting factor means more emphasis on historical data. Compared with the traditional LMS algorithm, the weighted recursive least squares filter has a faster convergence speed and can adapt to the dynamic changes of the signal faster. And the weighted recursive least squares filter has strong adaptability to non-stationary environments and stronger robustness to noise and abnormal signals.

[0114] The filter coefficient vector refers to a set of coefficient vectors used to control how the filter processes the input signal. By adjusting these coefficient vectors, the weighted recursive least squares filter can eliminate or enhance signal components at specific frequencies, thereby achieving echo cancellation.

[0115] The forgetting factor refers to a parameter between 0 and 1, which is used to control the influence degree of historical data on the update of filter coefficients. When λ = 1, the update of the filter coefficient vector only considers the error at the current moment and ignores historical data. When 0 < λ < 1, the update of the filter coefficient vector considers both the current moment and historical data, and the influence degree of historical data is determined by λ. The smaller λ is, the greater the influence degree of historical data is, the slower the filter convergence speed is, but the better the robustness is.

[0116] The inverse correlation matrix refers to the matrix used to describe the correlation between input signals. By calculating the correlation between input signals, the filter can estimate the error at the current moment, so as to accurately update the filter coefficient vector.

[0117] The server determines the current error signal according to the microphone signal, the filter coefficient vector at the previous moment and the reference frequency domain signal. The calculation formula of the current error signal is e(n) = d(n) - w(n - 1)x(n), where e(n) is the current error signal, d(n) is the desired signal, also known as the target signal, which is the microphone signal and includes echo and near-end speech, w(n - 1) is the filter coefficient vector at the previous moment, and x(n) is the input signal vector at the current moment, which is the reference frequency domain signal.

[0118] Next, the server determines the gain vector according to the pre-determined forgetting factor, the inverse correlation matrix at the previous moment and the reference frequency domain signal. The calculation formula of the gain vector is where k(n) is the gain vector, which represents the weighting degree of the filter for the input signal and is used to adjust the filter coefficients. λ is the forgetting factor, which controls the weight of historical data, and its value range is 0 < λ ≤ 1. P(n - 1) is the inverse correlation matrix at the previous moment, which represents the correlation between input signals. x(n) is the input signal vector at the current moment, which is the reference frequency domain signal. x T (n) is the transpose matrix of x(n).

[0119] Then, the server determines the current filter coefficient vector according to the filter coefficient vector at the previous moment, the current error signal and the gain vector. The calculation formula of the current filter coefficient vector is P(n) = λ -1 P(n - 1) - λ -1 k(n)x T (n)P(n - 1), where w(n - 1) is the filter coefficient vector at the previous moment, k(n) is the gain vector, which represents the weighting degree of the filter for the input signal and is used to adjust the filter coefficients, and e(n) is the current error signal.

[0120] Finally, the server determines the residual error signal based on the current filter coefficient vector, the microphone frequency-domain signal, and the reference frequency-domain signal. The calculation formula for the residual error signal is E(k) = Y(n) - w(n) * x(n), where w(n) is the current filter coefficient vector, and x(n) is the input signal vector at the current moment, which is the reference frequency-domain signal.

[0121] In this way, the server determines the current error signal based on the microphone signal, the filter coefficient vector at the previous moment, and the reference frequency-domain signal. Then, the server determines the gain vector based on the pre-determined forgetting factor, the inverse correlation matrix at the previous moment, and the reference frequency-domain signal. Next, the server determines the current filter coefficient vector based on the filter coefficient vector at the previous moment, the current error signal, and the gain vector. Finally, the server determines the residual error signal based on the current filter coefficient vector, the microphone frequency-domain signal, and the reference frequency-domain signal. In this way, it is possible to dynamically adjust the filter coefficients according to the dynamic changes in the in-vehicle acoustic environment, determine the residual error signal, thereby reducing the linear echo component in the acquired signal, adapting to different acoustic environments, and ensuring the echo cancellation effect.

[0122] Please refer to Figure 5 , in some embodiments, the method further includes:

[0123] 0235: Update the filter coefficient vector at the previous moment according to the current filter coefficient vector.

[0124] In some embodiments, the echo cancellation device further includes an update module, and the update module is further configured to update the filter coefficient vector at the previous moment according to the current filter coefficient vector.

[0125] In some embodiments, the processor is further configured to update the filter coefficient vector at the previous moment according to the current filter coefficient vector.

[0126] Specifically, the current filter coefficient vector is used to update the filter coefficient vector at the previous moment to make it closer to the optimal filter coefficient at the current moment, so as to be used as a parameter for subsequent echo cancellation processing.

[0127] In this way, the server updates the filter coefficient vector at the previous moment according to the current filter coefficient vector. In this way, by continuously updating the filter coefficients, it is possible to adapt to the changes in the in-vehicle acoustic environment, thereby accurately canceling the echo in the microphone signal, improving the quality of the voice signal, and maintaining a stable echo cancellation effect.

[0128] Please refer to Figure 6 , in some embodiments, the method further includes:

[0129] 0236: Determine the current inverse correlation matrix according to the inverse correlation matrix at the previous moment, the gain vector, the forgetting factor, and the reference frequency domain signal;

[0130] 0237: Update the inverse correlation matrix at the previous moment according to the current inverse correlation matrix.

[0131] In some embodiments, the determination module is further configured to determine the current inverse correlation matrix according to the inverse correlation matrix at the previous moment, the gain vector, the forgetting factor, and the reference frequency domain signal. The update module is further configured to update the inverse correlation matrix at the previous moment according to the current inverse correlation matrix.

[0132] In some embodiments, the processor is further configured to determine the current inverse correlation matrix according to the inverse correlation matrix at the previous moment, the gain vector, the forgetting factor, and the reference frequency domain signal. And update the inverse correlation matrix at the previous moment according to the current inverse correlation matrix.

[0133] Specifically, the server determines the current inverse correlation matrix according to the inverse correlation matrix at the previous moment, the gain vector, the forgetting factor, and the reference frequency domain signal. The calculation formula is P(n) = λ -1 P(n - 1) - λ -1 k(n)x T (n)P(n - 1), where P(n - 1) is the inverse correlation matrix at the previous moment. λ is the forgetting factor, which controls the weight of historical data. k(n) is the gain vector, which represents the weighting degree of the filter for the input signal and is used to adjust the filter coefficients. x(n) is the input signal vector at the current moment, that is, the reference frequency domain signal. x T (n) is the transpose matrix of x(n).

[0134] Next, the server updates the inverse correlation matrix at the previous moment according to the current inverse correlation matrix. That is, the current inverse correlation matrix is used to update the inverse correlation matrix at the previous moment to make it closer to the optimal inverse correlation matrix at the current moment. Updating the inverse correlation matrix is an important step in the wRLS filter, which ensures that the filter can adapt to the dynamically changing in-vehicle acoustic environment and maintain the best performance, thereby effectively eliminating echo and improving the quality of the voice signal.

[0135] In this way, the server determines the current inverse correlation matrix according to the inverse correlation matrix at the previous moment, the gain vector, the forgetting factor, and the reference frequency domain signal. Next, the server updates the inverse correlation matrix at the previous moment according to the current inverse correlation matrix. In this way, by continuously updating the inverse correlation matrix, the correlation between reference signals can be accurately reflected, thereby effectively eliminating echo in the microphone signal, improving the quality of the voice signal, and enhancing the accuracy of speech recognition.

[0136] Please refer to Figure 7, in some embodiments, the reference signal includes a first type of reference signal and a second type of reference signal. The first type of reference signal includes a plurality of first-type reference sub-signals, and each first-type reference sub-signal is determined based on a spatial position of a speaker in a vehicle cockpit. Step 03 (performing echo cancellation processing on the speaker signal according to the microphone signal, the first target signal, and the reference signal) includes:

[0137] 031: Performing a superposition process on the plurality of first-type reference sub-signals and the second type of reference signal to determine a target reference signal;

[0138] 032: Performing echo cancellation processing on the speaker signal according to the target reference signal, the microphone signal, and the first target signal.

[0139] In some embodiments, the determination module is further configured to perform a superposition process on the plurality of first-type reference sub-signals and the second type of reference signal to determine a target reference signal. The echo cancellation module is configured to perform echo cancellation processing on the speaker signal according to the target reference signal, the microphone signal, and the first target signal.

[0140] In some embodiments, the processor is further configured to perform a superposition process on the plurality of first-type reference sub-signals and the second type of reference signal to determine a target reference signal. And perform echo cancellation processing on the speaker signal according to the target reference signal, the microphone signal, and the first target signal.

[0141] Specifically, the first type of reference signal includes a plurality of first-type reference sub-signals, and each first-type reference sub-signal is determined based on a spatial position of a speaker in a vehicle cockpit. For example, the signals of the left speaker are mixed into one path, and the signals of the right speaker are mixed into another path. Then the left speaker signal and the right speaker signal are collectively referred to as the first type of reference signal. The left speaker signal is the left first-type reference sub-signal, and the right speaker signal is the right first-type reference sub-signal. Also, for example, for an in-vehicle 7.1.4 channel system, all the speaker signals on the left side in the vehicle cockpit are mixed into reference signal 1, all the speaker signals on the right side in the vehicle cockpit are mixed into reference signal 2, and the sky channel signal is separately extracted as reference signal 3. Reference signal 1, reference signal 2, and reference signal 3 are all first-type reference sub-signals.

[0142] The superposition process refers to performing a superposition process on the plurality of first-type reference sub-signals and the second type of reference signal to generate a single-path signal, that is, the target reference signal. The target reference signal generated after the superposition process is used to restore the original channel speaker signal and is used as the input of the non-linear processing module for further echo cancellation and noise reduction processing.

[0143] The server performs superposition processing on multiple first - type reference sub - signals and second - type reference signals to determine a target reference signal. Then, the server performs echo cancellation processing on the speaker signal based on the target reference signal, the microphone signal, and the first target signal.

[0144] In this way, the server performs superposition processing on multiple first - type reference sub - signals and second - type reference signals to determine a target reference signal. Then, the server performs echo cancellation processing on the speaker signal based on the target reference signal, the microphone signal, and the first target signal. Thus, by classifying the reference signals into two types and performing superposition processing, the echo in the in - vehicle signal can be accurately identified, thereby effectively eliminating the echo.

[0145] Please refer to Figure 8 , in some embodiments, step 032 (performing echo cancellation processing on the speaker signal based on the target reference signal, the microphone signal, and the first target signal) includes:

[0146] 0321: Perform feature extraction processing on the target reference signal, the microphone signal, and the first target signal to determine a target joint feature;

[0147] 0322: Perform encoding processing on the target joint feature to determine a high - dimensional embedded feature;

[0148] 0323: Perform temporal linking processing on the high - dimensional embedded feature to determine a temporal enhancement feature;

[0149] 0324: Perform decoding processing on the high - dimensional embedded feature and the temporal enhancement feature to determine a target reconstructed spectrum;

[0150] 0325: Perform mapping processing on the target reconstructed spectrum to achieve echo cancellation processing on the speaker signal.

[0151] In some embodiments, the determination module is used to perform feature extraction processing on the target reference signal, the microphone signal, and the first target signal to determine a target joint feature. And perform encoding processing on the target joint feature to determine a high - dimensional embedded feature. The determination module is used to perform temporal linking processing on the high - dimensional embedded feature to determine a temporal enhancement feature. And perform decoding processing on the high - dimensional embedded feature and the temporal enhancement feature to determine a target reconstructed spectrum. The echo cancellation module is used to perform mapping processing on the target reconstructed spectrum to achieve echo cancellation processing on the speaker signal.

[0152] In some embodiments, the processor is further configured to perform feature extraction processing on the target reference signal, the microphone signal, and the first target signal to determine the target joint feature. And perform encoding processing on the target joint feature to determine the high-dimensional embedding feature. And perform temporal linking processing on the high-dimensional embedding feature to determine the temporal enhancement feature. The processor is further configured to perform decoding processing on the high-dimensional embedding feature and the temporal enhancement feature to determine the target reconstructed spectrum. And perform mapping processing on the target reconstructed spectrum to implement echo cancellation processing of the speaker signal.

[0153] Specifically, performing echo cancellation processing on the speaker signal according to the target reference signal, the microphone signal, and the first target signal refers to performing non-linear echo cancellation processing and noise reduction processing on the input signal based on a grouped temporal convolutional recurrent network. That is, based on the grouped temporal convolutional recurrent network, echo cancellation processing is performed on the collected speaker signal according to the target reference signal, the microphone signal, and the first target signal.

[0154] The grouped temporal convolutional recurrent network is a deep learning model that combines the advantages of a temporal convolutional neural network (TCN) and a recurrent neural network (RNN), and introduces a grouped processing mechanism to reduce the computational complexity. The temporal convolutional neural network can capture short-term features in time series data, extract features through convolutional layers, and reduce the feature dimension through pooling layers. The recurrent neural network can model long-term dependencies in time series data, remember information in the sequence through recurrent units, and pass it to subsequent time steps.

[0155] Please refer to Figure 9 , Figure 9 FIG. [FIGURE NUMBER] is a schematic diagram of the processing flow of the grouped temporal convolutional recurrent network for the input signal. The grouped temporal convolutional recurrent network includes functional modules such as a feature extraction module, an encoder, a temporal linking module, a decoder, and an output layer. Among them, the feature extraction module is configured to perform feature extraction processing on the target reference signal, the microphone signal, and the first target signal to determine the target joint feature. The encoder is configured to perform encoding processing on the target joint feature to determine the high-dimensional embedding feature. The temporal linking module performs temporal linking processing on the high-dimensional embedding feature to determine the temporal enhancement feature. The decoder is configured to perform decoding processing on the high-dimensional embedding feature and the temporal enhancement feature to determine the target reconstructed spectrum. The output layer is configured to perform mapping processing on the target reconstructed spectrum to implement echo cancellation processing of the speaker signal.

[0156] Specifically, the feature extraction module can extract features from the target reference signal, the microphone signal, and the first target signal through a temporal convolutional layer and a pooling layer to obtain target joint features, including spectral features, time-domain features, etc., for describing the echo components and speech components in the signal. In some embodiments, the implementation process is as follows: First, group the target reference signal, the microphone signal, and the first target signal by time frames. Then, perform a convolution operation on each time frame group to extract features. Finally, perform a pooling operation on the convolved features to reduce the feature dimension.

[0157] The encoder maps the target joint features to a high-dimensional embedding space and downsamples the frequency axis size to determine the high-dimensional embedding features. The encoder consists of two convolutional (Conv) and three grouped temporal convolutional (GT-Conv) blocks, which extract features through convolution operations and reduce the feature dimension through pooling operations. Among them, each Conv block includes a convolutional layer, a batch normalization, and a PReLU activation. Batch normalization refers to a commonly used regularization technique in deep learning, which is used to accelerate the training of neural networks and improve the model performance. It standardizes the features of each mini-batch input data, making the mean of each feature 0 and the variance 1, thereby reducing the internal covariate shift and accelerating the model convergence. PReLU is a learnable activation function, which is an extended version of the ReLU activation function. Compared with ReLU, PReLU can automatically adjust the activation threshold according to the input data, so as to better adapt to different data distributions and improve the model performance. In some embodiments, the implementation process is as follows: First, the target joint features pass through multiple convolutional layers and GT-Conv layers to extract features and reduce the feature dimension. Then, the GT-Conv layer groups the target joint features and performs a convolution operation to reduce the computational complexity.

[0158] The temporal connection module consists of a grouped recurrent neural network (GRNN) and a dual path recurrent neural network (DPRNN), which memorize the information in the sequence through recurrent units and transmit it to subsequent time steps to determine the temporal enhancement features. Among them, GRNN uses a set of smaller recurrent layers to approximate a large standard recurrent layer to reduce the computational complexity. DPRNN connects the recurrent units into a dual path structure to further enhance the modeling ability of the model.

[0159] The grouped recurrent neural network is a variant of the recurrent neural network (RNN). By dividing the neurons in the RNN layer into multiple groups and only connecting within the groups, the model complexity is reduced and the training efficiency is improved.

[0160] The dual-path recurrent neural network is a network structure that combines a temporal convolutional network (TCN) and a recurrent neural network (RNN). The DPRNN processes the input data by splitting it into two paths. One path uses the TCN to extract local features, and the other path uses the RNN to extract global features. The outputs of the two paths are fused to improve the model performance.

[0161] The decoder is a mirror version of the encoder. Each Conv block is replaced by a DeConv block, which has the same components as the Conv block, replacing the convolutional layer with a transposed convolutional layer to restore the original size. A skip connection is adopted between the encoder and the decoder. The decoder decodes the high-dimensional embedded features and the temporal enhancement features through deconvolution operations, restores the feature dimension, and performs feature fusion through the GT-Conv layer to determine the target reconstructed spectrum.

[0162] The output layer uses a fully connected layer and an activation function to fuse the features of each group of signals, generate enhanced signal features, and map the enhanced signal features to a time-domain signal. That is, the target reconstructed spectrum is mapped to achieve echo cancellation processing of the speaker signal.

[0163] In this way, the server performs feature extraction processing on the target reference signal, the microphone signal, and the first target signal to determine the target joint features. Then, the server performs encoding processing on the target joint features to determine the high-dimensional embedded features. Next, the server performs temporal linking processing on the high-dimensional embedded features to determine the temporal enhancement features. Subsequently, the server performs decoding processing on the high-dimensional embedded features and the temporal enhancement features to determine the target reconstructed spectrum. Finally, the target reconstructed spectrum is mapped to achieve echo cancellation processing of the speaker signal. In this way, by capturing short-term features and modeling long-term dependencies, non-linear echo modeling and noise reduction are performed on in-vehicle signals, thereby balancing echo cancellation and speech retention, reducing speech distortion, increasing speech quality, and improving the speech interaction experience.

[0164] This application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the training method of the speech recognition model as described above are implemented.

[0165] It can be understood that a computer program includes computer program code. The computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable storage medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), and software distribution medium, etc.

[0166] In the description of this specification, the descriptions referring to terms such as "specifically", "further", "specially", "understandably", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiments or examples are included in at least one embodiment or example of the present application. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0167] Any process or method description shown in the flowchart or described in other ways herein can be understood as representing a module, segment, or part of executable request code including one or more steps for implementing a specific logical function or process, and the scope of the preferred embodiments of the present application includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in the reverse order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the art of the embodiments of the present application.

[0168] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.

Claims

1. A method for echo cancellation in a vehicle cabin, characterized in that: The method comprises: Determine a reference signal according to the loudspeaker signal; Determine a first target signal according to the microphone signal and the reference signal; The speaker signal is subjected to echo cancellation processing according to the microphone signal, the first target signal and the reference signal.

2. The method according to claim 1, characterized in that The step of determining the reference signal according to the loudspeaker signal comprises: Splitting the loudspeaker signal; Determining a first type of reference signal based on a spatial position of a speaker within the vehicle cabin; Determine the second type of reference signal based on the functional priority of the vehicle system; The reference signal is determined according to the first-type reference signal and the second-type reference signal.

3. The method according to claim 1, characterized in that The step of determining the first target signal according to the microphone signal and the reference signal comprises: Determine a microphone frequency domain signal according to the microphone signal; Determine a reference frequency domain signal according to the reference signal; Determine a residual error signal according to the microphone frequency domain signal and the reference frequency domain signal; The first target signal is determined according to the residual error signal.

4. The method according to claim 3, characterized in that The determining of the residual error signal according to the microphone frequency domain signal and the reference frequency domain signal comprises: Determine a current error signal according to the microphone signal, the filter coefficient vector at the previous moment and the reference frequency domain signal; Determine a gain vector according to a predetermined forgetting factor, an inverse correlation matrix at a previous moment and the reference frequency domain signal; Determine a current filter coefficient vector according to the filter coefficient vector at the previous moment, the current error signal and the gain vector; The residual error signal is determined according to the current filter coefficient vector, the microphone frequency domain signal and the reference frequency domain signal.

5. The method according to claim 4, characterized in that The method further comprises: The filter coefficient vector at the previous moment is updated according to the current filter coefficient vector.

6. The method according to claim 4, characterized in that The method further comprises: Determine a current inverse correlation matrix according to the inverse correlation matrix at the previous moment, the gain vector, the forgetting factor and the reference frequency domain signal; The inverse correlation matrix at the previous moment is updated according to the current inverse correlation matrix.

7. The method according to claim 1, characterized in that The reference signal includes a first type of reference signal and a second type of reference signal, the first type of reference signal includes a plurality of first type of reference sub-signals, each of the first type of reference sub-signals is determined based on a spatial position of a speaker in the vehicle cabin, and performing echo cancellation processing on the speaker signal according to the microphone signal, the first target signal and the reference signal, comprising: Performing superposition processing on a plurality of the first-type reference sub-signals and the second-type reference signal to determine a target reference signal; The speaker signal is subjected to echo cancellation processing according to the target reference signal, the microphone signal and the first target signal.

8. The method according to claim 1, characterized in that The performing echo cancellation processing on the loudspeaker signal according to the target reference signal, the microphone signal and the first target signal comprises: Performing feature extraction processing on the target reference signal, the microphone signal and the first target signal to determine a target joint feature; Encoding the target joint features to determine high-dimensional embedding features; Performing temporal linking processing on the high-dimensional embedding features to determine temporal enhancement features; Decoding the high-dimensional embedding features and the temporal enhancement features to determine a target reconstructed spectrum; Mapping processing is performed on the target reconstructed spectrum to achieve echo cancellation processing on the loudspeaker signal.

9. A server, characterized in that: The server includes a processor and a memory, wherein a computer program is stored in the memory. When the computer program is executed by the processor, the method according to any one of claims 1 to 8 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Cited By

  • Commercial vehicle cab active noise reduction weight automatic adjusting method and system

    CN121415755A