In-vehicle active noise control method and system based on convolutional cyclic network model

By proposing an in-vehicle active noise control method based on a convolutional recurrent network model, the real-time and multi-channel scalability issues of in-vehicle noise control are solved, achieving low-latency real-time noise reduction and multi-channel adaptability, thus improving the practicality and efficiency of in-vehicle noise control.

CN120977274APending Publication Date: 2025-11-18华研慧声(苏州)电子科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511096663.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing technologies for in-vehicle noise control suffer from poor real-time performance, inadequate noise reduction for mid-to-high frequencies, complex parameter tuning, and poor scalability. In particular, they are unable to meet the real-time requirements of in-vehicle systems when using multi-channel control.

Method used

An in-vehicle active noise control method based on a convolutional recurrent network model is adopted. By acquiring sensor signals in real time, performing short-time Fourier transform, caching and inputting them into a trained convolutional recurrent network model, and outputting control speaker signals, the method achieves dynamic storage and updating of time series, reducing caching requirements and model computation.

Benefits of technology

It achieves low-latency end-to-end real-time noise control, reduces cache requirements and model computation, is suitable for embedded AI deployment, has wide bandwidth and non-linear noise reduction capabilities, and supports multi-channel expansion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120977274A_ABST
    Figure CN120977274A_ABST
Patent Text Reader

Abstract

The invention discloses an in-vehicle active noise control method and system based on a convolutional recurrent network, and the method comprises the steps: obtaining reference signals collected by a sensor in real time, forming a cache data set by M reference signals, and enabling reference data in the cache data set to slide rightwards for a jump length after a short-time Fourier transform sliding window; the reference signals in the cache data set are subjected to short-time Fourier transform according to algorithm updating frequency, and one frame of time-frequency data is formed in each time of short-time Fourier transform; storing the N frames of time-frequency data by adopting a first-in first-out principle; and inputting the N frames of time-frequency data into the trained convolutional cyclic network model, and outputting an output signal for controlling a loudspeaker according to inverse Fourier transform. According to the method, the recurrent neural network is used for dynamically storing and updating time sequence steps, the noise reduction capability of the depth model is reserved, and the requirement for real-time noise reduction of the whole vehicle is met.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent noise control, and in particular to an in-vehicle active noise control method and system based on a convolutional recurrent network model. BACKGROUND

[0002] When riding in a car, especially at high speed, the noise in the car is very high, mainly engine compartment noise, tire noise, air vortex noise and several other aspects. Strong noise is not conducive to communication between passengers in the car, and driving for a long time in a noisy environment is very easy to cause fatigue.

[0003] In the prior art, classical FxLMS adaptive filtering is usually used to reduce road noise, as shown in the accompanying drawings, and the main process of this active noise reduction method is: using an accelerometer to pick up road information linear FIR filter error microphone feedback LMS update. Although this noise reduction method is real-time and low in computing power, it is limited to linear narrowband, strongly dependent on secondary path estimation, complex to adjust parameters, and poor in noise reduction effect for medium and high frequencies. When using deep learning models such as WaveNet-ANC and TCN-ANC, in order to obtain sufficient context, hundreds of milliseconds of cache are required, which does not meet the real-time requirements of vehicle-mounted devices, and the scalability is poor when controlling multiple channels, and the learning parameters increase geometrically. SUMMARY

[0004] To overcome the above-mentioned shortcomings, the purpose of the present application is to provide an in-vehicle active noise control method and system based on a convolutional recurrent network model, which uses recurrent neural network dynamic storage and updates time series steps, retaining the noise reduction ability of deep models and meeting the requirements of real-time noise reduction of the whole vehicle.

[0005] To achieve the above purpose, the technical solution adopted by the present application is: an in-vehicle active noise control method based on a convolutional recurrent network, comprising: real-time acquisition of reference signals collected by sensors, M reference signals forming a cache data set, the reference data in the cache data set being right-slid by a jump length after short-time Fourier transform sliding window; the reference signals in the cache data set are subjected to short-time Fourier transform according to an algorithm update frequency, and each short-time Fourier transform forms a frame of time-frequency data; N frames of time-frequency data are stored using the first-in-first-out principle; N frames of time-frequency data are input into a trained convolutional recurrent network model and an output signal of a loudspeaker is output according to inverse Fourier transform.

[0006] The beneficial effects of the present application are that: since only N frames of time-frequency data are cached and the update step is the hop length, the real-time response demand of vehicle noise reduction is met, the cache demand and model operation amount are reduced, the hop length in each short-time Fourier transform is completely aligned, and the end-to-end low delay effect is achieved. The parameter amount of the convolutional recurrent network model is controlled in the hundreds of thousands, the memory occupancy is less than 200MB during inference, and it is easy to deploy in embedded AI and run in real time.

[0007] Further, inputting the N frames of time-frequency data into the trained convolutional recurrent network model and outputting the output signal of the loudspeaker according to the inverse Fourier transform includes: inputting the N frames of time-frequency data into the trained convolutional recurrent network model, and outputting N frames of denoising spectrograms; performing inverse Fourier transform on the last frame of the N frames of denoising spectrograms, and adding the data corresponding to the hop length calculated by the overlap-add method as the output signal, wherein the inverse Fourier transform is triggered according to an algorithm update frequency.

[0008] Further, the algorithm update frequency = the sampling frequency of the sensor / hop length, and the hop length = the Fourier transform length-data overlap length.

[0009] Further, the cache data set is stored in a ring buffer, and the reference signal is first-in-first-out in the ring buffer.

[0010] Further, before inputting the N frames of time-frequency data into the convolutional recurrent network model, the time-frequency data is first decomposed into real and imaginary parts, and the real and imaginary parts are stored along the channel direction of the convolutional recurrent network model.

[0011] Further, the convolutional recurrent network model includes an encoding layer, a recurrent neural network layer and a decoding layer connected in sequence, and the loss function of the convolutional recurrent network model includes the influence of the secondary channel.

[0012] Further, the time-frequency data format input into the convolutional recurrent network model is frequency x time step x channel number*2, and when the number of loudspeakers is expanded, only the channel number of the convolutional recurrent network model needs to be expanded.

[0013] Further, the N is not less than 5.

[0014] The present application also discloses an in-vehicle active noise system based on a convolutional recurrent network, comprising: a sensor for acquiring a reference signal in real time; a circular buffer for storing M reference signal forming cache data sets, the reference data in the circular buffer sliding backward with a data overlap length after a short-time Fourier transform sliding window; a short-time Fourier transform module for performing a short-time Fourier transform on the reference signals in the cache data sets according to an algorithm update frequency to form a frame of time-frequency data; a first-in-first-out stack for storing N frames of the time-frequency data according to a first-in-first-out principle; a convolutional recurrent network module storing a trained convolutional recurrent network model and outputting N frames of spectrograms according to the N frames of the time-frequency data; an inverse Fourier transform module for performing an inverse Fourier transform on the last frame of the output spectrograms in the N frames of the spectrograms and adding data corresponding to a hop length point calculated by an overlap-add method as the output signal; a loudspeaker for outputting sound according to the output signal.

[0015] The application further discloses a storage medium having a computer program stored thereon, and the program is executed by a processor to implement the steps of the control method. BRIEF DESCRIPTION OF DRAWINGS

[0016] Figure 1 for a flowchart of the method in the embodiment of the application Figure 1 ; Figure 2 for a flowchart of the method in the embodiment of the application Figure 2 ; Figure 3 for a data schematic diagram of the circular buffer and the first-in-first-out stack in the embodiment of the application Figure 4 for a multi-channel expansion mode of the convolutional recurrent network model in the embodiment of the application Figure 5 for a system block diagram in the embodiment of the application Figure 6 for a comparison diagram of the prior art and the noise reduction of the application in the embodiment of the application DETAILED DESCRIPTION

[0017] The preferred embodiments of the application are described in detail below with reference to the accompanying drawings, so that the advantages and features of the application can be more easily understood by those skilled in the art, and the protection scope of the application can be more clearly defined.

[0018] A vehicle interior active noise control method based on a convolutional recurrent network, as shown in FIGS. 1 and 2, comprises the following steps. Figure 1 and FIG. 2 Figure 2 The method comprises the following steps. S100, real-time acquisition of reference signals collected by sensors, M reference signals form a cache data set, and the reference data in the cache data set is right-slid by a jump length after short-time Fourier transform sliding window.

[0019] The jump length is the number of cache data samples that the frame is right-slid after each short-time Fourier transform sliding window.

[0020] The jump length = Fourier transform length - data overlap length, and the Fourier transform length and the data overlap length are both parameters of the short-time Fourier transform.

[0021] The sensor is selected according to the type of noise reduction in the vehicle. For example, when road noise reduction is performed, the sensor is an acceleration sensor fixed on the chassis; when engine noise reduction is performed, the sensor is a speed sensor for collecting engine speed.

[0022] S200, the reference signals in the cache data set are subjected to short-time Fourier transform according to an algorithm update frequency, and each short-time Fourier transform forms a frame of time-frequency data.

[0023] The algorithm update frequency = sensor sampling frequency / jump length.

[0024] The sampling frequency of the sensor is determined by the nature of the sensor. For example, the sampling rate of the acceleration sensor is 48 kHz or 16 kHz. The algorithm update frequency needs to consider the operation capacity of the hardware system and the noise reduction effect. For road noise reduction, the algorithm update frequency is usually greater than 1500 Hz.

[0025] The algorithm update interval = 1 / algorithm update frequency = jump length / sensor sampling frequency, that is, the green data in the figure is updated. Figure 3 After the green data is updated, for example, the sampling frequency of the sensor = 16 kHz, and when the jump length = 160, the system is updated every 10 ms, and a short-time Fourier transform is performed every 10 ms. The 10 ms delay can meet the real-time demand of road noise control. The reference data in the cache data set in the figure (the sum of the yellow part and the green part, a total of M points) is processed by short-time Fourier transform, and after the next 160 points are updated, the data overlap length is slid, and a short-time Fourier transform is performed again to generate the next frame of time-frequency data. Similarly, the convolution recurrent network model in the controller and the inverse Fourier transform are also triggered only once at this time.

[0026] S300, N frames of time-frequency data are stored using the first-in-first-out principle.

[0027] N frames of time-frequency data are maintained in the time step frame, and the principle of first-in first-out is adopted. N usually needs to be selected optimally according to experiments. The smaller N is, the smaller the past time information carried is, but the faster the dynamic system responds, and the smaller the required computing power is. The larger N is, the better the system noise reduction performance performs for a steady state condition.

[0028] As shown by the yellow line, each cache data set calculates a current frame short-time Fourier transform, and splices N-1 frames of time-frequency data (the middle black line part) into time-frequency data composed of frequency and time step frame number [FxN]. Figure 3 Figure 3 As shown by the yellow line, each cache data set calculates a current frame short-time Fourier transform, and splices N-1 frames of time-frequency data (the middle black line part) into time-frequency data composed of frequency and time step frame number [FxN].

[0029] S400, inputting N frames of time-frequency data into a trained convolutional recurrent network model and outputting a control loudspeaker output signal according to inverse Fourier transform.

[0030] In the embodiment, the reference signal enters the controller in the form of a time domain signal. The controller includes a short-time Fourier transform module, a convolutional recurrent network module, and an inverse Fourier transform module. The time domain reference signal is converted into a time domain signal output by the loudspeaker after passing through the above modules. The output signal reaches the control point through a secondary path and cancels out the initial noise to achieve the purpose of noise reduction. In low delay, only a small amount of historical frames are cached. The time-frequency domain convolutional recurrent neural network captures the spatial position of the sensor and establishes an information relationship in the time sequence. The real-time output of the anti-noise signal makes the noise reduction effect take into account the wide frequency band and nonlinearity.

[0031] In the embodiment, since only N frames of time-frequency data are cached and the update step is the hop length, the end-to-end delay of the system can be controlled within the range of 20-25 ms, meeting the real-time response requirements of vehicle ANC.

[0032] Inputting N frames of time-frequency data into a trained convolutional recurrent network model and outputting a control loudspeaker output signal according to inverse Fourier transform specifically includes: S401, inputting N frames of time-frequency data into a trained convolutional recurrent network model to output N frames of denoising spectrum graphs.

[0033] S402, performing inverse Fourier transform on the last frame of denoising spectrum graph in the N frames of denoising spectrum graph, and adding the data corresponding to the hop length calculated by the overlap-add method as an output signal.

[0034] The convolutional recurrent network model output in step S401 contains information of the past N frames of data. At this time, the system output only needs to output a hop length size data to ensure that the sampling rate size does not change. Therefore, the inverse Fourier transform only needs to be performed on the last frame of denoising spectrum graph in the N frames of denoising spectrum graph. The overlap-add method is mainly to eliminate the influence of the window function on the sampling signal in the short-time Fourier transform.

[0035] ​Referring to Fig. 1, the cache dataset is stored in a ring buffer, the ring buffer stores a reference signal of a Fourier transform length M points, and the reference signal is in a first-in-first-out manner in the ring buffer. Figure 3

[0036] Before the N frames of time-frequency data are input into the convolution recurrent network model, the time-frequency data are disassembled into real parts and imaginary parts, and the real parts and the imaginary parts are stored along a channel direction of the convolution recurrent network model. At this time, the dimension becomes [F x N x 2], and the data processing format corresponds to Fig. 1. Figure 2

[0037] In one embodiment, the convolution recurrent network model comprises an encoding layer, a recurrent neural network layer (LSTM layer) and a decoding layer connected in sequence. The encoding layer extracts local time-frequency features by a plurality of 2D convolutions; the recurrent neural network layer models cross-frame dependencies in the time dimension; and the decoding layer reconstructs the F x N x 2 complex spectrum mapping by transposed convolution + jump connection. The convolution + LSTM simultaneously models the line and the loudspeaker nonlinearities, and compared with a complete batch time-frequency model, the parameter amount is halved, and the convolution recurrent network model is easy to embed.

[0038] The loss function of the convolution recurrent network model contains the influence of the secondary channel, that is, the network output is filtered by the secondary transfer function to minimize the mean square error.

[0039] The convolution recurrent network needs to be trained in advance, and only after the convolution recurrent network is trained, the convolution recurrent network model is arranged. The training of the convolution recurrent network is performed by using offline data acquisition.

[0040] The offline data includes a reference signal collected by an offline state sensor and a noise signal at a control point (error microphone), and when the system is in a real-time control stage, the error microphone is not needed as a system feedback, and compared with a traditional FxLMS method, the corresponding hardware cost is reduced. Taking a vehicle road noise as an example, the data acquisition needs to collect acceleration vibration data of a chassis selected point of the vehicle under various operating conditions as the reference signal and noise data at an in-vehicle control point as the target noise reduction signal. The more complete the data is collected, the more operating conditions are included, and the wider the range of the trained model is. The convolution recurrent network can effectively process time-frequency features, especially for time series with up-down dependencies.

[0041] In this embodiment, when the real-time noise control is performed by using the recurrent neural network model, a dynamic storage and update time sequence step mode is adopted, which not only retains the noise reduction capability of the deep model, but also meets the requirements of real-time noise reduction of the vehicle. At the same time, the data training method improves the convenience of extending the system from a single channel control to multiple channels.

[0042] Referring to Fig. 1, the cache dataset is stored in a ring buffer, the ring buffer stores a reference signal of a Fourier transform length M points, and the reference signal is in a first-in-first-out manner in the ring buffer. Figure 4 ​​As shown, the time-frequency data format of the input convolutional recurrent network (CRN) model is frequency × time step × number of channels × 2. When the number of speakers is increased, only the number of channels in the CRN model needs to be increased. When the reference is expanded from a single channel to a multi-channel reference, only the data input layer in the CRN model needs to be expanded in the direction of the number of channels; the CNN and LSTM layers in the CRN model do not need to be reset. When the speakers are expanded from a single output to multiple speaker outputs, only the network output layer needs to be expanded along the channel direction. For example, when the reference channels change from 1 to jNum and the number of speakers changes from 1 to kNum, the data changes in the network model are as follows: input changes from [F×N×2] to [F×N×2*jNum], and output changes from [F×N×2] to [F×N×2*kNum]. The CNN and LSTM parameters in the training model remain unchanged. This greatly improves the program's versatility and maintainability. Furthermore, the increase in network layer learning parameters is minimal when the system's reference and control channels increase.

[0043] Therefore, the recurrent integral network model in this embodiment can be easily extended to multi-channel noise reduction without major changes to the overall recurrent integral network model. Only a small increase in the number of convolution kernel channels is needed, without reconstructing the core network, which is more in line with actual noise reduction requirements.

[0044] In one embodiment, N is not less than 5. Experimental data shows that, due to the existence of the jump length, when N is more than 5 frames, a better noise reduction effect can be achieved, reducing the cache requirement and model computation.

[0045] See appendix Figure 6 As shown, this is a comparison of noise reduction between this embodiment and the prior art. In actual vehicle control, the method in this embodiment generates a noise attenuation of ≥5dBA in the 20-400Hz wide frequency band, which effectively reduces road noise. In other words, the method in this embodiment has a better noise reduction effect.

[0046] See appendix Figure 5 As shown, in one embodiment, an in-vehicle active noise system based on a convolutional recurrent network is also disclosed, including a sensor, a ring buffer, a controller, and a speaker.

[0047] The sensor is used to acquire reference signals in real time; the circular buffer is used to store M reference signals to form a buffered dataset, and the reference data in the circular buffer slides backward by the data overlap length after the short-time Fourier transform sliding window; the controller is used to generate control signals for controlling the loudspeaker based on the reference signals; the loudspeaker performs noise reduction based on the occurrence of the control signals.

[0048] The controller comprises a short-time Fourier transform module, a first-in first-out stack, a convolution recurrent network module and an inverse Fourier transform module, the short-time Fourier transform module is used for performing short-time Fourier transform on a reference signal in a cache data set according to an algorithm update frequency to form a frame of time-frequency data; the first-in first-out stack stores N frames of time-frequency data in a first-in first-out principle; the convolution recurrent network module stores a trained convolution recurrent network model, and outputs N frames of spectrum graphs according to the N frames of time-frequency data; and the inverse Fourier transform module performs inverse Fourier transform on a last frame of spectrum graph in the N frames of spectrum graphs, and adds data corresponding to a hop length point calculated by an overlap-add method as an output signal.

[0049] In one embodiment, a computer readable storage medium is also described, including computer instructions stored thereon, which, when executed by a processor, cause the processor to perform the active noise control method described above. It can be understood that the computer storage medium can be any tangible medium, such as a floppy disk, a CD-ROM, a DVD, a hard disk drive, or a network medium, etc.

[0050] The above embodiments are only for illustrating the technical concept and characteristics of the present application, and the purpose is to enable those skilled in the art to understand the content of the present application and implement it, and cannot limit the protection scope of the present application, and any equivalent changes or modifications made according to the spirit and essence of the present application should be covered within the protection scope of the present application.

[0051] The above embodiments are only for illustrating the technical concept and characteristics of the present application, and the purpose is to enable those skilled in the art to understand the content of the present application and implement it, and cannot limit the protection scope of the present application, and any equivalent changes or modifications made according to the spirit and essence of the present application should be covered within the protection scope of the present application.

Claims

1. A method for in-vehicle active noise control based on convolutional recurrent network, characterized in that: The method comprises: real-time acquisition of reference signals collected by a sensor, M reference signals forming a cache data set, the reference data in the cache data set being right-slid by a jump length after short-time Fourier transform sliding window; the reference signals in the cache data set being subjected to short-time Fourier transform according to an algorithm update frequency, each short-time Fourier transform forming a frame of time-frequency data; N frames of the time-frequency data being stored using a first-in-first-out principle; N frames of the time-frequency data being input into a trained convolution recurrent network model and an output signal for controlling a loudspeaker being output according to inverse Fourier transform.

2. The in-vehicle active noise control method based on a convolutional recurrent network according to claim 1, characterized in that: N frames of the time-frequency data being input into a trained convolution recurrent network model and an output signal for controlling a loudspeaker being output according to inverse Fourier transform specifically comprises: N frames of the time-frequency data being input into a trained convolution recurrent network model, and N frames of denoised spectrograms being output; inverse Fourier transform being performed on the last frame of the N frames of the denoised spectrograms, and data corresponding to a jump length calculated by overlap-add method being added to the output as the output signal, the inverse Fourier transform being triggered according to an algorithm update frequency. 3.The in-vehicle active noise control method based on a convolutional recurrent network according to claim 1, wherein: The algorithm update frequency = sampling frequency of the sensor / jump length, and the jump length = Fourier transform length - data overlap length.

4. The in-vehicle active noise control method based on a convolutional recurrent network according to claim 1, characterized in that: The cache data set is stored in a ring buffer, and the reference signals are stored in the ring buffer using a first-in-first-out principle.

5. The in-vehicle active noise control method based on a convolutional recurrent network according to claim 1, characterized in that: Before N frames of the time-frequency data are input into the convolution recurrent network model, the time-frequency data are first decomposed into real and imaginary parts, and the real and imaginary parts are stored along a channel direction of the convolution recurrent network model.

6. The in-vehicle active noise control method based on a convolutional recurrent network according to any one of claims 1-5, characterized in that: The convolution recurrent network model comprises an encoding layer, a recurrent neural network layer and a decoding layer connected in sequence, and a loss function of the convolution recurrent network model comprises an influence of a secondary channel.

7. The in-vehicle active noise control method based on a convolutional recurrent network according to claim 6, characterized in that: The time-frequency data input into the convolution recurrent network model are in a format of frequency x time step x channel number * 2, and when the number of the loudspeakers is expanded, only the channel number of the convolution recurrent network model needs to be expanded.

8. The in-vehicle active noise control method based on a convolutional recurrent network according to claim 1, characterized in that: The N is not less than 5.

9. An in-vehicle active noise system based on a convolutional recurrent network, characterized by: The method comprises: a sensor for real-time acquisition of reference signals; a ring buffer for storing M reference signals to form a cache data set, the reference data in the ring buffer being backward-slid by a data overlap length after short-time Fourier transform sliding window; a short-time Fourier transform module for performing short-time Fourier transform on the reference signals in the cache data set according to an algorithm update frequency to form a frame of time-frequency data; a first-in-first-out stack for storing N frames of the time-frequency data using a first-in-first-out principle; a convolution recurrent network module for storing a trained convolution recurrent network model and outputting N frames of spectrograms according to N frames of the time-frequency data; an inverse Fourier transform module for performing inverse Fourier transform on the last frame of the N frames of the output spectrograms, and adding data corresponding to a jump length calculated by overlap-add method to the output as the output signal; a loudspeaker for outputting sound according to the output signal.

10. A storage medium having stored thereon a computer program, characterized in that: The program is executed by a processor to implement the steps of the control method in any one of claims 1-8.