A circuit for intelligent noise reduction and signal enhancement of time series signals based on sensor array
Through the multi-stage noise reduction and signal enhancement circuit combined with STFT feature extraction and neural network, the problem of distinguishing interference sources and target sources of sensor array timing signals in complex environments is solved, and intelligent noise reduction and signal enhancement is achieved, which is suitable for diverse application scenarios.
Patent Information
- Application Number
- CN202411465211.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-21
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2044-10-21
AI Technical Summary
It is difficult for sensor array timing signals to effectively distinguish and process interfering signals and target signals in complex environments. Especially when the interference source and noise source exist at the same time and the energy is close, it is difficult for traditional methods to achieve effective noise reduction and signal enhancement, and traditional methods to reduce processing performance when the interference signal and target signals change.
The STFT feature extraction circuit, multi-source positioning circuit and multi-stage noise reduction and signal enhancement circuit are used to combine neural network technology for signal processing. Through STFT feature extraction, interference source and target source orientation information positioning, and signal processing is performed using multi-stage noise reduction and signal enhancement circuit based on multi-source orientation information to achieve suppression of interference signals and enhancement of target signals.
It realizes intelligent noise reduction and signal enhancement of sensor array timing signals in complex environments, can effectively distinguish interference sources from target sources, supports multiple processing modes, reduce hardware resource overhead, and is suitable for diversified application scenarios.
Smart Images

Figure CN119341868B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of array timing signal processing, and in particular to a timing signal intelligent noise reduction and signal enhancement circuit based on a sensor array. Background Art
[0002] With the development of sensor technology, sensor arrays have been widely used in various fields, such as radar, sonar, communications, and sound pickup. Sensor arrays can achieve functions such as positioning or signal acquisition by receiving signals from different directions. However, in practical applications, the time-series signals received by sensor arrays are often affected by various complex interferences, such as environmental noise, mixed signals, and multipath effects. These interferences reduce the signal-to-noise ratio and affect signal performance. To obtain effective information, the collected time-series signals from the sensor array need to be processed accordingly to reduce noise interference and enhance the target signal of interest. To address the problem of interference in sensor array time-series signals, traditional array noise reduction and signal enhancement methods mainly include frequency-domain filtering, sensor array beamforming, and adaptive filtering. The method based on frequency domain filtering can better process interference signals with fixed frequency domain energy distribution bands that do not overlap with the frequency domain energy distribution bands of target signals. However, when the frequency domain energy distribution bands of the interference have large fluctuations or cross or overlap with the frequency domain energy distribution bands of the target signals, there will often be a significant performance degradation. The beamforming based on the sensor array is usually based on the azimuth information of the interference signal, and has a good suppression effect on different types of interference. It can also distinguish and solve the situation where there is frequency domain intersection between the target signal and the interference signal that cannot be processed by frequency domain filtering in the spatial domain. However, for the situation where interference signals and target signals exist at the same time and the azimuths of the interference signals and target signals change, its processing performance will be greatly lost. Adaptive filtering is based on the statistical characteristics of the signal, and has a good suppression effect on different types of interference and can effectively reduce interference. However, it still has some limitations. In practical applications, interference signals often have non-statistical characteristics, such as pulse interference, non-Gaussian noise, etc. In the case where interference signals and noise signals exist at the same time, traditional noise reduction and signal enhancement methods usually distinguish target signals from interference signals based on the energy size of the signals. However, in actual applications, the energy of target signals and interference signals may be similar, making it difficult to effectively distinguish them, making it difficult to effectively perform noise reduction and signal enhancement. It is even more difficult to deal with the situation where the directions of noise signals and target signals change.
[0003] Therefore, there are several requirements and challenges for timing signal processing of sensor arrays in various application scenarios: 1) Effective noise reduction and signal enhancement of array timing signals in complex scenarios. For example, when interference sources and noise sources coexist with similar energies and azimuth fluctuations, traditional methods have difficulty effectively locating interference sources and target source signals, making it difficult to perform effective noise reduction and signal enhancement. 2) Flexible and configurable. For the diverse application scenarios of sensor array timing signals, it is usually necessary to support different configurable noise reduction and target enhancement processing algorithms or modes. Summary of the Invention
[0004] To address the above problems, the present invention proposes an intelligent noise reduction and signal enhancement circuit for time series signals based on a sensor array, which can achieve noise reduction and signal enhancement in complex environments.
[0005] The present invention provides an intelligent noise reduction and signal enhancement circuit for a time series signal based on a sensor array, comprising: an STFT feature extraction circuit, a multi-source positioning circuit for obtaining interference source and target source azimuth information, and a multi-stage noise reduction and signal enhancement circuit based on the multi-source azimuth information. The STFT feature extraction circuit is used to convert an input sensor array time series signal into a time-frequency domain to obtain its time-frequency signal; the multi-source positioning circuit for obtaining interference source and target source azimuth information is used to process the time-frequency signal processed by the STFT feature extraction circuit, and obtain the interference source and target source azimuths, i.e., angle values, of the currently processed array time series signal after processing; and the multi-stage noise reduction and signal enhancement circuit based on the multi-source azimuth information is used to process the time-frequency signal processed by the STFT feature extraction circuit in combination with the obtained interference source and target source azimuth information, thereby suppressing the interference signal and enhancing the target signal.
[0006] Furthermore, the working process of the intelligent noise reduction and signal enhancement circuit of the timing signal based on the sensor array is as follows: first, the parameters required for the multi-stage noise reduction and signal enhancement circuit based on multi-source azimuth information are initialized and stored in the corresponding storage circuit module; after detecting the multi-source positioning trigger signal sent by the external controller, the STFT result of the sensor array timing signal obtained by the STFT feature extraction circuit is sent to the multi-source positioning circuit for obtaining the azimuth information of the interference source and the target source for array signal processing; the interference source and the target source are located using the direction of arrival DOA estimation scheme based on the neural network, and the positioning result includes two dimensions of information, namely whether the interference source and the target source exist and their corresponding specific azimuth information; then the STFT result and the DOA result are sent to the multi-stage noise reduction and signal enhancement circuit based on the multi-source azimuth information for noise reduction and signal enhancement processing.
[0007] Furthermore, the STFT feature extraction circuit includes a pre-emphasis module, a framing module, a windowing module and a fast Fourier transform module. The pre-emphasis module processes the sensor array timing signal and sends the processed signal to the framing module to compensate for the loss of high-frequency components of the original timing signal. The specific processing method is: y(t) = x(t) - μx(t-1), where y(t), x(t) and x(t-1) respectively represent the output signal after pre-emphasis processing at time t, the original input signal and the input signal at the previous time, and μ is the pre-emphasis coefficient; the framing module frames the sensor array timing signal output by the pre-emphasis module according to the point number requirement of the fast Fourier transform module, and sends the framing results to the windowing module in sequence; the windowing module multiplies the signal output by the framing module by the windowing coefficient and sends the result to the fast Fourier transform module; the fast Fourier transform module performs fast Fourier transform processing on the signal output by the windowing module and outputs the STFT result of the sensor array timing signal.
[0008] Furthermore, the multi-source positioning circuit for obtaining the azimuth information of the interference source and the target source includes a complex phase difference calculation module, an inverse covariance matrix multiplication module, a DoA angle prediction module, a DoA histogram statistics module and a deep neural network module;
[0009] The complex phase difference calculation module uses the Q-order Cordic algorithm to calculate the phase of the complex signal and the phase difference between the reference sensor channel signal and the remaining sensor channel signals. The number of sensors N is greater than or equal to 2, and the order Q of the Cordic algorithm is selected according to actual needs. The phase difference calculation formula is:
[0010]
[0011] Where i∈[1, N-1], t represents time, f represents frequency, X0(t, f) is the STFT result of the reference sensor channel, X i (t, f) are the STFT results of the remaining sensor channels, Re represents the real part, and Im represents the imaginary part;
[0012] After obtaining the phase signal output by the complex phase difference calculation module, the inverse covariance matrix multiplication module converts the phase signal Multiplying with the inverse covariance matrix A, we can calculate the trigonometric function vector S of the sound source. Where α is the angle of the target or interference sound source, and the inverse covariance matrix is determined by the current sensor array;
[0013] The DoA angle prediction module uses the K-order Cordic algorithm to calculate the arctangent value of the trigonometric function vector output by the inverse covariance matrix module to obtain the sound source angle at each time-frequency point. The DoA angle prediction formula is as follows:
[0014]
[0015] Among them, DoA pred is the DoA angle prediction value, sinα and cosα are the elements in the trigonometric function vector;
[0016] The DoA histogram statistics module consists of a judgment logic circuit and an angle distribution histogram statistics register group. The DoA histogram statistics module collects statistics on the sound source angle signal output by the DoA angle prediction module and stores it in the histogram statistics register group.
[0017] The deep neural network module uses the statistical histogram output of the DoA histogram statistics module as the feature map input, which is processed by the deep neural network. The effective signals of the interference source and the target source as well as their specific orientation information, i.e., the angle value, are obtained through multiple output branches of the network.
[0018] Furthermore, the multi-stage noise reduction and signal enhancement circuit based on multi-source azimuth information includes an array-based frequency domain beamforming AFDBF module, a neural network-based beamforming NBBF module, a beamforming multi-stage fusion MFBF module, a control circuit module, and a vector coding module for vector coding of effective azimuth angles;
[0019] For the N-channel STFT input signal extracted by the STFT feature extraction circuit, the AFDBF module has N-1 AFDBF processing units corresponding to N-1 STFT outputs. One of the N-channel signals is a reference signal. A single AFDBF unit consists of a BF coefficient storage module and an arithmetic processing circuit. The BF coefficient storage module is used to store noise reduction and enhancement coefficients related to frequency, interference source, and target source orientation. The arithmetic circuit obtains coefficients based on the orientation information of the target source or interference source, and performs multiplication-addition enhancement or multiplication-subtraction noise reduction processing based on the existence information of the target source or interference source.
[0020] The NBBF module receives the data output from the STFT feature extraction circuit and the vector encoding circuit, and sequentially splices them as the feature map data of the neural network input, and sends them to the pre-trained neural network beamforming module to extract the time-frequency mask M with the target source enhancement and interference source noise reduction effect, that is, the neural network output result, and outputs N-1 groups of complex time-frequency masks M (c, ω i ), whose size is consistent with the AFDBF output size;
[0021] After receiving the STFT result Y processed by the AFDBF module and the time-frequency mask result M generated by the NBBF module, the MFBF module performs element-level multiplication fusion on the two, and then performs element-level addition fusion on the fused results of N-1 channels, and finally outputs the STFT result R after single-channel denoising and enhancement processing;
[0022] The control circuit module controls the four processing modes of the multi-stage noise reduction and signal enhancement circuit of multi-source azimuth information according to the values of the effective signal of the interference source and the effective signal of the target source;
[0023] The vector coding module receives scalar information of the interference source azimuth and the target source azimuth from the multi-source positioning circuit that obtains the interference source and target source azimuth information, as well as a control signal from the control circuit to determine whether the vector coding of the multi-source azimuth needs to be updated and whether it needs to be set to an invalid vector coding state.
[0024] Furthermore, the specific functions of the AFDBF unit are:
[0025] First, the obtained interference source azimuth information θ n , target source orientation information θ d And the corresponding FFT frequency ω i As the index, the frequency domain coefficient F related to the direction and frequency of the source signal, which is pre-generated and stored in the BF coefficient storage module, is taken out. The frequency domain coefficient is used to compensate for the difference in the target signal between the two channels by complex multiplication. Then, the frequency domain coefficient is used to compensate for the difference in the interference signal between the two channels by complex multiplication. Finally, the enhanced signal and the noise reduction signal are selectively added according to the effectiveness of the source signal to obtain the adaptive noise reduction and enhanced signal Y. The expression is:
[0026] Yd(c,ω i )=X0(ω i )+X(c,ω i )F(c,ω i ,θ d )
[0027] Yn(c,ω i )=X0(ω i )-X(c,ω i )F(c,ω i ,θ n )
[0028]
[0029] Where X0(*) represents the FFT result of a frame of sensor array timing signals as a reference channel input by the STFT feature extraction circuit, X(*) represents the FFT result of a frame of sensor array timing signals of other channels input by the STFT feature extraction circuit, F(*) represents the frequency domain coefficient used to compensate for the difference between the target source or interference source channels, Yd(*) represents the processing result of enhancing the target signal between channels, Yn(*) represents the processing result of suppressing the interference signal between channels, and Y(*) represents the final output result of the AFDBF processing unit. i represents the digital frequency index, ω i =2πf i , f i is the physical frequency point in FFT transformation, Among them, f s is the digital signal sampling rate, Num represents the number of FFT points, and the value range of i is θ d represents the target source azimuth from the multi-source positioning circuit, θ n It represents the azimuth of the interference source from the multi-source positioning circuit, d represents the effective signal of the target source, which is 1 if it is effective and 0 if it is invalid, n represents the effective signal of the interference source, which is 1 if it is effective and 0 if it is invalid, and c represents the STFT channel index, which is also the AFDBF processing unit index, and its value range is [1, N-1].
[0030] Furthermore, according to the interference source and target source information, the multi-stage noise reduction and signal enhancement circuit of the multi-source azimuth information is divided into four processing modes. The first is that when the interference source and the target source exist at the same time, multi-stage noise reduction is implemented on the interference source, multi-stage enhancement is implemented on the target source, and the processed STFT result is output; the second is that when only the interference source exists, multi-stage noise reduction is implemented on the interference source, and the processed STFT result is output; the third is that when only the target source exists, multi-stage enhancement is implemented on the target source, and the processed STFT result is output; the fourth is that when neither the interference source nor the target source exists, the STFT result of the reference sensor channel timing signal in the original multi-sensor channel is directly output.
[0031] The beneficial technical effects of the present invention are:
[0032] (1) The present invention proposes a multi-source positioning circuit that processes sensor array timing signals in the time and frequency domain to obtain the azimuth information of interference sources and target sources. The array signal processing technology and neural network processing technology used in the multi-source positioning circuit can effectively utilize the spatial information and statistical information of the array signal, and combine the learning ability of the neural network to effectively locate and distinguish the target signal and the interference signal.
[0033] (2) The present invention proposes a multi-stage noise reduction and signal enhancement circuit that combines the target source and interference source azimuth information to adaptively reduce the noise of the sensor array timing signal and enhance the target signal. The multi-stage noise reduction and signal enhancement circuit supports multiple processing modes such as noise reduction, enhancement, and simultaneous noise reduction and enhancement. In addition, the algorithm coefficients used for noise reduction or enhancement are highly configurable and can be expanded to support different noise reduction and enhancement algorithms by modifying the coefficients. The neural network technology used can support customized optimization networks for diverse application scenarios and is flexibly applicable to intelligent noise reduction and signal enhancement in various scenarios. The multi-source positioning circuit and the multi-stage noise reduction and signal enhancement circuit work together to achieve intelligent noise reduction and signal enhancement of the sensor array timing signal in a complex environment. In addition, both circuits process the signal in the time-frequency domain. Based on the short-term stability of the signal, a multiplexing design of the STFT feature extraction circuit is adopted to reduce hardware resource overhead. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0035] Figure 1 This is a structural diagram of a sensor array-based intelligent noise reduction and signal enhancement circuit for timing signals provided by an embodiment of the present invention;
[0036] Figure 2 is a structural diagram of an STFT feature extraction circuit provided by an embodiment of the present invention;
[0037] Figure 3 2. It is a structural diagram of a multi-source positioning circuit for obtaining position information of interference sources and target sources provided by an embodiment of the present invention;
[0038] Figure 4 1 is a structural diagram of a multi-stage noise reduction and signal enhancement circuit based on multi-source orientation information provided by an embodiment of the present invention;
[0039] Figure 5 is a schematic diagram of a four-sensor array provided by an embodiment of the present invention;
[0040] Figure 6 4 is a structural diagram of a DoA histogram statistics circuit provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0041] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0042] The structure diagram of the intelligent noise reduction and signal enhancement circuit of the timing signal based on the sensor array proposed by the present invention is as follows: Figure 1 As shown, the time series signal intelligent noise reduction and signal enhancement circuit includes an STFT (Short-Time Fourier Transform) feature extraction circuit, a multi-source positioning circuit for obtaining interference source and target source azimuth information, and a multi-stage noise reduction and signal enhancement circuit based on the multi-source azimuth information. The STFT feature extraction circuit is used to convert the input sensor array time series signal into the time-frequency domain to obtain its time-frequency signal; the multi-source positioning circuit for obtaining interference source and target source azimuth information is used to process the time-frequency signal processed by the STFT feature extraction circuit, and obtain the interference source and target source azimuth of the currently processed array time series signal, that is, the angle value after processing; the multi-stage noise reduction and signal enhancement circuit based on the multi-source azimuth information is used to process the time-frequency signal processed by the STFT feature extraction circuit in combination with the obtained interference source and target source azimuth information, so as to suppress the interference signal (that is, noise reduction) and enhance the target signal (that is, signal enhancement).
[0043] The working process of the intelligent noise reduction and signal enhancement circuit of the timing signal based on the sensor array is as follows: first, the parameters required for the multi-stage noise reduction and signal enhancement circuit based on the multi-source orientation information are initialized and stored in the corresponding storage circuit module; after detecting the multi-source positioning trigger signal sent by the external controller, the STFT result of the sensor array timing signal obtained by the STFT feature extraction circuit is sent to the multi-source positioning circuit that obtains the orientation information of the interference source and the target source for array signal processing, and the direction of arrival (DOA) based on the neural network is used. Abstract: In order to improve the positioning accuracy of multi-sensor channels, an interference source and a target source are proposed. The interference source and the target source are located by the DOA estimation scheme. The positioning result mainly contains two dimensions of information, namely, whether the interference source and the target source exist and their corresponding specific azimuth information. Then the STFT result and DOA result are sent to the multi-stage noise reduction and signal enhancement circuit based on multi-source azimuth information for processing. According to the interference source and target source information, the multi-stage noise reduction and signal enhancement circuit of multi-source azimuth information is divided into four processing modes: 1) The interference source and the target source exist at the same time: multi-stage noise reduction is implemented for the interference source, and multi-stage enhancement is implemented for the target source, and the processed STFT result is output; 2) Only the interference source exists: multi-stage noise reduction is implemented for the interference source, and the processed STFT result is output; 3) Only the target source exists: multi-stage enhancement is implemented for the target source, and the processed STFT result is output; 4) Neither the interference source nor the target source exists: the STFT result of the timing signal of the reference sensor channel in the original multi-sensor channel is directly output. The four processing modes of the multi-stage noise reduction and signal enhancement circuit based on multi-source orientation information can flexibly and intelligently process diverse low-SNR time-series signals from sensor arrays in different application scenarios, ensuring time-series signal performance. The proposed multi-source localization circuit for acquiring interference and target source orientation information and the multi-stage noise reduction and signal enhancement circuit based on multi-source orientation information both process sensor array signals in the time-frequency domain.
[0044] In order to meet the processing requirements of the multi-source positioning circuit for obtaining the interference source and target source position information and the multi-stage noise reduction and signal enhancement circuit based on the multi-source position information, a STFT feature extraction circuit is specifically designed, such as Figure 2 The STFT feature extraction circuit includes a pre-emphasis module, a framing module, a windowing module, and a Fast Fourier Transform (FFT) module, which can extract the STFT features of the sensor array timing signal.
[0045] Pre-emphasis module: processes the sensor array timing signal and sends the processed signal to the framing module. The purpose of the pre-emphasis module is to compensate for the loss of high-frequency components of the original timing signal. The specific processing method is as follows:
[0046] y(t)=x(t)-μx(t-1)
[0047] Where y(t), x(t), and x(t-1) represent the pre-emphasized output signal at time t, the original input signal, and the input signal at the previous time, respectively. μ is the pre-emphasis coefficient, with a typical value between 0.9 and 1.0.
[0048] Framing module: Based on the point number requirements of the fast Fourier transform module, the sensor array timing signal output by the pre-emphasis module is framed. For example, for an N-point FFT, each channel needs to divide the continuous timing signal into a group of data (a frame) every N points and send them to the subsequent windowing module in sequence.
[0049] Windowing module: multiplies the signal output by the framing module by the windowing coefficient and then sends it to the fast Fourier transform module.
[0050] Fast Fourier transform module: performs fast Fourier transform processing on the signal output by the windowing module and outputs the STFT result of the sensor array signal.
[0051] The structure of the multi-source positioning circuit for obtaining the position information of the interference source and the target source is as follows: Figure 3 As shown, it includes a complex phase difference calculation module, an inverse covariance matrix multiplication module, a DoA angle prediction module, a DoA histogram statistics module and a deep neural network module.
[0052] Complex phase difference calculation module: The FFT signal received from the STFT feature extraction circuit is a complex signal, containing real and imaginary parts. The complex phase difference calculation module uses a Q-order coordinate rotation digital computer (Cordic) algorithm to calculate the phase of the complex signal and the phase difference between the reference sensor channel signal and the remaining sensor channel signals. The reference sensor is a sensor in the sensor array, and the specific selection is related to the array algorithm design. The number of sensors N is greater than or equal to 2, and the order Q of the Cordic algorithm can be selected as needed, and the typical value is usually 8. The formula for calculating the phase difference is as follows:
[0053]
[0054] Where i∈[1, N-1], t represents time, f represents frequency, X0(t, f) is the STFT result of the reference sensor channel, Xi(t, f) is the STFT result of the remaining sensor channels, Re represents the real part, and Im represents the imaginary part.
[0055] Inverse covariance matrix multiplication module: After obtaining the phase signal output by the complex phase difference calculation module, the phase signal Multiplying it with the inverse covariance matrix A, we can calculate the trigonometric function vector S of the sound source. The calculation expression is as follows:
[0056]
[0057] Where α is the angle of the target (or interference) sound source, and the inverse covariance matrix is determined by the current sensor array. Here is a way to obtain the inverse covariance matrix:
[0058]
[0059] A=(P T P) -1 P T
[0060] Among them, P0, ..., P N-1 Represents the two-dimensional coordinate vector of each sensor, including X coordinate and Y coordinate, and P represents the sensor coordinate matrix. Give an example of the calculation of a four-sensor matrix, such as Figure 5 As shown in the figure, assuming that the receiving distance d0 between adjacent sensors is 0.04m and the reference sensor is sensor 1, P and A can be calculated as follows:
[0061]
[0062] The calculation of the inverse covariance matrix only needs to be completed off-chip, and the specific parameters can be configured in on-chip registers through initialization.
[0063] DoA Angle Prediction Module: This module uses the K-order Cordic algorithm to calculate the inverse tangent of the trigonometric function vector output by the Inverse Covariance Matrix module to obtain the sound source angle at each time-frequency point. The typical value of the Cordic algorithm's order K is 8. The DoA angle prediction formula is as follows:
[0064]
[0065] Among them, DoA pred is the DoA angle prediction value, sinα and cosα are the elements in the trigonometric function vector.
[0066] DoA histogram statistics module: The DoA histogram statistics module consists of a judgment logic circuit and an angle distribution histogram statistics register group, such as Figure 6 As shown in Figure 1, the DoA histogram statistics module collects statistics on the sound source angle signal output by the DoA angle prediction module and stores it in the histogram statistics register group. The length H of the register group is inversely proportional to the DoA angle resolution R, which can be expressed as the following formula:
[0067]
[0068] The i-th element of the statistical register group represents the number of sound source angles within the interval [i×R, (i+1)×R).
[0069] Deep neural network module: Use the statistical histogram output of the DoA histogram statistics module as the feature map input, and process it through the deep neural network. Through the multiple output branches of the network, the effective signal of the interference source and the effective signal of the target source and their specific direction information (i.e., angle value) can be obtained. The effective signal of the target source and the effective signal of the interference source take the values of 0 and 1, 0 represents that the current target source or interference source does not exist, and 1 represents that the current target source or interference source exists. The direction information refers to the angle value of the sound source in the reference sensor coordinate system, which ranges from 0 degrees to 360 degrees. For example, Figure 5 In the four-sensor array shown, the orientation information describes the angle θ between the sound source and the line connecting Sensors 1 and 2. When training a deep neural network, the training data includes four categories: the presence of both interference and target sources, the presence of only interference sources, the presence of only target sources, and the absence of neither interference nor target sources. The neural network training labels contain 0 / 1 encodings of the valid signals of the interference and target sources, as well as one-hot codes of the interference and target source angles. The length of the one-hot code is the same as the length of the histogram statistics register, H. Loss functions used in training include cross-entropy loss and binary cross-entropy loss. The neural network models used for training and the neural network circuits deployed can be various neural network model structures, including DNN, CNN, RNN, CRNN, MLP, Transform, LSTM, GRU, and others.
[0070] After obtaining the interference source and target source position and validity information, it is sent together with the multi-channel STFT results extracted by the STFT feature extraction circuit into the multi-stage noise reduction and signal enhancement circuit based on multi-source position information. The structure of the multi-stage noise reduction and signal enhancement circuit based on multi-source position information is as follows: Figure 4 As shown in the figure, it includes an array-based frequency domain beamforming (AFDBF) module, a neural network-based beamforming (NBBF) module, a multilevel fusion for beamforming (MFBF) module, a control circuit module, and a vector coding module for vector coding of effective azimuth angles. Its specific implementation and principles are as follows:
[0071] AFDBF module: For N-channel STFT input signals (N≥2) extracted by the STFT feature extraction circuit, the AFDBF module has N-1 AFDBF processing units corresponding to N-1 STFT outputs. One of the N-channel signals is a reference signal. Here, channel 0 is used as the reference signal, but other channels can also be used. When connecting to the external circuit, the interface can be swapped. Each AFDBF processing unit inputs and processes a set of dual-channel signals. A set of dual-channel signals consists of a reference signal and a channel signal that is different from that of other processing units.
[0072] A single AFDBF unit consists of a BF (Beamforming) coefficient storage module and an arithmetic processing circuit. The BF coefficient storage module is used to store noise reduction and enhancement coefficients related to frequency, interference source, and target source direction. The arithmetic circuit obtains coefficients based on the direction information of the target source or interference source, and performs multiplication and addition enhancement or multiplication and subtraction noise reduction processing based on the existence information of the target source or interference source. The arithmetic processing circuit includes a multiplier, an adder, and a subtractor. The specific function of the AFDBF unit is to first obtain the interference source direction information θ n , target source orientation information θ d And the corresponding FFT frequency ω i As an index, the frequency domain coefficient F related to the direction and frequency of the source signal, which is pre-generated and stored in the BF coefficient storage module, is taken out. The frequency domain coefficient is used to compensate for the difference in the target signal between the two channels (complex multiplication). The sum of the two channel signals (Yd) can enhance the target signal in the channel; in addition, the frequency domain coefficient is used to compensate for the difference in the interference signal between the two channels (complex multiplication). The difference (Yn) of the two channel signals can cancel the interference signal in the channel to achieve the purpose of noise reduction; finally, the enhanced signal and the noise reduction signal are selectively added according to the effectiveness of the source signal to obtain the adaptive noise reduction and enhanced signal Y. It can be expressed in the form of:
[0073] Yd(c,ω i )=X0(ω i )+X(c,ω i )F(c,ω i ,θ d )
[0074] Yn(c,ω i )=X0(ω i )-X(c,ω i )F(c,ω i ,θ n )
[0075]
[0076] Wherein, X0(*) represents the FFT result of a frame of sensor array timing signals as a reference channel input by the STFT feature extraction circuit, X(*) represents the FFT result of a frame of sensor array timing signals of other channels input by the STFT feature extraction circuit, F(*) represents the frequency domain coefficient used to compensate for the difference between the target source or interference source channels, Yd(*) represents the processing result of enhancing the target signal between channels, Yn(*) represents the processing result of suppressing the interference signal between channels, and Y(*) represents the final output result of the AFDBF processing unit; ω i represents the digital frequency index, ω i =2πf i , f i is the physical frequency point in FFT transformation, Among them, f s is the digital signal sampling rate, Num represents the number of FFT points, and the value range of i is θ d represents the target source azimuth from the multi-source positioning circuit, θ n It represents the azimuth of the interference source from the multi-source positioning circuit, d represents the effective signal of the target source, which is 1 if it is effective and 0 if it is invalid, n represents the effective signal of the interference source, which is 1 if it is effective and 0 if it is invalid, and c represents the STFT channel index, which is also the AFDBF processing unit index, and its value range is [1, N-1].
[0077] All frequency-domain coefficients can be generated off-chip and stored in the BF coefficient storage module. Generation methods include those based on differential sensor arrays and neural network training. Depending on the control circuit configuration, the value set of the array (d, n) is {(1, 1), (1, 0), (0, 1), (0, 0)}. In the (0, 0) state, meaning that no interference source or target source exists, the AFDBF module does not operate.
[0078] NBBF module: Receives the data output from the STFT feature extraction circuit and the vector encoding circuit, splices them in sequence as the feature map data of the neural network input, and sends them to the pre-trained neural network beamforming module to extract the time-frequency mask M with the target source enhancement and interference source noise reduction effect, that is, the neural network output result, and outputs N-1 groups of complex time-frequency masks M (c, ω i ), whose size is consistent with the AFDBF output size.
[0079] The NBBF neural network is trained using the timing signal training data of the original sensor array. The timing signal training data of the original sensor array consists of three parts: the target source signal, the interference source signal, and the background timing signal. It contains at least one type of target source or interference source signal. The signal-to-noise ratio distribution range of the target source signal in the training data can be adjusted according to the requirements of the actual application scenario. The training goal is to perform element-wise multiplication processing on the output mask signal and the AFDBF output signal to attenuate the interference source signal in the original signal or enhance the target signal in the original signal. The training adopts regression loss functions, including L1 loss and L2 loss. The neural network models used for training and the deployed neural network models can be various neural network model structures, including DNN, CNN, RNN, CRNN, MLP, Transform, LSTM, GRU, etc.
[0080] MFBF module: After receiving the STFT result Y processed by the AFDBF module and the time-frequency mask result M generated by the NBBF module, the two are effectively fused by element-level multiplication to achieve more effective interference signal noise reduction and target signal enhancement processing. Then, the results of the fusion of N-1 channels are element-level added and fused to further improve the signal-to-noise ratio of the target signal. Finally, the STFT result R after single-channel noise reduction and enhancement processing is output. The formula is expressed as:
[0081]
[0082] Control circuit module: Based on the values of the effective signal of the interference source and the effective signal of the target source, it can be divided into four circuit processing modes under the control of the control circuit, as follows:
[0083] Target source exists & interference source exists
[0084] The control circuit switches switch S0 to the left path, SS is in the off state, and the multi-channel STFT is sent to the AFDBF module and NBBF module. In addition, the control circuit closes switches S1, S2, S3, S4, and S6 in the AFDBF module, while switches S5 and S7 are in the off state, so as to simultaneously enhance the target signal and suppress the interference signal in the multi-channel STFT signal. Finally, switch S8 is also switched to the left path to receive and output the STFT result after multi-stage noise reduction and enhancement processing. The output result is expressed as:
[0085]
[0086] Target source exists & interference source does not exist
[0087] The control circuit switches switch S0 to the left path, SS is in the off state, and the multi-channel STFT is sent to the AFDBF module and NBBF module. In addition, the control circuit closes switches S2 and S7 in the AFDBF module, while switches S1, S3, S4, S5 and S6 are in the off state, so as to simultaneously enhance the target signal and suppress the interference signal in the multi-channel STFT signal. Finally, switch S8 is also switched to the left path to receive and output the STFT result after multi-stage noise reduction and enhancement processing. The output result is expressed as:
[0088]
[0089] The target source does not exist and the interference source exists
[0090] The control circuit switches switch S0 to the left path, SS is in the off state, and the multi-channel STFT is sent to the AFDBF module and NBBF module. In addition, the control circuit closes the S1 and S5 switches in the AFDBF module, while switches S2, S3, S4, S6 and S7 are in the off state, so as to simultaneously enhance the target signal and suppress the interference signal in the multi-channel STFT signal. Finally, switch S8 is also switched to the left path to receive and output the STFT result after multi-stage noise reduction and enhancement processing. The output result is expressed as:
[0091]
[0092] The target source does not exist & the interference source does not exist
[0093] The control circuit switches switch S0 to the right path. Under the control of the control circuit, SS selectively closes to filter out the STFT results of the reference channel in the multi-channel. Switches S1-S7 in the AFDBF module are all in the open state. Switch S8 is switched to the right path to receive the STFT results after single-channel extraction. The output result is expressed as:
[0094] R(ω i )=X0(ω i )
[0095] Vector encoding module: The vector encoding module receives the scalar information of the interference source azimuth and the target source azimuth from the multi-source positioning circuit that obtains the interference source and target source azimuth information, as well as the control signal from the control circuit to determine whether the vector encoding of the multi-source azimuth needs to be updated and whether it needs to be set to an invalid vector encoding state (when the corresponding source does not exist). The expression of the vector encoding V is:
[0096] Vd(ω i )=Vec d (ω i ,θd )
[0097] Vn(ω i )=Vec n (ω i ,θ n )
[0098] Among them, Vd(*) represents the vector code of the target source, Vn(*) represents the vector code of the interference source, and Vec d (*) represents the encoding function (circuit) of the target source, Vec n (*) represents the coding function (circuit) of the interference source, ω i represents the digital frequency index, ω i =2πf i , f i is the physical frequency point in FFT transformation, Among them, f s is the digital signal sampling rate, Num represents the number of FFT points, and the value range of i is θ d represents the target source azimuth from the multi-source positioning circuit, θ n It represents the azimuth of the interference source from the multi-source positioning circuit. The encoding function of the target source and the encoding function of the interference source can be the same. The circuit can be implemented using a look-up table (LUT), a register group or a static random-access memory (SRAM). The encoding methods include sine encoding, cosine encoding, sine-cosine encoding, etc. The sine-cosine encoding formula is given here as:
[0099] Vd(ω i )=cos(ω i θ d )
[0100] Vn(ω i )=sin(ω i θ n )
[0101] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A sensor array-based intelligent noise reduction and signal enhancement circuit for time series signals, characterized in that: The circuit includes: an STFT feature extraction circuit, a multi-source positioning circuit for obtaining interference source and target source azimuth information, and a multi-stage noise reduction and signal enhancement circuit based on the multi-source azimuth information. The STFT feature extraction circuit is used to convert the input sensor array time series signal into the time-frequency domain to obtain its time-frequency signal; the multi-source positioning circuit for obtaining interference source and target source azimuth information is used to process the time-frequency signal processed by the STFT feature extraction circuit, and obtain the interference source and target source azimuth of the currently processed array time series signal, that is, the angle value after the processing; the multi-stage noise reduction and signal enhancement circuit based on the multi-source azimuth information is used to process the time-frequency signal processed by the STFT feature extraction circuit in combination with the obtained interference source and target source azimuth information, so as to suppress the interference signal and enhance the target signal; The multi-source positioning circuit for obtaining the azimuth information of interference sources and target sources includes a complex phase difference calculation module, an inverse covariance matrix multiplication module, a DoA angle prediction module, a DoA histogram statistics module, and a deep neural network module; The complex phase difference calculation module uses the Q-order Cordic algorithm to calculate the phase of the complex signal and the phase difference between the reference sensor channel signal and the remaining sensor channel signals. The number of sensors N is greater than or equal to 2, and the order Q of the Cordic algorithm is selected according to actual needs. The phase difference calculation formula is: Where i∈[1, N-1], t represents time, f represents frequency, X0(t, f) is the STFT result of the reference sensor channel, X i (t, f) are the STFT results of the remaining sensor channels, Re represents the real part, and Im represents the imaginary part; After obtaining the phase signal output by the complex phase difference calculation module, the inverse covariance matrix multiplication module converts the phase signal Multiplying with the inverse covariance matrix A, we can calculate the trigonometric function vector S of the sound source. Where α is the angle of the target or interference sound source, and the inverse covariance matrix is determined by the current sensor array; The DoA angle prediction module uses the K-order Cordic algorithm to calculate the arctangent value of the trigonometric function vector output by the inverse covariance matrix module to obtain the sound source angle at each time-frequency point. The DoA angle prediction formula is as follows: Among them, DoA pred is the DoA angle prediction value, sinα and cosα are the elements in the trigonometric function vector; The DoA histogram statistics module consists of a judgment logic circuit and an angle distribution histogram statistics register group. The DoA histogram statistics module collects statistics on the sound source angle signal output by the DoA angle prediction module and stores it in the histogram statistics register group. The deep neural network module uses the statistical histogram output of the DoA histogram statistics module as the feature map input, which is processed by the deep neural network. The effective signals of the interference source and the target source as well as their specific orientation information, i.e., the angle value, are obtained through multiple output branches of the network.
2. The circuit according to claim 1, wherein: The working process of the intelligent noise reduction and signal enhancement circuit of the timing signal based on the sensor array is as follows: first, the parameters required for the multi-stage noise reduction and signal enhancement circuit based on multi-source azimuth information are initialized and stored in the corresponding storage circuit module; after detecting the multi-source positioning trigger signal sent by the external controller, the STFT result of the sensor array timing signal obtained by the STFT feature extraction circuit is sent to the multi-source positioning circuit that obtains the azimuth information of the interference source and the target source for array signal processing; the interference source and the target source are located using the direction of arrival (DOA) estimation scheme based on the neural network. The positioning result includes two dimensions of information, namely, whether there is an interference source and a target source and their corresponding specific azimuth information; then the STFT result and the DOA result are sent to the multi-stage noise reduction and signal enhancement circuit based on multi-source azimuth information for noise reduction and signal enhancement processing.
3. The circuit according to claim 1, wherein: The STFT feature extraction circuit includes a pre-emphasis module, a framing module, a windowing module, and a fast Fourier transform module. The pre-emphasis module processes the sensor array timing signal and sends the processed signal to the framing module to compensate for the loss of high-frequency components of the original timing signal. The specific processing method is: y(t) = x(t) - μx(t-1), where y(t), x(t), and x(t-1) represent the output signal after pre-emphasis processing at time t, the original input signal, and the input signal at the previous time, respectively, and μ is the pre-emphasis coefficient. The framing module frames the sensor array timing signal output by the pre-emphasis module according to the point number requirement of the fast Fourier transform module and sends the framing results to the windowing module in sequence. The windowing module multiplies the signal output by the framing module by the windowing coefficient and sends the result to the fast Fourier transform module. The fast Fourier transform module performs fast Fourier transform processing on the signal output by the windowing module and outputs the STFT result of the sensor array timing signal.
4. The circuit according to claim 1, wherein: The multi-stage noise reduction and signal enhancement circuit based on multi-source azimuth information includes an array-based frequency domain beamforming AFDBF module, a neural network-based beamforming NBBF module, a beamforming multi-stage fusion MFBF module, a control circuit module, and a vector coding module for vector coding of effective azimuth angles. For the N-channel STFT input signal extracted by the STFT feature extraction circuit, the AFDBF module has N-1 AFDBF processing units corresponding to N-1 STFT outputs. One of the N-channel signals is a reference signal. A single AFDBF unit consists of a BF coefficient storage module and an arithmetic processing circuit. The BF coefficient storage module is used to store noise reduction and enhancement coefficients related to frequency, interference source, and target source orientation. The arithmetic circuit obtains coefficients based on the orientation information of the target source or interference source, and performs multiplication-addition enhancement or multiplication-subtraction noise reduction processing based on the existence information of the target source or interference source. The NBBF module receives the data output from the STFT feature extraction circuit and the vector encoding circuit, and sequentially splices them as the feature map data of the neural network input, and sends them to the pre-trained neural network beamforming module to extract the time-frequency mask M with the target source enhancement and interference source noise reduction effect, that is, the neural network output result, and outputs N-1 groups of complex time-frequency masks M(c,ω i ), whose size is consistent with the AFDBF output size; After receiving the STFT result Y processed by the AFDBF module and the time-frequency mask result M generated by the NBBF module, the MFBF module performs element-level multiplication fusion on the two, and then performs element-level addition fusion on the fused results of N-1 channels, and finally outputs the STFT result R after single-channel denoising and enhancement processing; The control circuit module controls the four processing modes of the multi-stage noise reduction and signal enhancement circuit of multi-source azimuth information according to the values of the effective signal of the interference source and the effective signal of the target source; The vector coding module receives scalar information of the interference source azimuth and the target source azimuth from the multi-source positioning circuit that obtains the interference source and target source azimuth information, as well as a control signal from the control circuit to determine whether the vector coding of the multi-source azimuth needs to be updated and whether it needs to be set to an invalid vector coding state.
5. The circuit according to claim 4, characterized in that The specific functions of the AFDBF unit are: First, the obtained interference source azimuth information θ n , target source orientation information θ d And the corresponding FFT frequency ω i As the index, the frequency domain coefficient F related to the direction and frequency of the source signal, which is pre-generated and stored in the BF coefficient storage module, is taken out. The frequency domain coefficient is used to compensate for the difference in the target signal between the two channels by complex multiplication. Then, the frequency domain coefficient is used to compensate for the difference in the interference signal between the two channels by complex multiplication. Finally, the enhanced signal and the noise reduction signal are selectively added according to the effectiveness of the source signal to obtain the adaptive noise reduction and enhanced signal Y. The expression is: Yd(c,ω i )=X0(ω i )+X(c,ω i )F(c,ω i ,i d ) In (c,ω i )=X0(ω i )-X(c,ω i )F(c,ω i ,θ n ) Where X0(*) represents the FFT result of a frame of sensor array timing signals as a reference channel input by the STFT feature extraction circuit, X(*) represents the FFT result of a frame of sensor array timing signals of other channels input by the STFT feature extraction circuit, F(*) represents the frequency domain coefficient used to compensate for the difference between the target source or interference source channels, Yd(*) represents the processing result of enhancing the target signal between channels, Yn(*) represents the processing result of suppressing the interference signal between channels, and Y(*) represents the final output result of the AFDBF processing unit. i represents the digital frequency index, ω i =2πf i , f i is the physical frequency point in FFT transformation, Among them, f s is the digital signal sampling rate, Num represents the number of FFT points, and the value range of i is θ d represents the target source azimuth from the multi-source positioning circuit, θ n It represents the azimuth of the interference source from the multi-source positioning circuit, d represents the effective signal of the target source, which is 1 if it is effective and 0 if it is invalid, n represents the effective signal of the interference source, which is 1 if it is effective and 0 if it is invalid, and c represents the STFT channel index, which is also the AFDBF processing unit index, and its value range is [1, N-1].
6. The circuit according to claim 4, characterized in that According to the interference source and target source information, the multi-stage noise reduction and signal enhancement circuit of multi-source azimuth information is divided into four processing modes. The first is that when the interference source and target source exist at the same time, multi-stage noise reduction is implemented on the interference source, multi-stage enhancement is implemented on the target source, and the processed STFT result is output; the second is that when only the interference source exists, multi-stage noise reduction is implemented on the interference source, and the processed STFT result is output; Third, if only the target source exists, multi-level enhancement is implemented on the target source, and the processed STFT result is output; fourth, if neither the interference source nor the target source exists, the STFT result of the reference sensor channel timing signal in the original multi-sensor channel is directly output.
Citation Information
Patent Citations
Speech enhancement method and device based on dual-channel neural network time-frequency masking, and hearing-aid equipment
CN114078481A
Multi-sound-source direction-finding positioning method based on vector microphone in strong reverberation environment
CN115754898A