Underwater noise recognition method based on continuous wavelet transform and improved residual neural network
The time-frequency characteristics of underwater noise signals are extracted through variational mode decomposition and continuous wavelet transformation, combined with the improved residual neural network, the problem of insufficient underwater noise signal characteristic information is solved, the recognition accuracy is improved and the accuracy of the deep network is avoided.
Patent Information
- Application Number
- CN202310573352.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-20
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2043-05-20
AI Technical Summary
In the prior art, the one-dimensional representation of the underwater noise signal contains less characteristic information, resulting in low noise recognition accuracy and the problem of degradation of the accuracy of the deep neural network has not been effectively solved.
The variational modal decomposition algorithm is used to decompose the noise signal, calculate the Pearson correlation coefficient of the modal component and reconstruct the signal, combine the continuous wavelet transform to extract the time-frequency characteristics, and build an improved residual neural network model, and train it using the channel and spatial attention mechanism.
It effectively removes the redundant information of noise signals, improves the accuracy of noise recognition, avoids the reduction in the accuracy of deep networks, and realizes accurate identification of different types of noises.
Smart Images

Figure CN116740387B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of signal processing and deep learning, and in particular to an underwater noise recognition method based on continuous wavelet transform and improved residual neural network. Background Art
[0002] The ocean will be one of the primary battlefields of future high-tech warfare, and future naval equipment will inevitably move towards information technology. With the increasing development of electronic countermeasures (ECM) technology, its extensive use in warfare demonstrates its importance in naval warfare. Underwater acoustic countermeasures (UAS) utilize signal processing techniques to process underwater acoustic signals, enabling military tasks such as target detection, parameter estimation, and target identification. Underwater target detection is the foundation of ASA. Accurately detecting enemy targets is crucial for achieving numerous subsequent tasks.
[0003] The strength of an underwater vehicle's radiated noise and self-noise determines its stealth performance. Higher radiated and self-noise levels make it more susceptible to detection and tracking by opposing sonar systems, while also significantly interfering with its own sonar systems. Radiated and self-noise from underwater vehicles include vibration noise from propulsion systems such as propellers, propeller noise from various mechanical devices onboard the vehicle, and hydrodynamic noise. Vibration noise, specifically, refers to the noise radiated by the vibrations of various mechanisms onboard the vehicle. This noise severely impacts sonar equipment and is the primary source of vehicle noise. Therefore, research is needed to identify and separate vibration noise from underwater vehicles.
[0004] Recent research has shown that the noise radiation process of underwater targets such as surface ships and submarines is a non-Gaussian, nonlinear, and non-stationary process (the "three nons"). Combined with vibration noise processing methods, vibration noise feature extraction can be summarized into the following aspects:
[0005] (1) Time domain waveform feature extraction
[0006] The time-domain waveform structure of the noise radiated by underwater acoustic targets, such as ships, contains rich target characteristic information. From the time domain, characteristic parameters reflecting the waveform structure can be directly extracted. For example, some scholars have studied characteristics such as the mean, variance, peak index, kurtosis, and skewness of underwater acoustic time-domain signals. Combining FFT power spectrum and STFT analysis, they have constructed waveform indicators in the frequency domain as underwater acoustic signal features. These have been applied to transient sonar signal classification with excellent results. Some domestic scholars have studied the direct extraction of waveform structural features such as zero-crossing point distribution, peak-to-peak amplitude distribution, wavelength difference distribution, and wave train area distribution of target signals.
[0007] (2) Power spectrum, line spectrum and modulation spectrum feature extraction
[0008] Power spectrum estimation is a fundamental method for obtaining the second-order statistical characteristics of underwater acoustic signals. Transforming from the time domain to the frequency domain converts complex waveforms in the time domain into a relatively simple single frequency component distribution in the frequency domain. Therefore, low-frequency line spectrum features and wide-band spectrum characteristics in the power spectrum become effective features for underwater acoustic target detection and recognition. Furthermore, propeller noise is a major source of noise for underwater acoustic targets such as surface ships and submarines. Its cavitation noise often produces amplitude or frequency modulation. Demodulated modulation spectra contain numerous discrete line spectrum components corresponding to the propeller shaft frequency, blade frequency, and their harmonics. Utilizing these frequency modulation features can provide an effective basis for passive target detection. High-order spectra are capable of effectively processing non-Gaussian and non-stationary signals and suppressing Gaussian and non-Gaussian colored noise. Therefore, they can also be used to extract features from underwater acoustic signals such as ship noise.
[0009] (3) Auditory feature extraction
[0010] Since 2001, international researchers have conducted extensive research on extracting auditory features from noise radiated by underwater targets. Based on the human hearing mechanism, Wang Yang et al. extracted auditory spectral features, speech features, and psychological parameter characteristics of underwater targets, achieving a series of research results. Li Zhaohui et al. conducted in-depth research on auditory models, combining them with the characteristics of underwater acoustic signals to demonstrate the model's applicability to underwater acoustics.
[0011] Currently, most research on underwater noise processing relies on one-dimensional representations of the signal. However, one-dimensional signals contain less feature information than two-dimensional images, making them ineffective in extracting important features of noise signals. This, in turn, hinders the neural network's ability to learn these features. Furthermore, optimizing the neural network is crucial for improving recognition accuracy. Summary of the Invention
[0012] In order to effectively solve the problem that one-dimensional underwater noise signals contain few features and have low noise recognition accuracy, the present invention proposes an underwater noise recognition method based on continuous wavelet transform and improved residual neural network. The underwater noise is reconstructed using variational mode decomposition and correlation principle, and the noise features are extracted using continuous wavelet transform. Finally, a noise recognition model is constructed to achieve accurate recognition of different noises.
[0013] To achieve the above object, the present invention proposes an underwater noise recognition method based on continuous wavelet transform and improved residual neural network, the method comprising the following steps:
[0014] Step 1: Decompose the original sample noise signal using the variational mode decomposition algorithm to obtain several modal components;
[0015] Step 2: Calculate the Pearson correlation coefficient between each modal component and the original sample noise signal, and add the modal components whose correlation coefficient is greater than the set threshold K to obtain the reconstructed noise signal;
[0016] Step 3: Perform continuous wavelet transform on the reconstructed noise signal to extract its features in the time domain and frequency domain, and obtain a two-dimensional time-frequency image containing the time-frequency features of the noise signal;
[0017] Step 4: Improve the residual neural network using channel and spatial attention mechanisms to build a noise recognition model;
[0018] Step 5: Using the two-dimensional time-frequency diagram of the noise signal obtained in step 3 and the identification label corresponding to the original sample noise signal, the noise recognition model constructed in step 4 is trained to obtain a trained noise recognition model;
[0019] Step 6: Process the noise signal actually collected according to the noise signal processing method in steps 1 to 3 to obtain a two-dimensional time-frequency diagram of the noise signal actually collected; input the two-dimensional time-frequency diagram of the noise signal actually collected into the trained noise recognition model for recognition.
[0020] Furthermore, the specific process of decomposing the noise signal using the variational mode decomposition algorithm in step 1 is as follows:
[0021] Step 1.1: Assume that the original signal f is decomposed into K modal components and construct the constrained variational model as follows:
[0022]
[0023] where u k ={u1, u2, ..., u k} is the set of modal functions, ω k ={ω1,ω2,…,ω k} is the set of center frequencies of each mode, u k (t) is the narrowband mode function, K is the number of modes, is the sum of all modal components, To adjust the exponential term of the center frequency of each modal component, δ(t) is the unit pulse function, for u k (t) Partial derivative with respect to t, * represents convolution operation;
[0024] Step 1.2: Introduce the augmented Lagrange function to solve the constrained variational model in step 1 and obtain u k The best choice for (t):
[0025]
[0026] Where α is the quadratic penalty factor and λ(t) is the Lagrange multiplier.
[0027] Furthermore, the specific process of step 2 is:
[0028] Calculate the Pearson correlation coefficient between each modal component and the original sample noise signal:
[0029]
[0030] in is the standard deviation of the modal component, is the standard deviation of the original sample noise signal; the modal components with correlation coefficients greater than K are summed to obtain the reconstructed noise signal f′(t).
[0031] Furthermore, in step 2, the threshold K is set to 0.3.
[0032] Furthermore, the specific process of step 3 is:
[0033] According to the waveform of the noise signal in the time domain, a wavelet with a similar waveform is selected as the wavelet basis function; for the reconstructed noise signal f′(t), its continuous wavelet transform is:
[0034]
[0035] Where ψ(t) is the wavelet basis function, It is the complex conjugate of ψ(t), a is the scale factor, which controls the lateral expansion and contraction of the wavelet basis function, and b is the time shift factor, which controls the movement of the wavelet basis function on the time axis;
[0036] The wavelet coefficients at different frequencies are combined and represented as an RGB three-channel image to obtain a two-dimensional time-frequency image containing the time-frequency characteristics of the noise signal.
[0037] Furthermore, Morlet wavelet is adopted as the wavelet basis function.
[0038] Furthermore, in step 4, the noise recognition model processes the input signal as follows:
[0039] The input signal is converted from the original RGB three channels to 64 channels after the first convolution layer. At this time, the data dimension is 32×64×112×112. The 64-channel data is input to the channel attention module. In the channel attention module, the input is globally pooled:
[0040] Calculate the average value of all pixel values in the channel as the eigenvalue of the channel, and output it as a 1×1×64 tensor; then pad the tensor with a size of 3; perform a one-dimensional convolution on the padded tensor using a convolution kernel of size 7; compress the convolution result to between (0, 1) using the sigmoid function and reconstruct it into a 32×64×1×1 tensor, which is then multiplied by the input signal to obtain the output data with channel attention weights;
[0041] The output data of the channel attention module is input into the spatial attention module; in the spatial attention module, maximum pooling and average pooling are performed in the plane dimension to obtain two 112×112×1 feature maps, which are then concatenated into a 112×112×2 feature map in the channel dimension. This map is then reduced to one channel through a 7×7 convolutional layer, and the sigmoid function is used to generate the spatial weight coefficient and multiplied by the input of the spatial attention module to obtain the final output.
[0042] Furthermore, in step 5, the process of training the noise recognition model is as follows:
[0043] The two-dimensional time-frequency diagram of the original sample noise signal is divided into a training set and a validation set. The two-dimensional time-frequency diagram of the samples in the training set is rotated 90° and 180°, and the enhanced samples are added to the training set.
[0044] The noise recognition model was trained using the training set, using the Adam optimizer, setting the learning rate to 0.0001, and the BatchSize to 32. The training set was input into the neural network for 100 rounds of training. After each round of training, the validation set was tested, and the parameters of the round with the highest validation set accuracy during the training process were saved.
[0045] Beneficial effects
[0046] The present invention first performs variational modal decomposition on the underwater noise signal to obtain several modal components. Then, the Pearson correlation coefficient between each modal component and the original signal is calculated, and the modal components with coefficients greater than K are summed to obtain the reconstructed signal. The reconstructed signal is subjected to a continuous wavelet transform to extract the time-frequency characteristics of the noise signal and obtain a two-dimensional time-frequency graph of the noise signal. The residual neural network is improved using an attention mechanism to construct a recognition model based on deep learning. The improved neural network can better learn the characteristics of the noise signal during training, ignore non-important information, and improve the training efficiency and recognition accuracy.
[0047] Aiming at the characteristic that noise signals are accompanied by a large amount of redundant information, the present invention proposes to use variational mode decomposition method to reconstruct the noise signals, which has a significant effect of removing redundancy.
[0048] To address the problem that one-dimensional signals contain less information, the present invention proposes using a continuous wavelet transform method to obtain a two-dimensional time-frequency diagram containing the time-frequency characteristics of the noise signal, which can effectively extract the characteristics of the noise signal.
[0049] The recognition model in the present invention includes residual skip connections, which can avoid the accuracy degradation problem of deep networks to a certain extent and achieve effective recognition of noise types.
[0050] Additional aspects and advantages of the present invention will be set forth in part in the description which follows and, in part, will be obvious from the description which follows, or may be learned by practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments with reference to the accompanying drawings, in which:
[0052] Figure 1 This is the overall steps of an underwater noise recognition method based on continuous wavelet transform and improved residual neural network of the present invention;
[0053] Figure 2 Three noise signal waveforms of an underwater noise recognition method based on continuous wavelet transform and improved residual neural network according to the present invention;
[0054] Figure 3 The modal components obtained by using the underwater noise identification method based on continuous wavelet transform and improved residual neural network of the present invention; (a) return vibration noise signal, (b) submerged vibration noise signal, (c) upwelling noise signal;
[0055] Figure 4 Two-dimensional time-frequency images of noise signals obtained by using the underwater noise recognition method based on continuous wavelet transform and improved residual neural network of the present invention; (a) return vibration noise signal, (b) submerged vibration noise signal, (c) upwelling noise signal;
[0056] Figure 5 This is a network architecture in an underwater noise recognition method based on continuous wavelet transform and improved residual neural network of the present invention;
[0057] Figure 6 This is a schematic diagram of the attention module in the underwater noise recognition method based on continuous wavelet transform and improved residual neural network of the present invention. DETAILED DESCRIPTION
[0058] The following describes in detail embodiments of the present invention. The embodiments are exemplary and intended to explain the present invention, but are not to be construed as limiting the present invention.
[0059] like Figure 1As shown, the underwater noise recognition method based on continuous wavelet transform and improved residual neural network in this embodiment includes the following steps:
[0060] Step 1: Use the variational mode decomposition algorithm to decompose the original sample noise signal to obtain several modal components.
[0061] The specific process of decomposing noise signals using the variational mode decomposition algorithm consists of two parts: constructing the variational problem and solving the variational problem:
[0062] Step 1.1: Assume that the original signal f is decomposed into K modal components, ensuring that the decomposition sequence is a modal component with a finite bandwidth and a center frequency, and that the sum of the estimated bandwidths of each mode is minimized. The constraint is that the sum of all modes is equal to the original signal. The constrained variational model is constructed as follows:
[0063]
[0064] where u k ={u l ,u2,…,u k ] is the set of modal functions, ω k ={ω1,ω2…,ω k} is the set of center frequencies of each mode, u k (t) is the narrowband mode function, K is the number of modes, is the sum of all modal components, To adjust the exponential term of the center frequency of each modal component, δ(t) is the unit pulse function, for u k (t) Partial derivative with respect to t, * represents convolution operation;
[0065] Step 1.2: Introduce the augmented Lagrange function to solve the constrained variational model in step 1 and obtain u k The best choice for (t):
[0066]
[0067] Where α is the quadratic penalty factor and λ(t) is the Lagrange multiplier.
[0068] Step 2: Calculate the Pearson correlation coefficient between each modal component and the original sample noise signal, and add the modal components with correlation coefficients greater than a set threshold K=0.3 to obtain a reconstructed noise signal.
[0069] Calculate the Pearson correlation coefficient between each modal component and the original sample noise signal:
[0070]
[0071] in is the standard deviation of the modal component, is the standard deviation of the original sample noise signal; the modal components with correlation coefficients greater than 0.3 are summed to obtain the reconstructed noise signal f′(t).
[0072] Step 3: Perform continuous wavelet transform on the reconstructed noise signal to extract its features in the time domain and frequency domain to obtain a two-dimensional time-frequency image containing the time-frequency features of the noise signal.
[0073] According to the waveform of the noise signal in the time domain, the Morlet wavelet with a similar waveform is selected as the wavelet basis function; the mathematical expression of the Morlet wavelet is:
[0074]
[0075] In the time domain, the mother wavelet is moved on the time axis, and the wavelet coefficients are obtained by calculating the convolution of the wavelet function and the window signal. The larger the wavelet coefficient, the better the fit between the wavelet and the signal segment. In the frequency domain, the length and frequency of the wavelet are changed by scaling the length of the wavelet, and the wavelet coefficients at different frequencies are obtained by convolution with the window signal. The translation and scaling of the wavelet basis function can be expressed as the following formula:
[0076]
[0077] Among them, a is the scale factor, which controls the horizontal expansion and contraction of the wavelet basis function, and b is the time shift factor, which controls the movement of the wavelet basis function on the time axis. 2 (R), its continuous wavelet transform is defined as follows:
[0078]
[0079] Where ψ(t) is the wavelet basis function, is the complex conjugate of ψ(t); the wavelet coefficients at different frequencies are combined and represented as an RGB three-channel image to obtain a two-dimensional time-frequency image containing the time-frequency characteristics of the noise signal.
[0080] Step 4: Improve the residual neural network using channel and spatial attention mechanisms to build a noise recognition model;
[0081] Construct a residual neural network based on the attention mechanism. The network structure is shown in the following table.
[0082]
[0083] Channel and spatial attention mechanisms are added after the first convolutional layer of the neural network. First, the underwater noise signal data passes through the first convolutional layer, converting the original RGB three-channel data into 64 channels, resulting in a data dimension of 32×64×112×112. These 64-channel data are input to the channel attention module. The channel attention module performs global pooling on the input. Specifically, this module calculates the average value of all pixel values in the channel, which serves as the feature value for that channel. The output is a 1×1×64 tensor. To maintain the shape of the input data, it is padded with a size of 3. This tensor is then convolved one-dimensionally with a convolution kernel of size 7 to capture information between different channels. The convolution result is compressed to the range (0, 1) using a sigmoid function and reconstructed into a 32×64×1×1 tensor. This is then multiplied by the input x to produce the output data with channel attention weights. The output of the channel attention module is then input to the spatial attention module. In the spatial attention module, maximum pooling and average pooling are performed in the plane dimension to obtain two 112×112×1 feature maps. These two feature maps are concatenated in the channel dimension to form a 112×112×2 feature map. A 7×7 convolution layer is used to reduce it to one channel. The sigmoid function is then used to generate the spatial weight coefficient and multiply it by the input to obtain the final output.
[0084] Step 5: Using the two-dimensional time-frequency diagram of the noise signal obtained in step 3 and the identification label corresponding to the original sample noise signal, the noise recognition model constructed in step 4 is trained to obtain a trained noise recognition model;
[0085] The time-frequency graph samples of the noise signal were divided into a training set and a validation set. The training set samples were rotated 90° and 180° to obtain enhanced data samples. The neural network used the Adam optimizer with a learning rate of 0.0001 and a batch size of 32. The training set was fed into the neural network for 100 rounds of training. After each round of training, the validation set was tested, and the parameters of the round with the highest validation set accuracy were saved.
[0086] Step 6: Use a UUV to collect actual noise signals at the Danjiangkou Reservoir. Process the noise signals in the same way as in Steps 1 to 3 to obtain a two-dimensional time-frequency diagram of the noise signals. Input the two-dimensional time-frequency diagram of the noise signals into the trained noise recognition model for recognition.
[0087] like Figures 2 to 6As shown, through the method of the present invention, the noise signal containing a large amount of redundant information is filtered, the two-dimensional time-frequency diagram of the noise signal is obtained by continuous wavelet transform, and the underwater noise is classified using a recognition model based on deep learning.
[0088] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are illustrative and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention without departing from the principles and purpose of the present invention.
Claims
1. A method for underwater noise recognition based on continuous wavelet transform and improved residual neural network, characterized by: The following steps are involved: Step 1: Decompose the original sample noise signal using the variational mode decomposition algorithm to obtain several modal components; Step 2: Calculate the Pearson correlation coefficient between each modal component and the original sample noise signal, and add the modal components whose correlation coefficients are greater than a set threshold K to obtain a reconstructed noise signal; Step 3: Perform continuous wavelet transform on the reconstructed noise signal to extract its features in the time domain and frequency domain, and obtain a two-dimensional time-frequency image containing the time-frequency features of the noise signal; Step 4: Improve the residual neural network using channel and spatial attention mechanisms to build a noise recognition model; The noise recognition model processes the input signal as follows: The input signal is converted from the original RGB three channels to 64 channels after the first convolution layer. At this time, the data dimension is ; Input the 64-channel data into the channel attention module; Perform global pooling on the input in the channel attention module: Calculate the average value of all pixel values in the channel as the feature value of the channel, and the output is Tensor; and pad the tensor with a size of 3; use a convolution kernel of size 7 to perform a one-dimensional convolution on the padded tensor; compress the convolution result to between (0,1) using the sigmoid function and reconstruct it as The tensor is then multiplied by the input signal to obtain the output data with channel attention weights; The output data of the channel attention module is input into the spatial attention module; in the spatial attention module, maximum pooling and average pooling are performed on the plane dimension to obtain two The two feature maps are concatenated into one in the channel dimension. The feature map of The convolution layer is reduced to 1 channel, and then the sigmoid function is used to generate the spatial weight coefficient and multiplied by the input of the spatial attention module to obtain the final output; Step 5: Using the two-dimensional time-frequency diagram of the noise signal obtained in step 3 and the identification label corresponding to the original sample noise signal, the noise recognition model constructed in step 4 is trained to obtain a trained noise recognition model; Step 6: Process the noise signal actually collected according to the noise signal processing method in steps 1 to 3 to obtain a two-dimensional time-frequency diagram of the noise signal actually collected; input the two-dimensional time-frequency diagram of the noise signal actually collected into the trained noise recognition model for recognition.
2. The underwater noise recognition method based on continuous wavelet transform and improved residual neural network according to claim 1, characterized in that: The specific process of decomposing the noise signal using the variational mode decomposition algorithm in step 1 is as follows: Step 1.1: Assume the original signal is decomposed into modal components, the constrained variational model is constructed as: in is the set of modal functions, is the set of center frequencies of each mode, is the narrowband mode function, is the number of modes, is the sum of all modal components, To adjust the exponential term of the center frequency of each modal component, is the unit pulse function, for right The partial derivative of Represents the convolution operation; Step 1.2: Introduce the augmented Lagrange function to solve the constrained variational model in step 1 and obtain The best choice: in is the quadratic penalty factor, is the Lagrange multiplier.
3. The underwater noise recognition method based on continuous wavelet transform and improved residual neural network according to claim 1, characterized in that: The specific process of step 2 is: Calculate the Pearson correlation coefficient between each modal component and the original sample noise signal: in is the standard deviation of the modal component, is the standard deviation of the original sample noise signal; the modal components with correlation coefficients greater than K are summed to obtain the reconstructed noise signal .
4. The underwater noise recognition method based on continuous wavelet transform and improved residual neural network according to claim 3, characterized in that: In step 2, the threshold K is set to 0.
3.
5. The underwater noise recognition method based on continuous wavelet transform and improved residual neural network according to claim 1, characterized in that: The specific process of step 3 is as follows: According to the waveform of the noise signal in the time domain, a wavelet with a similar waveform is selected as the wavelet basis function; the reconstructed noise signal , its continuous wavelet transform is: in is the wavelet basis function, yes The complex conjugate of is the scale factor, which controls the horizontal expansion and contraction of the wavelet basis function. is the time shift factor, which controls the movement of the wavelet basis function on the time axis; The wavelet coefficients at different frequencies are combined and represented as an RGB three-channel image to obtain a two-dimensional time-frequency image containing the time-frequency characteristics of the noise signal.
6. The underwater noise recognition method based on continuous wavelet transform and improved residual neural network according to claim 5, characterized in that: Morlet wavelet is used as the wavelet basis function.
7. The underwater noise recognition method based on continuous wavelet transform and improved residual neural network according to claim 1, characterized in that: In step 5, the process of training the noise recognition model is as follows: The two-dimensional time-frequency diagram of the original sample noise signal is divided into a training set and a validation set. The two-dimensional time-frequency diagram of the samples in the training set is rotated 90° and 180°, and the enhanced samples are added to the training set. The noise recognition model was trained using the training set, using the Adam optimizer, setting the learning rate to 0.0001, and the BatchSize to 32. The training set was input into the neural network for 100 rounds of training. After each round of training, the validation set was tested, and the parameters of the round with the highest validation set accuracy during the training process were saved.
Citation Information
Patent Citations
Environmental sound recognition method based on pseudo-color time-frequency image and convolutional network
CN112652326A
Underwater acoustic communication signal modulation mode identification method based on improved gating network and residual network
CN113269077A