An integrated method and system for noise reduction and identification of ship targets under marine environmental noise interference.
By using a deep noise reduction and recognition integrated network, a dual-decoder structure and a multi-scale recognition module, combined with a noise reduction-classification joint training loss function, the accuracy and robustness of ship target classification under marine environmental noise interference are solved, achieving more efficient recognition performance.
Patent Information
- Application Number
- CN202510611225.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-05-13
AI Technical Summary
Existing technologies lack the accuracy and robustness for classifying and identifying ship targets under marine environmental noise interference. Traditional noise reduction methods cannot effectively improve recognition performance, and deep learning models have limited generalization ability in complex noise environments.
Design a deep noise reduction and recognition integrated network. The deep noise reduction network module with dual decoder structure processes the amplitude spectrum and phase spectrum of the signal in parallel. Combined with a multi-scale recognition module and a noise reduction-classification joint training loss function, end-to-end optimization is achieved.
It significantly improved the accuracy and signal-to-noise ratio of ship target classification, with an average improvement of 25.73% and 12.31%, respectively. The signal-to-noise ratio was improved by 3.97 dB, the root mean square error was reduced by 35.95%, and the mean absolute error was reduced by 35.59%.
Smart Images

Figure CN120561643B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of ship target recognition technology, and particularly relates to an integrated method and system for noise reduction and recognition of ship targets under marine environmental noise interference. Background Technology
[0002] Ship target classification and identification plays a vital role in marine monitoring and maritime traffic management, primarily through two methods: manual acoustic identification and automated machine identification. For a long time, manual acoustic identification relied mainly on sonar operators. Through years of experience, they were able to identify the acoustic signatures of different ships. However, training sonar operators is a lengthy and complex process, requiring extensive practical experience to achieve a high level of identification proficiency.
[0003] Machine recognition not only automatically learns useful features from massive amounts of data, reducing reliance on prior knowledge, but also avoids interference from human factors, improving classification efficiency. The machine recognition process mainly involves signal preprocessing, feature extraction, and classifier design. Extracted features typically include time-domain features, frequency-domain features, time-frequency features, nonlinear features, auditory features, and visual features. These features describe the characteristics of ships from different perspectives, providing crucial information for subsequent classification stages. Classifier design is based on the features of the training dataset, minimizing the error rate or loss when classifying test data by determining the optimal decision rule. Currently, traditional machine learning-based methods for ship target classification and recognition mainly include Naive Bayes, decision trees, support vector machines, K-nearest neighbors, and hidden Markov models. However, ship target recognition models based on traditional machine learning algorithms are inherently shallow structures with limited expressive and learning capabilities. With the increasing complexity of the marine environment, traditional classifiers are struggling to meet the accuracy requirements of practical applications.
[0004] Current research on ship target classification and recognition has shifted towards deep learning algorithms. Deep neural network models integrate feature learning, feature optimization, and classification decisions within a single framework. Through the cooperation between these stages, they better handle complex feature representations and improve recognition accuracy, demonstrating significant application potential in this field. Researchers have used methods such as Long Short-Term Memory networks and attention-based feature enhancement networks to demonstrate good recognition capabilities in various ship target classification tasks. However, single feature inputs often fail to fully exploit the potential of deep neural networks, limiting further performance improvements. Therefore, more research is exploring strategies for combining multiple features to more comprehensively represent ship information. Researchers have used fused features such as amplitude spectrum and phase spectrum to train neural networks, achieving accurate differentiation of ship targets.
[0005] In harsh marine environments, received ship signals are easily interfered with by marine ambient noise. Marine ambient noise is characterized by a wide frequency range, complex sound sources, and spatiotemporal dynamics. Based on frequency range, it can be categorized into extremely low frequency (ULF) noise (less than 10Hz), ultra-low frequency (ULF) noise (10Hz–300Hz), very low frequency (VLF) noise (300Hz–3kHz), and high frequency noise (greater than 10kHz). ULF noise mainly originates from crustal movement and underwater blasting operations; ULF noise originates from ship navigation and industrial activities; VLF noise originates from wind, waves, rainfall, and marine life; and high frequency noise mainly originates from thermal noise generated by molecular disturbances. This noise severely limits the propagation range of target sound waves at sea, increasing the difficulty of ship classification and identification. For example, noise from wind, waves, and rainfall often distorts ship signals; the vocal frequencies of marine life partially overlap with ship signals, easily leading to misjudgments; ship navigation itself is a major source of low-frequency noise, making target ship identification susceptible to interference from surrounding vessels.
[0006] In low signal-to-noise ratio (SNR) conditions, existing deep learning-based identification methods tend to overemphasize the dominant noise component, leading to a decline in the ability to extract features from target vessels and severely impacting decision accuracy. Furthermore, existing methods typically rely on relatively ideal SNR conditions during model training, resulting in insufficient generalization ability when applied to complex and variable marine environments. Therefore, the objective of this invention is to design a model capable of resisting marine environmental noise interference, which is crucial for further improving vessel classification and identification performance.
[0007] Currently, two main methods are used to address noise interference and improve ship identification performance. One method involves introducing multi-condition training and data augmentation strategies during the training process of the identification model. For example, training the classification model with mixed data of different signal-to-noise ratios increases the amount of training data; random perturbations such as time-domain distortion and frequency-domain masking are added in the time and frequency domains to enhance the diversity of training data. However, this method significantly increases the computational cost of model training, limiting its practical application value. The other method involves introducing a separate noise reduction front-end, placing it upstream of the identification module as a preprocessing unit to maximize the signal-to-noise ratio before classification. Traditional noise reduction methods include spectral subtraction, Wiener filtering, and wavelet thresholding. However, when the frequency ranges of the signal and noise overlap significantly, spectral subtraction introduces significant signal distortion; Wiener filtering relies on noise estimation and performs poorly in non-stationary noise conditions; and wavelet thresholding's noise reduction performance is easily affected by the basis function and threshold setting. The noise sources in the marine environment are complex, exhibiting high uncertainty and dynamic changes, making traditional noise reduction methods unsuitable for such environments. Compared with traditional methods, deep learning-based denoising algorithms have improved the dimensionality of feature extraction. The working method of neural networks does not require pre-setting the characteristics of noise, can adapt to various noise types, and improves the flexibility of denoising.
[0008] However, the optimization objectives of noise reduction processing and the recognition module are fundamentally different. Although the signal-to-noise ratio is improved after noise reduction, it may not necessarily provide optimal results for classification and recognition tasks. Existing research shows that while noise reduction processing can effectively suppress noise interference, it also introduces signal distortion, thus negatively impacting classification and recognition. Therefore, it is urgent to address the technical problems arising from the inconsistency between the optimization objectives of noise reduction processing and target recognition tasks in existing technologies. Summary of the Invention
[0009] The purpose of this invention is to overcome the shortcomings of the prior art and to propose an integrated method and system for noise reduction and identification of ship targets under marine environmental noise interference.
[0010] In view of this, the present invention proposes an integrated noise reduction and identification method for ship target classification under marine environmental noise interference, comprising:
[0011] The radiated noise signal of the ship to be identified is subjected to a short-time Fourier transform to obtain the corresponding time spectrum, which is then input into a trained integrated ship target noise reduction and classification model to obtain the ship classification; wherein, the integrated ship target noise reduction and classification model includes:
[0012] The deep noise reduction network module incorporates the MP-SENet parallel decoder structure into DBSA-Net to enable separate reconstruction of the amplitude spectrum and phase spectrum of the ship signal in the time and frequency domains.
[0013] The inverse Fourier transform module is used to perform inverse Fourier transform on the amplitude spectrum and phase spectrum of the ship signal to obtain the denoised ship signal to be identified in the time domain.
[0014] The multi-scale recognition module is used to classify ships by extracting multi-scale features, channel-aware squeezing-excitation, and hierarchical feature fusion from the denoised ship signals.
[0015] Preferably, the input of the deep noise reduction network module is the time spectrum of the ship radiated noise signal to be identified, and the output is the amplitude spectrum and the phase spectrum. The deep noise reduction network module includes an encoder, a phase decoder and an amplitude decoder; wherein the encoder is followed by four two-stage convolutional enhancement Transformers and then connected in parallel with the amplitude decoder and the phase decoder.
[0016] Preferred,
[0017] The encoder includes three cascaded convolutional modules, each of which includes a two-dimensional convolutional layer, a BN layer, and a PReLU activation function.
[0018] The phase decoder includes three convolutional modules and one parallel phase estimation module. Each convolutional module includes a two-dimensional transposed convolutional layer, a BN layer, a PReLU activation function, and a two-dimensional convolutional layer. The parallel phase estimation module first extracts the real and imaginary components through two parallel two-dimensional convolutional layers, and then activates these two components using the arctangent function to predict the denoised phase spectrum.
[0019] The amplitude decoder includes three cascaded convolutional modules, each of which includes a two-dimensional transposed convolutional layer, a BN layer, a PReLU activation function, and a two-dimensional convolutional layer.
[0020] Preferably, the input of the multi-scale recognition module is a combined feature, which is a three-channel combined feature obtained by stacking the Mel spectrum, CQT spectrum and Gammatone spectrum extracted from the denoised ship signal to be identified. The multi-scale recognition module includes: a one-dimensional convolution with ReLU activation function and BN layer, three parallel SE-Res2 modules connected by weighted skip connections and then passed through a dense layer, and outputting a high-level semantic representation for statistical pooling layer.
[0021] Preferably, the SE-Res2 module adopts the Res2Net module, which uses dilated convolution to extract multi-scale ship features. The squeeze-excitation module follows the Res2Net module to dynamically enhance the response of the target ship's relevant channels.
[0022] Preferably, the statistical pooling layer introduces a channel-based attention mechanism, the processing of which includes:
[0023] The extracted attention information is projected into a smaller dimension representation and activated by a non-linear function, as shown in the following equation:
[0024] u t =tanh(W g h t +b g )
[0025] in, It is the feature vector of the t-th frame. and Here, denoted as the projection matrix and bias, respectively; tanh(·) is the non-linear activation function used; and C and R are the number of channels in the input feature vector and the feature dimension of the output after projection, respectively.
[0026] The obtained projection is represented as u t This is converted into channel-related attention scores, as shown in the following formula:
[0027]
[0028] Among them, s t,c It is the attention score for time step t and channel c. and The weights and biases of channel c;
[0029] Attention score s t,c The attention weight ω is obtained by normalization using the Softmax(·) function. t,c .
[0030] Preferably, the multi-scale recognition module further includes a fully connected layer for distributing the attention weights ω. t,c The index of the maximum value is obtained by using the Argmax() function. The label corresponding to this index in the ship dictionary is the final classification result.
[0031] Preferably, the method further includes training the integrated noise reduction and classification model for ship targets using a joint training loss function for noise reduction and classification, wherein the joint training loss function for noise reduction and classification satisfies the following formula:
[0032]
[0033]
[0034] in, For the joint training loss function, To reduce noise loss, For classification loss, μ is a weighting factor. For time domain loss, For magnitude loss, For complex losses, The consistency loss is given by N, where N is the number of samples. C Let y be the number of categories. If the true category of sample i is c, then y ic Equals 1, otherwise y ic p is 0 ic It is the predicted probability that sample i belongs to category c.
[0035] On the other hand, the present invention provides an integrated noise reduction and identification system for ship target classification under marine environmental noise interference, comprising:
[0036] The short-time Fourier transform module is used to perform a short-time Fourier transform on the radiated noise signal of the ship to be identified, so as to obtain the corresponding time spectrum;
[0037] A classification output module is used to input the time-spectrum data into the trained integrated ship target noise reduction and classification model to obtain ship classifications; wherein, the integrated ship target noise reduction and classification model includes:
[0038] The deep noise reduction network module incorporates the MP-SENet parallel decoder structure into DBSA-Net to enable separate reconstruction of the amplitude spectrum and phase spectrum of the ship signal in the time and frequency domains.
[0039] The inverse Fourier transform module is used to perform inverse Fourier transform on the amplitude spectrum and phase spectrum of the ship signal to obtain the denoised ship signal to be identified in the time domain.
[0040] The multi-scale recognition module is used to classify ships by extracting multi-scale features, channel-aware squeezing-excitation, and hierarchical feature fusion from the denoised ship signals.
[0041] Compared with the prior art, the advantages of the present invention are:
[0042] 1. To address the problem that inconsistent optimization objectives between signal denoising and classification recognition lead to models with good denoising performance failing to effectively improve ship recognition accuracy, this invention proposes a deep denoising and recognition integrated network and designs a joint training loss function for denoising and classification to optimize the overall model. The advantage lies in its end-to-end optimization strategy, achieving deep coupling between denoising and recognition. On one hand, the loss of the recognition module not only backpropagates to its own parameters but also further propagates to the deep denoising network module, subjecting it to the direct constraint of the classification loss and allowing it to more effectively retain feature information beneficial to classification. Thus, the deep denoising network module is no longer an independent preprocessing step but an integrated sub-module driven by the target recognition task. On the other hand, the collaborative optimization mechanism allows the recognition module to fully engage with and learn more feature distributions under noise interference during training, thereby significantly improving recognition performance. Experimental results on the Shipsear dataset show that compared to direct classification recognition on noisy data, the integrated network achieves an average improvement in recognition accuracy of 25.73%. Compared to the traditional two-stage denoising-classification method, the integrated network achieves an average improvement in recognition accuracy of 12.31%.
[0043] 2. To address the shortcomings of traditional denoising methods in handling marine background noise interference, which suffer from poor performance and insufficient flexibility, this invention designs a deep denoising network module with a dual-decoder structure, capable of processing the amplitude and phase spectra of the signal in parallel. The independently constructed phase decoder effectively alleviates the estimation difficulties caused by the entanglement and unstructured nature of the phase, improving the accuracy of phase recovery and thus enhancing denoising performance. Experimental results on the Shipsear dataset show that the denoising module achieves an average signal-to-noise ratio improvement of 3.97 dB, an average peak signal-to-noise ratio improvement of 3.81 dB, a root mean square error reduction of 35.95%, and a mean absolute error reduction of 35.59%.
[0044] 3. To address the problem that single features are insufficient to fully exploit the performance potential of neural networks, thus limiting classification and recognition performance, this invention employs a multi-feature combination approach to input into a multi-scale recognition module. By comprehensively utilizing the complementary information of different features, the richness of feature representation is enhanced. The recognition module dynamically adjusts the weights of feature channels through multi-scale feature extraction, a channel-aware squeezing-excitation module, and hierarchical feature fusion, thereby enhancing the expressive power of key features and improving classification and recognition performance. Experimental results on the Shipsear dataset show that, without added noise, the module achieves a recognition accuracy of 86.34%. Attached Figure Description
[0045] Figure 1 shows a comparison of time-frequency diagrams, where Figure 1(a) is the original signal, Figure 1(b) is the signal superimposed with rainfall noise, and Figure 1(c) is the noise reduction result;
[0046] Figure 2 It is a network structure for an integrated model of ship target noise reduction and classification;
[0047] Figure 3 It is a D2D-Net network structure;
[0048] Figure 4 It is the ECAPA-TDNN network structure. Detailed Implementation
[0049] This invention proposes an integrated noise reduction and identification method for ship target classification under marine environmental noise interference, comprising:
[0050] The radiated noise signal of the ship to be identified is subjected to a short-time Fourier transform to obtain the corresponding time spectrum, which is then input into a trained integrated ship target noise reduction and classification model to obtain the ship classification; wherein, the integrated ship target noise reduction and classification model includes:
[0051] The deep noise reduction network module incorporates the MP-SENet parallel decoder structure into DBSA-Net to enable separate reconstruction of the amplitude spectrum and phase spectrum of the ship signal in the time and frequency domains.
[0052] The inverse Fourier transform module is used to perform inverse Fourier transform on the amplitude spectrum and phase spectrum of the ship signal to obtain the denoised ship signal to be identified in the time domain.
[0053] The multi-scale recognition module is used to classify ships by extracting multi-scale features, channel-aware squeezing-excitation, and hierarchical feature fusion from the denoised ship signals.
[0054] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and embodiments.
[0055] Example 1
[0056] Embodiment 1 of the present invention proposes an integrated method for noise reduction and identification of ship targets under marine environmental noise interference.
[0057] In the ship classification and identification process, the quality of feature extraction directly affects model performance and classification results. This invention uses three time-frequency transforms to extract Mel spectrum, Constant Q Transform (CQT) spectrum, and Gammatone spectrum, respectively. These three complement each other to form richer classification features. Among them, the Mel spectrum provides frequency distribution information that better matches human auditory perception, the CQT spectrum has advantages in high-frequency resolution, and the Gammatone spectrum highlights low-frequency characteristics.
[0058] However, in complex marine environments, noise interference severely impacts the feature extraction process. As shown in Figure 1, Figure 1(a) is the time-frequency diagram of the original signal from a fishing vessel, and Figure 1(b) is the time-frequency diagram of the same signal after adding rainfall noise. It can be observed that the time-frequency structure of the original signal is significantly disrupted after adding rainfall noise. Figure 1(c) shows the result after processing with a denoising algorithm. Compared with the signal after adding rainfall noise in Figure 1(b), most of the noise is effectively suppressed. Compared with the original signal in Figure 1(a), the denoised signal shows significant information loss, and some structural features are weakened, as shown in the red box in Figure 1(c). This phenomenon indicates that the traditional two-stage denoising-classification method may damage key information with classification discriminative power during denoising, affecting the integrity of subsequent feature extraction and thus impacting the performance of the recognition model.
[0059] To address this issue and further enhance the noise resistance of ship identification systems in marine environments, this invention proposes a deep noise reduction and identification integrated network, such as... Figure 2 As shown, during the training phase, the ship radiated noise signal is first transformed from the time domain to the time-frequency domain using a Short Time Fourier Transform (STFT). Then, a deep denoising network module is used to process the signal to generate a clearer ship signal. Based on this, Mel spectrum, CQT spectrum, and Gammatone spectrum are extracted from the denoised ship signal, and these three time-frequency features are stacked to form a three-channel combined feature, which is input into the multi-scale recognition module to predict ship classification. The entire network achieves gradient information sharing and parameter collaborative updating between the denoising and recognition modules through a jointly designed denoising-classification training loss function. During the testing phase, a noisy signal with a low signal-to-noise ratio is input, and the ship classification result is directly output through forward propagation calculation of the integrated network.
[0060] The integrated deep denoising and recognition network is designed with a unified optimization objective to achieve joint training of denoising and recognition. To realize this concept, specific modules need to be designed as carriers. The following sections will provide a detailed introduction to the deep denoising network module, the inverse Fourier transform module, the multi-scale recognition module, and the joint training loss for denoising and classification in the integrated deep denoising and recognition network.
[0061] (1) Dual decoder deep noise reduction network module
[0062] This invention proposes a dual-decoder denoising network module (D2D-Net) for denoising ship signals. This module incorporates the MP-SENet parallel decoder structure idea into DBSA-Net, enabling separate reconstruction of the amplitude and phase spectra of the ship signal in the time and frequency domains. While DBSA-Net indirectly optimizes the phase through complex spectra, the improved D2D-Net directly predicts the phase spectrum of the ship signal through an independent phase decoder, thus overcoming the estimation difficulties caused by the entanglement and unstructured nature of phase.
[0063] The structure of the network module is as follows: Figure 3 As shown, it mainly includes an encoder, an amplitude decoder, and a phase decoder. The encoder and decoder are connected by four two-stage convolutional enhanced Transformers (TS-Conformers) to capture long-term dependencies in the time-frequency domain and enhance the joint modeling capability of amplitude and phase.
[0064] The original signal first undergoes a short-time Fourier transform to convert it into a time-frequency domain representation, which is then input into the encoder. The encoder consists of three cascaded convolutional modules. Each convolutional module contains a two-dimensional convolutional layer, a batch normalization (BN) layer, and a parametric rectified linear unit (PReLU) activation function. The two-dimensional convolutional layer is used to downsample the features, with the number of output channels increasing progressively layer by layer. In one embodiment, these channels are 16, 32, and 64, respectively, to gradually extract higher-dimensional feature representations.
[0065] The amplitude decoder predicts the amplitude spectrum mask from the time-frequency domain representation and multiplies it by the noisy amplitude spectrum to reconstruct the denoised amplitude spectrum. This decoder consists of three cascaded convolutional modules, each containing a 2D transposed convolutional layer, a Batch Normalization (BN) layer, a PReLU activation function, and a 2D convolutional layer. The transposed convolutional layer performs the upsampling operation, with the number of input channels decreasing layer by layer; in one embodiment, this is 64, 32, and 16 respectively. The other 2D convolutional layer learns the spectral gain to further optimize the quality of the reconstructed signal.
[0066] The phase decoder is used to predict the denoised phase spectrum from the time-frequency domain representation. This decoder consists of three convolutional modules and a parallel phase estimation module. Similar to the amplitude decoder, each convolutional module sequentially contains a 2D transposed convolutional layer, a BN layer, a PReLU activation function, and another 2D convolutional layer, with upsampling achieved through the transposed convolutional layer. Since the phase information of natural acoustic signals such as ship radiated noise often exhibits unstructured and convoluted characteristics, this can affect the accurate estimation of the denoised phase spectrum. Therefore, a parallel phase estimation module follows the convolutional modules. In this module, the real and imaginary components are first extracted through two parallel 2D convolutional layers, and then the arctangent function is used to activate these two components to predict the denoised phase spectrum. This parallel processing approach can capture phase features more comprehensively and improve the accuracy of phase prediction.
[0067] (2) Inverse Fourier transform module, used to perform inverse Fourier transform on the amplitude spectrum and phase spectrum of the ship signal to obtain the denoised ship signal in the time domain.
[0068] (3) ECAPA-TDNN multi-scale recognition module
[0069] This invention uses ECAPA-TDNN for ship target classification and identification, with the structure as follows: Figure 4 As shown.
[0070] ECAPA-TDNN first uses one-dimensional convolutions with ReLU activation and BN layers to segment the temporal spectrum of the denoising network output into frame-level features. Then, a Res2Net module is used to extract multi-scale ship features using dilated convolutions. Each Res2Net module is followed by a squeeze-excitation module to obtain the SE-Res2 module, which dynamically enhances the response of the target ship-related channels. The outputs of all SE-Res2 modules are fused through weighted skip connections for multi-level feature fusion. Dense layers are used to reduce the dimensionality of the fused features, and the output is used for high-level semantic representations in statistical pooling layers.
[0071] Furthermore, to differentiate the importance of different frames to the classification results, ECAPA-TDNN introduces a statistical pooling module based on a channel attention mechanism. This module calculates the score of the attention mechanism to determine the importance of different frames within a given channel, enabling the network to dynamically adjust attention based on global attributes. First, the extracted attention information is projected into a smaller dimension representation and activated by a nonlinear function, as shown in Equation (1). This dimensionality reduction representation is shared by the channels, which can reduce the risk of overfitting during training.
[0072] u t =tanh(W g h t +b g(1)
[0073] in, It is the feature vector of the t-th frame. and Let be the projection matrix and the bias, respectively, and tanh(·) be the nonlinear activation function used. Then, the obtained projection representation is converted into channel-related attention scores, as shown in Equation (2).
[0074]
[0075] Among them, s t,c It is the attention score for time step t and channel c. and Let be the weights and biases of channel c. The attention scores are normalized using the Softmax(·) function to obtain the attention weights, the magnitude of which reflects the importance of each frame in channel c, as shown in equation (3).
[0076] ω t,c =Softmax(s t,c (3)
[0077] Where, ω t,c This represents the normalized attention weights.
[0078] Finally, the output is obtained through a fully connected layer. The output is then passed through the Argmax() function to obtain the index of the maximum value. The label corresponding to this index in the ship dictionary is the final classification result.
[0079] (4) Noise reduction-classification joint training loss
[0080] This invention proposes a joint training loss function for noise reduction and classification. This is used to train a deep denoising and recognition integrated network model. The loss function consists of a denoising loss and a classification loss, with the weights of the two parts balanced by setting the parameter μ to optimize the overall framework. As shown in equation (4), this is the joint training loss function. The mathematical expression for .
[0081]
[0082] in, To reduce noise loss, For classification loss, μ is a weighting factor with a value between 0 and 1.
[0083] Noise reduction loss uses time domain loss magnitude loss Complex loss Consistency loss and phase loss The definitions of temporal loss, amplitude loss, and complex loss are derived from CMGAN. The temporal loss is calculated for both the clean ship signal x and the denoised ship signal. The L1 norm distance between them is shown in Equation (5).
[0084]
[0085] in, Indicates the relationship between x and To find the expectation of the joint distribution, i.e., to calculate the expectation of all sample pairs. The L1 norm distance is then averaged. A frequency domain representation of the clean ship signal is obtained through STFT, with its complex spectrum denoted as X and its amplitude spectrum as X0. m The phase spectrum is X p As shown in equations (6), (7), and (8).
[0086]
[0087] Where x[n+mH] is the signal frame, m represents the time frame index, L represents the frame length, H represents the frame shift, w[n] represents the window function, N represents the number of FFT points, k represents the frequency index, and X r X represents the real part. i Indicates the imaginary part.
[0088] Amplitude loss was used to calculate the clean amplitude spectrum X. m and the amplitude spectrum after noise reduction The L2 norm distance between them is shown in Equation (9). The complex loss is used to calculate the clean complex spectrum X and the denoised complex spectrum. The L2 norm distance between the real and imaginary parts is shown in Equation (10). The L2 norm is more sensitive to smaller errors than the L1 norm and is more suitable for adjusting the amplitude and phase of a signal.
[0089]
[0090] Building upon this, a consistency loss is further incorporated. This loss function calculates the enhanced complex spectrum directly generated by the model. and the complex spectrum regenerated after passing through iSTFT and STFT The L2 norm distance is shown in Equation (11).
[0091]
[0092] Furthermore, phase loss optimization is employed to overcome the difficulties caused by phase winding and unstructured characteristics. Phase winding refers to the phase value cycling within the range [-π, π], which leads to errors when directly calculating the phase difference. An anti-winding function is defined as shown in Equation (12) to avoid the error propagation problem caused by phase winding. Unstructured characteristics mean that the phase information lacks a clear structure, making direct optimization difficult.
[0093]
[0094] Phase loss includes instantaneous phase loss Group delay loss and instantaneous angular frequency loss As shown in equation (13). The instantaneous phase loss is used to calculate the clean phase spectrum X. p and the phase spectrum after noise reduction The L1 norm distance between them is shown in Equation (14). The group delay loss is used to calculate the clean phase spectrum X. p and the phase spectrum after noise reduction The L1 norm distance of the difference on the frequency axis is shown in Equation (15). The instantaneous angular frequency loss is used to calculate the clean phase spectrum X. p and the phase spectrum after noise reduction The L1 norm distance of the difference on the time axis is shown in Equation (16).
[0095]
[0096] Among them, Λ DF and Λ DT These represent the differential operators along the frequency axis and the time axis, respectively.
[0097] Finally, the above loss functions are linearly combined to obtain the final noise reduction loss. As shown in equation (17).
[0098]
[0099] Among them, γ1, γ2, γ3, γ4, and γ5 are hyperparameters that need to be adjusted during the experiment.
[0100] Classification loss Let be the cross-entropy loss function, expressed as in equation (18).
[0101]
[0102] Where N is the sample size, N C y represents the number of categories. ic The value is either 1 or 0. If the true class of sample i is c, the value is 1; otherwise, it is 0. icThis is the predicted probability that sample i belongs to category c.
[0103] Example 2
[0104] Embodiment 2 of the present invention provides an integrated noise reduction and identification system for ship target classification under marine environmental noise interference, comprising:
[0105] The short-time Fourier transform module is used to perform a short-time Fourier transform on the radiated noise signal of the ship to be identified, so as to obtain the corresponding time spectrum;
[0106] A classification output module is used to input the time-spectrum data into the trained integrated ship target noise reduction and classification model to obtain ship classifications; wherein, the integrated ship target noise reduction and classification model includes:
[0107] The deep noise reduction network module incorporates the MP-SENet parallel decoder structure into DBSA-Net to enable separate reconstruction of the amplitude spectrum and phase spectrum of the ship signal in the time and frequency domains.
[0108] The inverse Fourier transform module is used to perform inverse Fourier transform on the amplitude spectrum and phase spectrum of the ship signal to obtain the denoised ship signal to be identified in the time domain.
[0109] The multi-scale recognition module is used to classify ships by extracting multi-scale features, channel-aware squeezing-excitation, and hierarchical feature fusion from the denoised ship signals.
[0110] It is worth noting that in the embodiments of the above system, the modules included are divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional module are only for easy differentiation and are not used to limit the scope of protection of the present invention.
[0111] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to the embodiments, those skilled in the art should understand that modifications or equivalent substitutions to the technical solutions of the present invention do not depart from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A noise reduction and identification integrated method for ship target classification under marine environmental noise interference, comprising: The radiated noise signal of the ship to be identified is subjected to a short-time Fourier transform to obtain the corresponding time spectrum, which is then input into a trained integrated ship target noise reduction and classification model to obtain the ship classification; wherein, the integrated ship target noise reduction and classification model includes: The deep noise reduction network module incorporates the MP-SENet parallel decoder structure into DBSA-Net to enable separate reconstruction of the amplitude spectrum and phase spectrum of the ship signal in the time and frequency domains. The inverse Fourier transform module is used to perform inverse Fourier transform on the amplitude spectrum and phase spectrum of the ship signal to obtain the denoised ship signal to be identified in the time domain. The multi-scale recognition module is used to classify ships by extracting multi-scale features, channel-aware squeezing-excitation, and hierarchical feature fusion from the denoised ship signals.
2. The integrated noise reduction and identification method for ship target classification under marine environmental noise interference as described in claim 1, characterized in that, The input of the deep noise reduction network module is the time spectrum of the ship radiated noise signal to be identified, and the output is the amplitude spectrum and the phase spectrum. The deep noise reduction network module includes an encoder, a phase decoder and an amplitude decoder; wherein the encoder is followed by four two-stage convolutional enhancement Transformers and then connected in parallel with the amplitude decoder and the phase decoder.
3. The integrated noise reduction and identification method for ship target classification under marine environmental noise interference as described in claim 2, characterized in that, The encoder includes three cascaded convolutional modules, each of which includes a two-dimensional convolutional layer, a BN layer, and a PReLU activation function. The phase decoder includes three convolutional modules and one parallel phase estimation module. Each convolutional module includes a two-dimensional transposed convolutional layer, a BN layer, a PReLU activation function, and a two-dimensional convolutional layer. The parallel phase estimation module first extracts the real and imaginary components through two parallel two-dimensional convolutional layers, and then activates these two components using the arctangent function to predict the denoised phase spectrum. The amplitude decoder includes three cascaded convolutional modules, each of which includes a two-dimensional transposed convolutional layer, a BN layer, a PReLU activation function, and a two-dimensional convolutional layer.
4. The integrated noise reduction and identification method for ship target classification under marine environmental noise interference as described in claim 1, characterized in that, The input to the multi-scale recognition module is a combined feature, which is a three-channel combined feature obtained by stacking the Mel spectrum, CQT spectrum and Gammatone spectrum extracted from the denoised ship signal to be identified. The multi-scale recognition module includes: a one-dimensional convolution with ReLU activation function and BN layer, three parallel SE-Res2 modules connected by weighted skip connections and then passed through a dense layer, and outputting a high-level semantic representation for statistical pooling layer.
5. The integrated noise reduction and identification method for ship target classification under marine environmental noise interference as described in claim 4, characterized in that, The SE-Res2 module adopts the Res2Net module, which uses dilated convolution to extract multi-scale ship features. The squeeze-excitation module follows the Res2Net module to dynamically enhance the response of the target ship's relevant channels.
6. The integrated noise reduction and identification method for ship target classification under marine environmental noise interference as described in claim 4, characterized in that, The statistical pooling layer introduces a channel-based attention mechanism, and the processing includes: The extracted attention information is projected into a smaller dimension representation and activated by a non-linear function, as shown in the following equation: u t =tanh(W g h t +b g ) in, It is the feature vector of the t-th frame. and Here, denoted as the projection matrix and bias, respectively; tanh(·) is the non-linear activation function used; and C and R are the number of channels in the input feature vector and the feature dimension of the output after projection, respectively. The obtained projection is represented as u t This is converted into channel-related attention scores, as shown in the following formula: Among them, s t,c It is the attention score for time step t and channel c. and The weights and biases of channel c; Attention score s t,c The attention weight ω is obtained by normalization using the Softmax(·) function. t,c .
7. The integrated noise reduction and identification method for ship target classification under marine environmental noise interference as described in claim 6, characterized in that, The multi-scale recognition module also includes a fully connected layer for applying attention weights ω. t,c The index of the maximum value is obtained by using the Argmax() function. The label corresponding to this index in the ship dictionary is the final classification result.
8. The integrated noise reduction and identification method for ship target classification under marine environmental noise interference as described in claim 1, characterized in that, The method further includes training the integrated noise reduction and classification model for ship targets using a joint training loss function for noise reduction and classification, wherein the joint training loss function for noise reduction and classification satisfies the following equation: in, For the joint training loss function, To reduce noise loss, For classification loss, μ is a weighting factor. For time domain loss, For magnitude loss, For complex losses, The consistency loss is given by N, where N is the number of samples. C Let y be the number of categories. If the true category of sample i is c, then y ic Equals 1, otherwise y ic p is 0 ic It is the predicted probability that sample i belongs to category c.
9. A noise reduction and identification integrated system for classifying ship targets under marine environmental noise interference, characterized in that, include: The short-time Fourier transform module is used to perform a short-time Fourier transform on the radiated noise signal of the ship to be identified, so as to obtain the corresponding time spectrum; and A classification output module is used to input the time-spectrum data into the trained integrated ship target noise reduction and classification model to obtain ship classifications; wherein, the integrated ship target noise reduction and classification model includes: The deep noise reduction network module incorporates the MP-SENet parallel decoder structure into DBSA-Net to enable separate reconstruction of the amplitude spectrum and phase spectrum of the ship signal in the time and frequency domains. The inverse Fourier transform module is used to perform inverse Fourier transform on the amplitude spectrum and phase spectrum of the ship signal to obtain the denoised ship signal to be identified in the time domain. The multi-scale recognition module is used to classify ships by extracting multi-scale features, channel-aware squeezing-excitation, and hierarchical feature fusion from the denoised ship signals.
Citation Information
Patent Citations
Target noise extraction and evaluation method and device
CN118197353A
Real-time dynamic noise reduction using convolutional networks
US20210012767A1