Underwater acoustic signal denoising method based on time-frequency adaptive dual-path Conformer network

Through the water acoustic signal denoising method based on time-frequency adaptive dual-path Conformer network, the problem of poor noise removal effect in complex water acoustic environments is solved, efficient noise cancellation and signal reconstruction are achieved, and the performance of the water acoustic signal processing system is improved.

CN120296311APending Publication Date: 2025-07-11SOUTH CHINA UNIV OF TECH
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510336961.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing water acoustic signal processing methods have limited denoising effects in complex water acoustic environments, making it difficult to effectively model multiple noise characteristics, lack the ability to adapt to different noise types, and have high computational complexity, making it difficult to deploy.

Method used

The water acoustic signal denoising method based on the time-frequency adaptive dual-path Conformer network is adopted. By constructing a multi-signal-to-noise ratio, the time-frequency adaptive dual-path Conformer network is used for feature extraction and denoising, including the time path Conformer and the frequency path Conformer, and signal reconstruction is carried out by combining the multi-scale fusion dynamic gated network.

Benefits of technology

It improves the denoising effect of water acoustic signals and the generalization ability of the model, reduces the computational complexity, enhances the recognition ability of water acoustic target classification tasks, and adapts to signal processing under different noise conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296311A_ABST
    Figure CN120296311A_ABST
Patent Text Reader

Abstract

The invention discloses an underwater acoustic signal denoising method based on a time-frequency adaptive dual-path Conformer network, and the method comprises the steps: carrying out the preprocessing of real marine environment noise and underwater acoustic target signals, and generating a multi-signal-to-noise-ratio noisy data set; performing short-time Fourier transform on the noisy data, and extracting real part and imaginary part features to form a feature tensor; extracting features by using a feature encoder, and generating intermediate feature representation; respectively extracting a time path feature and a frequency path feature through a time-frequency adaptive dual-path Conformer network; performing multi-scale convolution processing and weighted fusion on the extracted time-frequency features by adopting a multi-scale fusion dynamic gating network; performing nonlinear mapping on the fused features by using a feature decoder to generate a mask matrix; and restoring the complex frequency spectrum based on the mask matrix, and restoring the denoised underwater acoustic time domain signal. According to the method, time-frequency information is fully mined in combination with underwater sound noise characteristics, and the underwater sound signal denoising effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of underwater acoustic signal processing, and particularly relates to an underwater acoustic signal denoising method based on a time-frequency adaptive dual-path Conformer network. Background Art

[0002] Underwater acoustic signals are easily interfered by environmental noise, equipment noise, reverberation, etc. during the propagation process, resulting in a decline in signal quality and affecting the accuracy of subsequent processing. Traditional denoising methods include time-domain filtering, frequency-domain filtering, and time-frequency analysis, etc. However, the denoising effects of these methods in complex underwater acoustic environments are limited, and it is easy to cause target signal loss or residual noise, affecting practical applications.

[0003] In recent years, deep learning technology has been widely applied to the denoising field and achieved good results. However, the existing methods still have the following problems: 1. The underwater acoustic environment is complex and the noise types are diverse. Underwater noise includes non-stationary noises such as wind noise, sea wave noise, rain noise, etc. The time-frequency characteristics of different noises vary greatly, and it is difficult for existing methods to effectively model the characteristics of various noises, resulting in insufficient generalization ability of the denoising model.

[0004] 2. There is a lack of high-quality noisy underwater acoustic datasets. Most of the publicly available underwater acoustic datasets are pure underwater acoustic signals, lacking noisy underwater acoustic data with annotations, resulting in difficulty in effectively training and evaluating deep learning models and affecting the applicability of the models in actual complex environments.

[0005] 3. Insufficient analysis of underwater acoustic characteristics. In current research, most methods directly use the technologies of speech denoising without fully considering the particularity of underwater acoustic signals and without deeply modeling the diversity and variation laws of underwater acoustic noises, resulting in the denoising model lacking the adaptability to different noise types.

[0006] 4. The computational complexity is relatively high and it is difficult to deploy. Deep learning models usually have a large number of parameters and high computational overhead. Summary of the Invention

[0007] The main purpose of the present invention is to overcome the deficiencies of the prior art and provide an underwater acoustic signal denoising method based on a time-frequency adaptive dual-path Conformer network, which can effectively eliminate noise, improve signal quality and accuracy for the scenario of low signal-to-noise ratio of ocean environmental noise.

[0008] To achieve the above purpose, the present invention adopts the following technical solutions: An underwater acoustic signal denoising method based on a time-frequency adaptive dual-path Conformer network, the underwater acoustic signal denoising method includes the following steps: S1. Collect the real marine environmental noise and underwater acoustic target time-domain signals and perform preprocessing. After energy threshold filtering, perform additive superposition to generate a noisy underwater acoustic target dataset with multiple signal-to-noise ratios; S2. Perform feature processing on the noisy underwater acoustic target dataset with multiple signal-to-noise ratios. Obtain time-frequency information through short-time Fourier transform, split it into real and imaginary parts, and stack the real and imaginary parts of the time-frequency information along the channel dimension to form a feature tensor; S3. Input the feature tensor into the feature encoder. The feature encoder performs high-level feature encoding on the feature tensor and extracts an intermediate feature representation with global context information; S4. After reshaping the size of the intermediate feature representation, input it into the time-frequency adaptive dual-path Conformer network for learning and training. The time-frequency adaptive dual-path Conformer network includes a time-path Conformer and a frequency-path Conformer. Among them, the time-path Conformer is composed of a first feed-forward network layer, an improved multi-head window attention module, a convolutional enhancement module, and a second feed-forward network layer connected in sequence. The frequency-path Conformer is composed of a third feed-forward network layer, a band-limited multi-head attention module, a convolutional enhancement module, and a fourth feed-forward network layer connected in sequence. The intermediate feature is input into the time-path Conformer and the frequency-path Conformer simultaneously to extract the time-path feature and the frequency-path feature respectively; S5. Input the time-path feature and the frequency-path feature into the multi-scale fusion dynamic gating network. Perform multi-scale convolution on the time-path feature and the frequency-path feature respectively and then splice them. Through weighted fusion by the dynamic gating module in the multi-scale fusion dynamic gating network, obtain the weighted fusion time-frequency feature; S6. Input the weighted fusion time-frequency feature into the feature decoder. The feature decoder performs non-linear mapping on the weighted fusion time-frequency feature to generate a mask matrix; S7. Perform complex spectrum restoration processing on the mask matrix. Perform product and addition / subtraction operations on the mask matrix and the feature tensor to obtain the denoised real and imaginary components, and then perform inverse short-time Fourier transform to restore the denoised underwater acoustic time-domain target signal.

[0009] Furthermore, the process of generating the noisy underwater acoustic target dataset with multiple signal-to-noise ratios in step S1 is as follows: S101. Select the recorded marine environmental noise, select the natural environmental noise mainly composed of sea wind, sea waves, and rain. Perform preprocessing on the real marine environmental noise signal and the original underwater acoustic target signal. After resampling the signal, segment it with a fixed duration, retain a certain overlapping duration between each segment, and normalize the signal; S102. Perform energy threshold filtering on the processed real marine environmental noise signal and the original target signal, calculate the energy spectrum of the signal, set an energy threshold, and remove the low-energy components in the energy spectrum. S103. Set a certain signal-to-noise ratio range, randomly select real marine environmental noise and signal-to-noise ratio, and additively superimpose the target signal and the real marine environmental noise to obtain a multi-signal-to-noise ratio noisy underwater acoustic target dataset.

[0010] Furthermore, the process of the preprocessing in step S2 and obtaining the feature tensor is as follows: S201. Perform frame segmentation on the multi-signal-to-noise ratio noisy underwater acoustic target dataset, segment it according to a fixed frame length and frame shift to obtain a series of short-time frames; apply a window function to each frame signal to reduce spectral leakage and improve frequency resolution, and the window function is selected as the Hamming window. S202. Use the short-time Fourier transform to extract STFT features, and the expression of the short-time Fourier transform is as follows: ; where is the length of the FFT, is the frame shift, is the frame index, is the frequency index, is the moving window function; S203. Expand the formula of the short-time Fourier transform to obtain the expressions of the real part and the imaginary part . The expressions of the real part and the imaginary part are:

[0011] .

[0012] Concatenate the real part and the imaginary part in the channel dimension to obtain the feature tensor, and the expression of the feature tensor is:

[0013] where and respectively represent the frequency dimension and the time dimension, and the feature tensor is a complex matrix with the real part in the first channel and the imaginary part in the second channel, and the matrix size is .

[0014] Furthermore, the feature encoder adopts a multi-layer convolutional structure, and introduces skip connections between different convolutional modules to effectively fuse shallow and deep feature information and alleviate the possible information loss problem. Each convolutional module consists of a convolutional layer, a batch normalization layer, and an activation layer connected in sequence.

[0015] The convolutional layer is responsible for extracting time-frequency features. Since the sea breeze, sea waves, and rain noise in the underwater acoustic scenario have long-term correlations, ordinary convolutions may have difficulty effectively learning long-term dependence relationships. Multiple layers need to be stacked, resulting in a large computational load and low computational efficiency. Therefore, dilated convolutions with gradually increasing dilation rates are adopted. While ensuring computational efficiency and not increasing the number of parameters, the receptive field is effectively expanded, enabling the feature encoder to capture feature information in a larger range. After the convolutional layer, there is a batch normalization layer, which can accelerate convergence and improve the stability of training. Compared with the layer normalization layer, the batch normalization layer has stronger generalization ability and is more adaptable to different types of underwater acoustic noise. The activation layer can introduce non-linear mappings, improving the expression ability of the network and enabling the model to learn more discriminative features. After five convolutional modules with the same structure but different parameter settings, the number of channels is gradually increased, and the feature encoder outputs a high-dimensional intermediate feature representation.

[0016] Furthermore, the time-frequency adaptive dual-path Conformer module consists of a time-path Conformer and a frequency-path Conformer. The intermediate feature representation output by the feature encoder will enter both the time-path Conformer and the frequency-path Conformer for learning simultaneously. The Conformer structure is proposed based on the convolutional neural network (CNN) and the Transformer network, combining the advantages of both, with better performance in feature extraction and higher computational efficiency. On this basis, the Conformer structure is further improved to obtain a structure more suitable for the underwater acoustic scenario.

[0017] First, the Conformer structure is designed as a parallel dual-path of a time path and a frequency path, which can separately learn and model time and frequency features, making full use of the features of underwater acoustic signals in these two dimensions, thus obtaining a better denoising effect. The time path captures the dynamic changes and transient information of the signal, while the frequency path focuses on the spectral pattern and noise distribution. Combining the two can more comprehensively model the complexity of underwater acoustic signals.

[0018] Second, the specific time-path Conformer and frequency-path Conformer are respectively improved and optimized. The multi-head attention module in the time-path Conformer is optimized to form an improved multi-head window attention module. The traditional global self-attention is replaced with window attention for two main reasons: First, to improve the ability to capture time information. For the time path, local time correlations are more important, and important features in underwater acoustic signals usually have sudden changes within a short time window. By dividing the attention into windows, sudden noise can be learned, thus better denoising. Second, to significantly reduce the computational complexity. The computational complexity of the traditional global self-attention is , and by dividing the window, the attention complexity is reduced to , where is the window length, which is usually much smaller than the sequence length , which enables the network to improve the computational efficiency while ensuring performance. The calculation process of the improved multi-head window attention is as follows: After the intermediate representation is processed by the feed-forward network layer, it enters the improved multi-head window attention module. First, the input is divided into windows with a certain window length, and then each window is separately mapped to the query (Query, Q) matrix, the key (Key, K) matrix, and the value (Value, V) matrix, and the attention scores are calculated in combination with the relative position encoding. After each window is calculated, SoftMax normalization is used, and the attention scores of each window are merged and linearly projected to obtain the window attention enhanced feature. The expression for calculating the attention scores is:

[0019] where, is the relative position bias matrix. Using relative position encoding, by paying attention to the relative positions of elements in the sequence and calculating the distances between elements, their mutual relationships are represented. This method can represent the structural information in the sequence more flexibly and is more suitable for use on the time path.

[0020] The multi-head attention module in the frequency path Conformer is optimized to form a band-limited multi-head attention module. Since the underwater acoustic target signal and the real marine environmental noise signal often concentrate in different frequency ranges, the underwater acoustic target signal information is concentrated in the low frequency, while the frequency domain distribution of the real marine environmental noise signal is relatively wide. Using traditional global attention may introduce irrelevant noise correlations, resulting in poor denoising effects, and global attention focuses more on the entire frequency band and cannot model specific frequency regions. Therefore, it is improved to a band-limited multi-head attention module, which focuses on a specific frequency region by restricting the calculation range of each attention, avoiding interference from irrelevant frequency bands. At the same time, the band-limited multi-head attention module also reduces the computational complexity. The computational complexity of the traditional multi-head attention is , where is the size of the frequency dimension, is the feature dimension of the attention calculation, and the computational complexity of the band-limited multi-head attention module is , where is the bandwidth of each head. When When this is the case, the computational complexity is significantly reduced. The computational process of the band-limited multi-head attention is as follows: The output after passing through the feed-forward network layer enters the band-limited multi-head attention module. First, a dynamic band mask matrix is generated to divide the features into multiple bands. Each band is separately mapped to the QKV matrices. When calculating the attention scores, a mask matrix is introduced to obtain the limited attention scores. After normalization by SoftMax, multiple bands are combined and linearly projected to obtain the band-limited attention enhanced features.

[0021] In addition to the improvement of the attention module, the convolutional modules of the temporal path Conformer and the frequency path Conformer are also improved. A convolutional enhancement module is designed, which is composed of a dilated convolutional layer, an activation layer, a depth convolutional layer, a normalization layer, and a residual connection connected in sequence. Compared with the traditional convolutional module, the dilated convolution brings a larger receptive field, and the depth convolution effectively reduces the number of parameters and the computational amount.

[0022] Furthermore, the multi-scale fusion dynamic gating network is a module designed to better fuse the temporal path features and the frequency path features. Among them, the multi-scale fusion dynamic gating network is composed of a multi-scale convolutional module and a dynamic gating module. The multi-scale convolutional module in the multi-scale fusion dynamic gating network is composed of convolutional layers with different-sized convolutional kernels connected in sequence, which is used to extract features in different receptive field ranges and enhance the fusion ability of the temporal path features and the frequency path features. The dynamic gating module in the multi-scale fusion dynamic gating network is used to adaptively control the fusion ratio of the temporal path features and the frequency path features, and is composed of two fully connected layers, a normalization layer, and an activation layer connected in sequence. After the features passed through the multi-scale convolutional module are concatenated, the gating coefficients can be obtained through the dynamic gating module. The temporal path features and the frequency path features are weighted using the gating coefficients to obtain the weighted fusion time-frequency features.

[0023] Furthermore, the feature decoder adopts a symmetric multi-layer transposed convolution structure, and skip connections are introduced between different transposed convolution modules to effectively fuse the shallow and deep feature information, alleviate the possible information loss problem, and thus achieve efficient underwater acoustic signal reconstruction. Each transposed convolution module is composed of a transposed convolutional layer, a batch normalization layer, and an activation layer connected in sequence.

[0024] The number of channels of the decoder is symmetric with that of the encoder and decreases layer by layer to adapt to the reduction requirements of the features. The transposed convolution uses dilated transposed convolution, and the dilation rate increases layer by layer, which can greatly increase the receptive field without significantly increasing the computational amount, enabling the decoder to utilize the time-frequency information in a larger range. Finally, a mask matrix is output. The mask matrix is a complex matrix with the real part in the first channel and the imaginary part in the second channel, and the matrix size is , and the expression is:

[0025] Furthermore, the process of complex spectrum restoration processing in step S7 is as follows: S701. Multiply the mask matrix output by the feature decoder with the real and imaginary parts of the feature tensor to obtain the real and imaginary parts of the denoised underwater acoustic target signal. The expression is as follows:

[0026]

[0027]

[0028] Among them, and are the real and imaginary parts of the feature tensor respectively, and are the real and imaginary parts of the mask matrix output by the feature decoder. The obtained through the operation and are the real and imaginary parts of the denoised underwater acoustic target signal.

[0029] S702. Perform inverse short-time Fourier transform on the real and imaginary parts of the denoised underwater acoustic target signal to restore it into a time-domain signal. The expression is as follows:

[0030] The calculated is the denoised underwater acoustic target signal.

[0031] Compared with the prior art, the present invention has the following advantages and beneficial effects: 1. The present invention constructs a noisy underwater acoustic target dataset with multiple signal-to-noise ratios by randomly combining various measured natural noises in the ocean with underwater acoustic signals, enabling the model to learn richer noise characteristics and improving its denoising effect and generalization ability under different noise conditions. At the same time, the noisy underwater acoustic target dataset is not only applicable to the denoising task, but also can provide higher-quality data input for the underwater acoustic target classification task, enabling the classification network to still maintain high-efficiency and accurate recognition ability when facing the denoised signal in practical applications, thereby improving the performance of the entire underwater signal processing system.

[0032] 2. The feature encoder module of the present invention learns the distribution characteristics of underwater acoustic signals, fully considering the complexity and variability of underwater acoustic noise. The time-frequency adaptive dual-path Conformer module efficiently models the time path and frequency path to improve the adaptability to underwater acoustic scenarios and denoising performance.

[0033] 3. The present invention improves the attention module in the Conformer network, uses a combination of window attention and relative position encoding in the time path, and uses band-limited attention in the frequency path. The local attention can significantly reduce the computational complexity and the number of parameters. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0035] Figure 1 is a flowchart of a method for denoising underwater acoustic signals based on a time-frequency adaptive dual-path Conformer network disclosed in the present invention; Figure 2 is a schematic diagram of the structure of the feature encoder in the present invention; Figure 3 is a schematic diagram of the structure of the time-frequency adaptive dual-path Conformer network in the present invention; Figure 4 is a schematic diagram of the improved window attention mechanism of the time path in the present invention; Figure 5 is a schematic diagram of the band-limited attention mechanism of the frequency path in the present invention; Figure 6 is a schematic diagram of the multi-scale fusion dynamic gating network in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0036] In order to enable those skilled in the art to better understand the solutions of the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts fall within the scope of protection of the present application.

[0037] Referring to "embodiments" in the present application means that the specific features, structures, or characteristics described in conjunction with the embodiments can be included in at least one embodiment of the present application. The phrase appears in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described in the present application can be combined with other embodiments.

[0038] Embodiment 1 Refer to Figure 1, the present invention provides an underwater acoustic signal denoising method based on a time-frequency adaptive dual-path Conformer network, comprising the following steps: Using real ocean environmental noise and underwater acoustic target time-domain signals, selecting natural noises such as sea wind, sea waves, and rain, resampling, segmenting, and normalizing the noise and signals, then performing energy threshold filtering, and generating a noisy dataset by superimposing at different signal-to-noise ratios; frame the noisy dataset, obtain the time-frequency spectrogram through short-time Fourier transform, split it into real and imaginary parts and stack them into a feature tensor; input the feature tensor into a feature encoder composed of multiple convolutional modules, and through dilated convolution, batch normalization, and activation processing, extract intermediate feature representations; after reshaping the dimensions of the intermediate features, input them into the time-path Conformer and the frequency-path Conformer respectively. The time path extracts the time-path features through the first feed-forward network layer, improved multi-head window attention, convolutional enhancement module, and the second feed-forward network layer; the frequency path extracts the frequency-path features through the third feed-forward network layer, band-limited multi-head attention module, convolutional enhancement module, and the fourth feed-forward network layer; input the time-path features and the frequency-path features into a multi-scale fusion dynamic gating network, and through multi-scale convolution and dynamic gating module weighted fusion, obtain the first weighted fusion time-frequency features; input the weighted fusion time-frequency features into a feature decoder composed of deconvolution modules, and generate a mask matrix through non-linear mapping; perform complex spectrum restoration on the mask matrix, operate with the feature tensor to obtain the denoised real and imaginary components, and restore the denoised underwater acoustic time-domain target signal through inverse short-time Fourier transform.

[0039] This embodiment conducts the generation and validity verification of a noisy underwater acoustic target dataset.

[0040] The real ocean environmental noise signal and the original underwater acoustic target signal adopted in this embodiment both originate from the ShipsEar dataset. This dataset divides the collected ship targets into four categories, A, B, C, and D, according to the types of ship targets. At the same time, ocean natural noise is obtained through actual measurement and classified into category E.

[0041] In this embodiment, the original underwater acoustic target signal is selected as the C-class passenger ferry signal. This is because the C-class passenger ferry has the most abundant sample quantity in the dataset, and most of the sample collection distances are within 50 meters, with a low degree of noise interference, and can be approximately regarded as a clean signal, which helps the accuracy and stability of the experiment.

[0042] The noise dataset comes from the E-class environmental noise of the ShipsEar dataset. Among them, three environmental noise data numbered 82, 83, and 84 are selected. Because the collection environments of these three sample data are representative, corresponding to the situations of maximum wind speed, maximum wave speed, and maximum rainfall respectively, and there are detailed collection data records.

[0043] In this embodiment, all audio data are resampled to 16 kHz and segmented into 2 - second segments. In the data pre - processing stage, first, the marine environmental noise and the original underwater acoustic target signal are analyzed and cleaned to ensure data quality. Subsequently, the two types of signals are normalized, and the amplitude is normalized to the interval to reduce the amplitude difference between different signals and ensure effective fusion under different noise conditions.

[0044] In addition, to enhance the effectiveness of the data, an energy - threshold filtering method is used to screen the processed marine noise and target signals. Specifically, first, the energy spectrum of each signal is calculated, and a reasonable energy threshold is set to remove low - energy components and improve the effectiveness of the signal. In this experiment, the energy threshold is set to 0.1, that is, the part of the energy spectrum below 10% of the maximum energy value will be regarded as the low - energy region and weakened or removed by the filter to highlight the main signal features, reduce weak interference, and thus improve the robustness and generalization ability of the denoising model.

[0045] The marine noise and target signals processed by the energy - threshold filtering method are superimposed, and the signal - to - noise ratios are set to - 5 dB, 0 dB, and 5 dB to obtain a noisy underwater acoustic target dataset.

[0046] (1) Validation of effectiveness on the denoising network To evaluate the impact of the constructed noisy underwater acoustic target dataset on the denoising model, in this embodiment, a comparative experiment is conducted to analyze whether this dataset can improve the performance of the denoising model and its generalization ability in different noise environments. In this embodiment, different datasets and different denoising methods are used for comparison, and the specific division is shown in Table 1 below.

[0047] Table 1. Comparison table of different datasets and different denoising methods

[0048] DCCRN is a classic denoising network model. Its main structure includes an encoder, a convolutional layer, a recurrent neural network layer, and a decoder. The network uses the real and imaginary parts of the complex STFT spectrum of the noisy signal as inputs. The encoder consists of multiple complex convolutional layers, and each complex convolutional layer includes two convolutional operations, which act on the real and imaginary parts respectively. The recurrent neural network further processes the output of the encoder to obtain temporal features. The decoder also uses complex convolutional layers to convert the features back to the complex STFT spectrum of the signal, thereby recovering the denoised signal.

[0049] The evaluation metrics for the experiment are the mean squared error (MSE) and the Structural Similarity Index Measure (SSIM). The mean squared error (MSE) measures the Euclidean distance between the predicted value and the target value. By calculating the real and imaginary parts obtained from the short-time Fourier transform of the signal before and after denoising respectively, the similarity of the signal before and after denoising is evaluated. The smaller the MSE, the closer the signal before and after denoising, and the better the denoising effect. SSIM measures similarity by calculating the mean, variance, and covariance. Calculating the SSIM of the enhanced spectrum and the clean spectrum can evaluate whether the spectrum structure is damaged. The calculation formula of SSIM is as follows:

[0050] where, and represent the spectrograms before and after denoising respectively, and are the means, representing the brightness, and are the variances, representing the contrast, is the covariance, representing the structure, and are the introduced constants. The range of SSIM is 0 - 1. When SSIM is closer to 1, it indicates that the spectrum structures are more similar.

[0051] The following are the performance performances of three different training data on the DCCRN network, as shown in Table 2 below. The best results are marked in bold. The experiments were trained on the same epoch, learning rate, optimizer, and dataset to ensure fairness.

[0052] Table 2. Performance Performance of Three Different Training Data on the DCCRN Network

[0053] Through the comparative experiments, the following conclusions can be drawn: The noisy underwater acoustic dataset proposed in the present invention shows significantly lower mean squared error and higher Structural Similarity Index Measure (SSIM) under the same denoising network. This result indicates that the denoising model trained using this dataset can more effectively reduce noise interference, thereby significantly improving the quality of the denoised signal.

[0054] The experimental results show that the noisy underwater acoustic dataset not only helps the training of the denoising model, but also improves the generalization ability of the model in practical applications. Especially when facing different noise types and varying signal-to-noise ratio conditions, it can effectively cope with diverse underwater acoustic environments.

[0055] (2)Validation of Effectiveness on the Classification Network In order to verify that the constructed noisy underwater acoustic target dataset can also play a role in downstream underwater acoustic target classification tasks, provide higher-quality data input for it, optimize the classification ability of the classification network, make the classification network more robust and have better generalization performance, the following experiments were designed.

[0056] The experiment generated multiple mixed datasets with different proportions by mixing the original dataset and the noisy underwater acoustic target dataset in different proportions. Then, the same convolutional neural network was used to perform classification tasks on these mixed datasets, aiming to observe the improvement effect of the noisy underwater acoustic target dataset on classification performance during the training process. This experiment can effectively evaluate the impact of the introduction of noisy data on the robustness and accuracy of the classification network, thereby verifying the important role of the noisy underwater acoustic target dataset in improving underwater acoustic target classification tasks. The mixed datasets are set as shown in Table 3 below.

[0057] Table 3. Division Table of Mixed Datasets

[0058] The experiment used a convolutional neural network as the classification network. The network was composed of three convolutional modules connected in sequence. After the convolutional modules, there were three consecutive fully connected layers. The convolutional module was composed of a convolutional layer, an activation layer, and a pooling layer connected in sequence; among them, the convolutional kernel size of the convolutional layer was , the ReLU activation function was selected, and the max pooling was used in the pooling layer.

[0059] The accuracy (Accuracy) and precision (Precision) were used as evaluation metrics for the classification network. The accuracy represents the proportion of the number of correctly predicted samples in the total number of samples, and the calculation method is to divide the number of correctly classified samples by the total number of samples. The precision represents the proportion of the samples that truly belong to a certain category among those determined to be in that category, and the calculation method is to divide the number of samples correctly classified into a certain category by the total number of samples assigned to this category. The precision pays more attention to the amount of misjudgment.

[0060] Classification experiments were carried out on different mixed datasets using the same convolutional neural network. The results are shown in Table 4 below, and the best results are marked in bold. The experiment was trained on the same epoch, learning rate, optimizer, and dataset to ensure fairness.

[0061] Table 4. Results Table of Classification Experiments on Different Mixed Datasets Using the Same Convolutional Neural Network

[0062] It can be seen from the experimental data that the addition of the noisy underwater acoustic target dataset has an obvious effect on improving the accuracy and precision of the classification task. The introduction of the noisy underwater acoustic target dataset can not only supplement more noise information, but also effectively improve the performance of the classification network, enabling it to maintain a high recognition ability in complex environments. At the same time, the addition of augmented data enhances the generalization ability and robustness of the classification network, enabling it to stably and efficiently classify underwater acoustic signals with different signal-to-noise ratios.

[0063] Example 2 In this example, a denoising performance ablation experiment is carried out based on the time-frequency adaptive dual-path Conformer network.

[0064] The obtained noisy underwater acoustic target dataset is input into the time-frequency adaptive dual-path Conformer network for training to verify the denoising performance of the network.

[0065] First, the noisy underwater acoustic target dataset with multiple signal-to-noise ratios is subjected to feature processing. Time-frequency information is obtained through short-time Fourier transform, split into real and imaginary parts, and the real and imaginary parts of the time-frequency information are stacked in the channel dimension to form a feature tensor. Among them, the size (n_fft) of the FFT of the short-time Fourier transform is set to 512, and the window shift step (hop_length) is set to 128. The size of the obtained feature tensor is , where 247 is the F dimension and 257 is the T dimension; The feature tensor is input into the feature encoder. The structure diagram of the feature encoder is as Figure 2 shown, which contains 5 identical convolutional modules. Each convolutional module consists of a convolutional layer, a normalization layer, and an activation layer connected in sequence. Among them, the normalization layer selects batch normalization, and the activation function of the activation layer selects PReLU. After the output of each convolutional layer is summed with the output of the module through skip connection, it is used as the input of the next convolutional module. Dilated convolutions are used in the convolutional module to gradually increase the receptive field, and the dilation coefficient is set to , the number of channels is set to , the size of the convolutional kernels in the first three layers is set to , and the convolutional kernels in the last two layers are set to . After being processed by the feature encoder, the feature tensor obtains an intermediate feature representation; The intermediate feature representation output by the feature encoder will simultaneously enter the time-path Conformer and the frequency-path Conformer for learning. The structure is as Figure 3As shown. The first feed-forward network layer and the second feed-forward network layer in the Temporal Path Conformer have the same structure, which consists of a normalization layer, a linear expansion layer, an activation layer, a linear compression layer, an activation layer, and a Dropout layer connected in sequence. Among them, the normalization method selects layer normalization, the linear expansion layer expands the input dimension to 256 dimensions, the activation functions of the two activation layers both select GELU, the linear compression layer compresses the 256 dimensions back to the original dimension, and the coefficient of the Dropout layer is set to 0.1; the window size of the improved multi-head window attention is set to 5, the number of attention heads is set to 8, and the calculation process of the improved multi-head window attention is as Figure 4 shown; the convolutional kernel sizes of the dilated convolutional layer and the depth convolutional layer in the convolutional enhancement module are both , and the dilation rate of the dilated convolutional layer is 2; The third feed-forward network layer and the fourth feed-forward network layer in the Frequency Path Conformer have the same structure as that in the Temporal Path Conformer. The input dimension is expanded to 512 dimensions in the linear expansion layer to retain more feature information. The activation functions of the two activation layers select Swish, the linear compression layer compresses the 512 dimensions back to the original dimension, and the coefficient of the Dropout layer is set to 0.2; the width of the banded attention mask in the band-limited multi-head attention module is set to 4, the number of attention heads is set to 8, and the calculation process of the band-limited multi-head attention module is as Figure 5 shown; the convolutional kernel sizes of the dilated convolutional layer and the depth convolutional layer in the convolutional enhancement module are both , and the dilation rate of the dilated convolutional layer is 2; After passing through the Temporal Path Conformer and the Frequency Path Conformer, the obtained temporal path features and frequency path features are input into the multi-scale fusion dynamic gating network for feature fusion. The multi-scale fusion dynamic gating network includes a multi-scale convolutional module and a dynamic gating module, as Figure 6 shown; the first layer of the multi-scale convolutional module uses a convolutional kernel with a size of , and the second layer selects a convolutional kernel with a size of ; both fully connected layers of the dynamic gating module reduce the input dimension to its , the activation function selects Sigmoid, and the output is normalized to between to dynamically adjust the weights; The structure of the feature decoder is similar to that of the feature encoder, including 5 identical convolutional modules, except that the convolutional layer is replaced with a transposed convolutional layer;. The convolutional module uses dilated convolution to gradually increase the receptive field, the dilation coefficient is set to , the number of channels is set to , the convolutional kernel sizes of the first two layers are set to , and the convolutional kernels of the last three layers are set to When performing weighted fusion, the time-frequency features pass through the feature decoder to obtain the mask matrix; Multiply the mask matrix with the real and imaginary parts of the feature tensor to obtain the real and imaginary parts of the denoised underwater acoustic target signal. Finally, perform the inverse short-time Fourier transform to obtain the denoised underwater acoustic time-domain target signal. The size of the FFT (n_fft) is set to 512, and the window shift step (hop_length) is set to 128.

[0066] The loss function of the training network is a hybrid function that combines the time-domain loss, amplitude loss, real-part loss, and imaginary-part loss. The total loss is calculated by adding and taking the logarithm. The calculation formula is as follows:

[0067]

[0068]

[0069]

[0070]

[0071]

[0072] Among them, the L1 loss is used to calculate the time-domain loss, directly constraining the error of the time-domain signal to ensure that the predicted signal is as close as possible to the target signal in the time domain. The amplitude, real part, and imaginary part are calculated using the mean square error (MSE) method, which can make large errors contribute more gradients to promote network optimization. Finally, taking the logarithm of the loss is to prevent a certain loss term from being too large and affecting optimization, and to still make small loss terms contribute significantly.

[0073] To verify the impact of each key module in the time-frequency adaptive dual-path Conformer network on the denoising performance, a series of ablation experiments were designed. Different modules were gradually replaced to analyze their impact on the model performance, and the SSIM was selected as the evaluation index. The specific parameter settings in each module are kept consistent with those in Embodiment 2. All experiments were trained on the same epoch, learning rate, optimizer, and dataset to ensure fairness. Ablation experiments were conducted on datasets with different signal-to-noise ratios, and the results are shown in Table 5 below. The best results are marked in bold.

[0074] Table 5. Results of ablation experiments on datasets with different signal-to-noise ratios

[0075] The results of the above embodiments show that the complete version of the model achieves the optimal SSIM score at all signal-to-noise ratios, verifying its effectiveness. Among them, the dual-path structure, the time-path Conformer, and the frequency-path Conformer have a significant impact on the final performance, while the multi-scale fusion dynamic gating network also plays a positive role in improving SSIM. The effectiveness of the denoising method based on the time-frequency adaptive dual-path Conformer network proposed in the present invention is verified through ablation experiments.

[0076] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0077] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.

Claims

1. An underwater acoustic signal denoising method based on a time-frequency adaptive dual-path Conformer network, characterized in that, The underwater acoustic signal denoising method comprises the following steps: S1. Collect the real ocean environment noise and underwater acoustic target time domain signals and pre-process them. After energy threshold filtering, perform additive superposition to generate a noisy underwater acoustic target data set with multiple signal-to-noise ratios. S2. Perform feature processing on the noisy underwater acoustic target data set with multiple signal-to-noise ratios, obtain the time-frequency information through short-time Fourier transform, split it into real and imaginary parts, and stack the real and imaginary parts of the time-frequency information in the channel dimension to form a feature tensor; S3, inputting the feature tensor into the feature encoder, the feature encoder performs high-level feature encoding on the feature tensor, and extracts an intermediate feature representation with global context information; S4. After resizing the intermediate feature representation, the representation is input into the time-frequency adaptive dual-path Conformer network for learning and training. The time-frequency adaptive dual-path Conformer network includes a time path Conformer and a frequency path Conformer. The time path Conformer is composed of a first feedforward network layer, an improved multi-head window attention module, a convolution enhancement module, and a second feedforward network layer connected in sequence, and the frequency path Conformer is composed of a third feedforward network layer, a band-limited multi-head attention module, a convolution enhancement module, and a fourth feedforward network layer connected in sequence. The intermediate features are simultaneously input into the time path Conformer and the frequency path Conformer to extract the time path features and the frequency path features respectively. S5, inputting the time path features and the frequency path features into the multi-scale fusion dynamic gating network, performing multi-scale convolution on the time path features and the frequency path respectively and then splicing them, and weighted fusion through the dynamic gating module in the multi-scale fusion dynamic gating network to obtain weighted fusion time-frequency features; S6, inputting the weighted fusion time-frequency features into a feature decoder, and the feature decoder performs nonlinear mapping on the weighted fusion time-frequency features to generate a mask matrix; S7. Perform complex spectrum restoration processing on the mask matrix, multiply and add and subtract the mask matrix and the feature tensor to obtain the real and imaginary components after denoising, and then perform inverse short-time Fourier transform to restore the denoised underwater acoustic time domain target signal.

2. The underwater acoustic signal denoising method based on a time-frequency adaptive dual-path Conformer network according to claim 1, wherein, The process of generating a noisy underwater acoustic target data set with multiple signal-to-noise ratios in step S1 is as follows: S101, selecting the recorded ocean environmental noise, selecting natural environmental noise mainly composed of sea breeze, sea waves and rain, pre-processing the real ocean environmental noise signal and the original underwater acoustic target signal, resampling the signal, segmenting it with a fixed time length, retaining a certain overlap time between each segment, and normalizing the signal; S102, performing energy threshold filtering on the processed real ocean environment noise signal and the original target signal, calculating the energy spectrum of the signal, setting the energy threshold, and removing the low-energy components in the energy spectrum; S103, setting a certain signal-to-noise ratio range, randomly selecting real ocean environment noise and signal-to-noise ratio, additively superimposing the target signal with the real ocean environment noise, and obtaining a noisy underwater acoustic target data set with multiple signal-to-noise ratios.

3. The underwater acoustic signal denoising method based on the time-frequency adaptive dual-path Conformer network according to claim 1, wherein, The feature encoder is composed of multiple convolutional modules. Skip connections are used between the convolutional modules. Each convolutional module is composed of a convolutional layer, a batch normalization layer, and an activation layer connected in sequence. Each convolutional layer uses dilated convolutions with gradually increasing dilation factors.

4. A method for denoising underwater acoustic signals based on a time-frequency adaptive dual-path Conformer network according to claim 1, characterized in that In step S4, after the intermediate features are reshaped in size, they are simultaneously input into the temporal path Conformer and the frequency path Conformer. The first feed-forward network layer and the second feed-forward network layer in the temporal path Conformer have the same structure, which is composed of a normalization layer, a linear expansion layer, an activation layer, a linear compression layer, an activation layer, and a Dropout layer connected in sequence, but with different parameters and activation functions. The features output by the first feed-forward network layer enter the improved multi-head window attention module. First, window partitioning is performed, and then each window is separately mapped to a query matrix, a key matrix, and a value matrix. Relative position encoding is combined to calculate the attention scores, and after SoftMax normalization, they are merged and projected to obtain window attention enhanced features. The window attention enhanced features are input into the convolutional enhancement module. The convolutional enhancement module is composed of a dilated convolutional layer, an activation layer, a depth convolutional layer, a normalization layer, and a residual connection connected in sequence. The output passes through the second feed-forward network layer to obtain the temporal path features. The third feed-forward network layer and the fourth feed-forward network layer in the frequency path Conformer have the same structure, which is composed of a normalization layer, a linear expansion layer, an activation layer, a linear compression layer, an activation layer, and a Dropout layer connected in sequence, but with different parameters and activation functions. The output of the third feed-forward network layer enters the band-limited multi-head attention module. First, a dynamic band mask matrix is generated, the features are divided into multiple frequency bands, and each frequency band is separately mapped to a query matrix, a key matrix, and a value matrix. The mask matrix is introduced when calculating the attention scores to obtain the limited attention scores. After SoftMax normalization, they are merged and projected to obtain band-limited attention enhanced features. These features are processed by the convolutional enhancement module. The convolutional enhancement module is composed of a dilated convolutional layer, an activation layer, a depth convolutional layer, a normalization layer, and a residual connection connected in sequence. The output passes through the fourth feed-forward network layer to obtain the frequency path features.

5. The underwater acoustic signal denoising method based on the time-frequency adaptive dual-path Conformer network according to claim 4, characterized in that, The multi-scale fusion dynamic gating network includes a multi-scale convolutional module and a dynamic gating module. Among them, the multi-scale convolutional module is composed of convolutional layers with different-sized convolutional kernels connected in sequence, which processes the temporal path features and the frequency path features respectively. After the outputs are concatenated, they are input into the dynamic gating module. The dynamic gating module is composed of two fully connected layers, a normalization layer, and an activation layer connected in sequence. The concatenated features pass through the dynamic gating module to obtain the gating coefficients, and the temporal path features and the frequency path features are weighted using the gating coefficients to obtain the weighted fusion time-frequency features.

6. The underwater acoustic signal denoising method based on a time-frequency adaptive dual-path Conformer network according to claim 1, characterized in that The feature decoder contains multiple transposed convolutional modules. Each transposed convolutional module is composed of a transposed convolutional layer, a batch normalization layer, and an activation layer connected in sequence. Skip connections are used between the transposed convolutional modules, and they also have skip connections with the corresponding convolutional modules in the feature encoder.

Citation Information

Cited By

  • UUV self-noise suppression method based on dual-channel collaborative noise reduction neural network

    CN120748425A

  • Time sequence prediction method and system based on time-frequency characteristics and dual-channel processing

    CN121278365A

  • Time series prediction method and system based on time-frequency characteristics and double-channel processing

    CN121278365B

  • Low signal-to-noise ratio underwater acoustic line spectrum enhancement system based on deep learning

    CN121636928A

  • Partial discharge signal noise reduction method and system based on time-frequency domain cooperation

    CN121743680A