Pulse signal reconstruction method based on deep learning in non-gaussian environment

By using a dual-path Transformer network based on a time-frequency domain attention mechanism, combined with dilated convolution and jumper structures, the problem of pulse signal reconstruction under non-Gaussian interference in marine environments was solved, achieving high-precision signal reconstruction and improved signal-to-interference-plus-noise ratio.

CN119377887BActive Publication Date: 2025-11-18HARBIN ENG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411501068.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-25
Publication Date
2025-11-18
Estimated Expiration
2044-10-25

AI Technical Summary

Technical Problem

In marine environments, non-Gaussian interference leads to low accuracy in pulse signal reconstruction using traditional methods in the context of low input signal-to-interference-plus-noise ratio. Furthermore, the characteristics of pulse signals are similar to those of non-Gaussian interference, making them difficult to separate effectively.

Method used

A dual-path Transformer network based on a time-frequency domain attention mechanism is adopted, which combines dilated convolution and jumper structure to extract the time-frequency domain feature differences between the pulse signal and non-Gaussian interference through encoder and decoder, thereby realizing signal reconstruction.

Benefits of technology

In non-Gaussian interference environments, it significantly improves the reconstruction accuracy and signal-to-interference-plus-noise ratio gain of pulse signals, demonstrating good generalization ability and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119377887B_ABST
    Figure CN119377887B_ABST
Patent Text Reader

Abstract

The application relates to a deep learning-based pulse signal reconstruction method in a non-Gaussian environment, and relates to a deep learning-based pulse signal reconstruction method.The application aims to solve the problem that the existing marine environment noise often has non-Gaussian statistical characteristics, the traditional method based on the Gaussian distribution assumption has low pulse signal reconstruction accuracy in the non-Gaussian interference and low input signal-to-noise ratio background, and the deep learning-based pulse signal reconstruction method in the non-Gaussian environment is proposed.The specific process of the deep learning-based pulse signal reconstruction method in the non-Gaussian environment is as follows: a training set is constructed; a deep neural network model is constructed; the constructed deep neural network model is trained based on the training set, and a trained deep neural network model is obtained; received underwater acoustic pulse signals are input into the trained deep neural network model, and the trained deep neural network model outputs reconstructed pulse signals.The application is used in the field of pulse signal reconstruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a pulse signal reconstruction method based on deep learning. Background Technology

[0002] The marine environment contains a variety of interference sources, such as the sound emitted by marine organisms such as fish and shrimp, ship navigation and seabed drilling. The interference signals generated by these high-energy interference sources often have non-Gaussian statistical characteristics [1-4]([1]Y.Li,X.Ma,L.Wang,and Y.Liu,“Using deep learning to de-noise and reconstruct pulse signals in non-Gaussian environment,”Journal of Applied Acoustics,vol.40,no.1,pp.131-141,Jan.,2021,doi:10.11684 / j.issn.1000-310X.2021.01.016.[2]X.Mo,H.Wen,andY.Yang,“A parameter estimation method of αstabledistribution and its application in the statistical modeling of ice-generated noise," ActaAcustica, vol.48, no.2, pp.319-326, Mar., 2023, doi:10.15949 / j.cnki.0371-0025.2023.02.001. [3] G. Song, X. Guo, and L. Ma, "αstabledistribution model in oceanambient noise," Acta Acustica, vol.44, no.2, pp.177-188, Mar., 2019, doi:10.15949 / j.cnki.0371-0025.2019.02.004. [4] F. Traverso, G. Vernazza, and A. Trucco, "Simulation of non-White and non-Gaussian underwater ambient Noise, “Oceans-Yeosu, 2012, pp. 1-10.) and leads to a significant reduction in the signal-to-interference-plus-noise ratio of the received signal. Since the statistical characteristics of non-Gaussian interference deviate significantly from the normal distribution and are accompanied by non-stationarity, this causes its second moment to diverge, thus leading to the traditional method based on the Gaussian distribution assumption [5-6] ([5] Z. Zhao, Q. Li, Z. Xia, and D.Shang, "A Single-HydrophoneCoherent-Processing Method for Line-Spectrum Enhancement," Remote Sens., vol.15, no.3, pp.659, Jan., 2023, doi:10.3390 / rs15030659. [6] C.Xing, Y.Wu, L.Xie, and D.Zhang, "Asparse dictionary learning-based denoising method for underwateracoustic Sensors, “Appl.Acoust.,vol.180,pp.108140.1–108140.13,Apr.,2021,doi:10.1016 / j.apacoust.2021.108140.) When facing non-Gaussian interference, its weight iteration process may not converge, and it cannot effectively achieve interference suppression and signal reconstruction. The signal processing performance of traditional methods is significantly reduced [7-9]([7]H.Yang,Y.Cheng,and G.Li,“A denoising method for ship radiated noise based on Spearman variational mode decomposition,spatial-dependence recurrence sampleentropy,improved wavelet threshold denoising,and Savitzky-Golay filter," Alex.Eng.J., vol.60, no.3, pp.3379-3400, Jan., 2021, doi:10.1016 / j.aej.2021.01.055. [8] J.Wang, J.Li, S.Yan, W.Shi, X.Yang, Y.Guo, and T.Gulliver, "Anovel underwateracoustic signal denoising algorithm for Gaussian / Non-Gaussian impulsivenoise," IEEE Trans.Veh.Technol., vol.70, no.1, pp.429-445, Jan., 2021, doi:10.1109 / TVT.2020.3044994. [9] Y. Li, and L.Wang, “A novel noise reduction technique for underwater acoustic signals based on complete ensemble empirical mode decomposition with adaptive noise, minimum mean square variance criterion and least mean square adaptive filter,” Def. Technol., vol. 16, no. 3, pp. 543-554, Jun., 2020, doi:10.1016 / j.dt.2019.07.020.). Furthermore, non-Gaussian interference exhibits significant spike pulse characteristics in its time waveform, closely resembling the characteristics of the target pulse signal, making pulse signal reconstruction even more difficult. How to fully utilize the characteristics of underwater acoustic pulse signals to solve the pulse signal reconstruction problem under low input signal-to-interference-plus-noise ratio and non-Gaussian interference background remains to be addressed.

[0003] Deep neural networks have excellent automatic feature extraction and learning capabilities, enabling them to deeply explore and understand the potential structure and features in complex signals. Through multi-layer nonlinear transformations, they can capture fine information that is difficult to capture by traditional methods. In recent years, signal denoising and reconstruction methods based on deep neural networks [10-15](

[10] X.Wang,Y.Zhao,X.Teng,andW.Sun,“A stacked convolutional sparse denoising autoencoder model for underwater heterogeneous information data,”Appl.Acoust.,vol.167,Oct.,2020,Art.no.107391,doi:10.1016 / j.apacoust.2020.107391.

[11] Y.Song,F.Liu,and T.Shen,“A novel noise reduction technique for underwater acoustic signals based on dual-path recurrent neural network,”IET Commun.,vol.17,no.2,pp.135-144,Oct.,2022,doi:10.1049 / cmu2.12518.

[12] X.Zhou,and K.Yang, "Adenoising representationframework for underwater acoustic signal recognition," J.Acoust.Soc.Am., vol.147, no.4, pp.EL377-EL383, Apr., 2020, doi:10.1121 / 10.0001130.

[13] D.Ju, C.Chi, Z.Li, Y.Li, C.Zhang, and H.Huang, "Deep-learning-based line enhancer for passivesonar systems," IET Radar Sonar Nav., vol.16, no.3, pp.589-601, Dec., 2021, doi:10.1049 / rsn2.12205.

[14] J.Yin, W.Luo, L.Li, X.Han, L.Guo, and J.Wang, “Enhancement of underwater acoustic signal based on denoising automatic-encoder,” Journal on Communications, vol.40, no.10, pp.119-126, Oct., 2019, doi:10.11959 / j.issn.1000-436x.2019181.

[15] R. Zaheer, I. Ahmad, Q. Viet Phung, and D. Habibi, “Blind Source Separation and Denoising of Underwater Acoustic Signals,” IEEE Access, vol.12, pp.80208-80222, Jun., 2024, doi:10.1109 / ACCESS.2024.3410276.) is developing rapidly. Denoising techniques based on deep learning are mainly divided into two categories: direct mapping and mask separation methods. The direct mapping method takes the clean signal as the learning target and directly recovers the clean signal from the noisy signal through an end-to-end neural network model. For example, the literature

[16] (

[16] AANair, and K. Koishida, “Cascaded time+time-frequency Unet for speech enhancement: jointly addressing clipping, codec distortions, and gaps,” in ICASSP, 2021, pp. 7153-7157.) applies the UNet network

[17] (

[17] O. Ronneberger, P. Fischer, and T. Brox, “U-Net: convolutional networks for biomedical image segmentation,” in MICCAI, 2015, pp. 234-241.) to the denoising of the signal. The U-shaped network structure with jumpers effectively integrates shallow and deep features, which significantly improves the denoising effect of the signal. Furthermore, the CBDNet network

[18] (

[18] S.Guo,Z.Yan,K.Zhang,W.Zuo,andM.Wang,“Toward convolutional blind denoising of realphotographs,”in CVPR,2019,pp.1712-1722).The fully convolutional network is combined with the UNet network. The former is responsible for estimating the noise components in the input data, while the latter uses the noise estimate obtained by the fully convolutional network for noise reduction. The RIDNet network

[19] (

[19] S. Anwar, and N. Barnes, “Real image denoising with feature attention,”in ICCV,2019,pp.3155-3164.) introduces residual structure for feature extraction and combines it with attention mechanism to achieve noise reduction of input data, demonstrating the potential of attention mechanism in signal processing. The mask separation method focuses on estimating the mask between the noisy signal and the clean signal, and uses the mask to achieve noise reduction of the input signal and recover a clearer acoustic signal. For example, in reference

[20] (

[20] Y.Song,F.Liu,and T.Shen,“Method of UnderwaterAcoustic Signal Denoising Based on Dual-Path Transformer Network,”IEEEAccess,vol.12,pp.81483-81494,Nov.,2024,doi:10.1109 / ACCESS.2022.3224752.), a mask was constructed using a dual-path Transformer network, and the sequence modeling capability of the Transformer network was used to capture long-term dependencies, thus successfully achieving underwater acoustic signal denoising in the time domain. To further improve the denoising performance of the model, the literature

[21] (J.Fan, J.Yang, X.Zhang, and C.Zheng, “Monaural speech enhancement using U-netfused with multi-head self-attention,”Acta Acustica, vol.47, no.6, pp.703-716, Nov., 2022, doi:10.15949 / j.cnki.0371-0025.2022.06.007.) proposed the TUNet model. This model integrates the multi-scale feature fusion capability of the UNet network with the dual-path Transformer structure of the multi-head attention mechanism. The Transformer structure is used to construct a mask separation structure to achieve time-domain-based single-channel signal denoising. Summary of the Invention

[0004] The purpose of this invention is to address the problem that existing marine environmental noise often has non-Gaussian statistical characteristics, and that traditional methods based on the Gaussian distribution assumption have low accuracy in pulse signal reconstruction under non-Gaussian interference and low input signal-to-interference-plus-noise ratio backgrounds. Therefore, this invention proposes a pulse signal reconstruction method based on deep learning in non-Gaussian environments.

[0005] The specific process of the deep learning-based pulse signal reconstruction method in non-Gaussian environments is as follows:

[0006] Step 1: Construct the training set;

[0007] Step 2: Construct a deep neural network model;

[0008] Step 3: Train the constructed deep neural network model based on the training set to obtain the trained deep neural network model;

[0009] Step 4: Input the received underwater acoustic pulse signal into the trained deep neural network model, and the trained deep neural network model outputs the reconstructed pulse signal.

[0010] The beneficial effects of this invention are as follows:

[0011] Marine environmental noise often exhibits non-Gaussian statistical characteristics due to natural phenomena or human interference. To address the performance degradation of traditional methods based on the Gaussian distribution assumption under non-Gaussian interference and low input signal-to-interference-plus-noise ratio (SINR) conditions, this invention proposes a dual-path Transformer network signal reconstruction method based on a time-frequency domain attention mechanism. The network introduces dilated convolutional layers and jumper structures into the encoder and decoder structures to fuse shallow and deep features in the input data. An attention mechanism combined with the Transformer network structure is used to obtain a separator for mask feature extraction. Considering the differences in the distribution characteristics of pulse signals and non-Gaussian interference in the time and frequency domains, the network reconstructs pulse signals at both global and local scales. Results show that when the input SINR is low, the network performs well in pulse signal reconstruction under both non-Gaussian interference and Gaussian noise environments, demonstrating good generalization ability. Simulation results show that when the input signal-to-interference ratio (SIR) and signal-to-noise ratio (SNR) are -10 dB, the SIR gain of the network proposed in this invention is 38.65 dB and 37.32 dB, respectively, and the signal distortion ratio is 5.80 dB and 4.80 dB, respectively, both of which are superior to the comparison algorithms. Experimental results show that the method of this invention can effectively reconstruct underwater acoustic pulse signals in real physical scenarios, demonstrating the feasibility and effectiveness of the method.

[0012] To further improve signal reconstruction performance, this invention proposes a dual-path Transformer network signal reconstruction method based on a time-frequency domain attention mechanism to address the signal reconstruction problem of underwater acoustic pulse signals under low input signal-to-interference-plus-noise ratio (SINR) non-Gaussian interference. In the time-frequency domain, classic underwater acoustic pulse signals such as continuous wave (CW), linear frequency modulation (LFM), and hyperbolic frequency modulation (HFM) pulse signals all have finite pulse widths and bandwidths, with valuable feature information concentrated in specific time-frequency regions. Unlike the former, non-Gaussian pulse interference has a narrower pulse width and a wider bandwidth. This invention fully utilizes these signal characteristics to design a dual-path Transformer network based on a time-frequency domain attention mechanism, aiming to distinguish between non-Gaussian interference and pulse signals in the time-frequency domain, ultimately achieving high-precision pulse signal reconstruction under non-Gaussian interference backgrounds. The innovations of this invention are as follows:

[0013] (1) Traditional CNN networks are limited by the limited receptive field of convolutional layers and can only capture local information. Therefore, the network of this invention uses dilated convolution combined with the Transformer network structure to expand the receptive field and better capture global information in the input data.

[0014] (2) Due to the information concentration characteristics of underwater acoustic pulse signals in the time and frequency domain, the network of this invention places greater emphasis on the application of attention mechanisms. This invention designs a new attention mechanism to increase the model's time and frequency domain analysis capabilities and the ability to extract signal features in specific time and frequency domain regions, while effectively suppressing non-Gaussian interference.

[0015] (3) Compared with the comparison method, the method of the present invention aims to improve the information extraction capability of the network at both the global and local scales by using dilated convolution, attention mechanism, and jumper structure combined with Transformer network structure to target the time-frequency characteristics of pulse signals, so as to better realize the signal reconstruction of underwater acoustic pulse signals under non-Gaussian interference background.

[0016] This invention simulates two common pulse signals, namely CW signals and LFM pulse signals, superimposed with alpha-stabilized distributed interference and Gaussian white noise to obtain a training set for network training. The generalization ability of the network is verified using simulated HFM pulse signals. This invention verifies the feasibility and effectiveness of the network in pulse signal reconstruction tasks through simulation and experimental data. Attached Figure Description

[0017] Figure 1 This is a flowchart of the present invention;

[0018] Figure 2 This is a physical scene model diagram;

[0019] Figure 3 A deep neural network model structure diagram is constructed for this invention;

[0020] Figure 4 This is a structural diagram of the encoder and decoder of the present invention. Skip Connection is a skip connection structure.

[0021] Figure 5 This is a structural diagram of the jumper structure of the present invention. Skip_Connection_Input is the input of the jumper connection structure, and Skip_Connection_Output is the output of the jumper connection structure.

[0022] Figure 6 This is a structural diagram of the DPSAT Block of the present invention;

[0023] Figure 7 Here is a diagram of the Transformer architecture;

[0024] Figure 8 A diagram of the attention mechanism structure;

[0025] Figure 9 The following diagrams show the signal processing results when the desired signal is a CW signal: (a) input signal time domain, (b) desired signal time domain, and (c) signal time domain of the network reconstructed by the present invention.

[0026] Figure 10 The following diagrams show the signal processing results when the desired signal is an LFM signal: (a) input signal time domain, (b) desired signal time domain, and (c) signal time domain of the network reconstructed by the present invention.

[0027] Figure 11 The following diagrams show the signal processing results when the desired signal is an HFM signal: (a) input signal time domain, (b) desired signal time domain, and (c) signal time domain reconstructed by the network of this invention.

[0028] Figure 12 This is a graph showing the relationship between the input signal-to-interference ratio (SINR) and the output signal-to-interference-plus-noise ratio (SINR) gain.

[0029] Figure 13 The graph shows the relationship between the input signal-to-interference ratio and the signal distortion ratio.

[0030] Figure 14 This is a graph showing the relationship between input signal-to-noise ratio and output signal-to-interference-plus-noise ratio gain.

[0031] Figure 15 A graph showing the relationship between input signal-to-noise ratio and signal distortion ratio;

[0032] Figure 16This is a time-domain waveform diagram of the measured data;

[0033] Figure 17 The following are the signal processing results of the measured data: (a) the signal time domain reconstructed by the network of the present invention, (b) the signal time domain reconstructed by UNet, (c) the signal time domain reconstructed by CBDNet, and (d) the signal time domain reconstructed by RIDNet. Detailed Implementation

[0034] Specific Implementation Method 1: The specific process of the pulse signal reconstruction method based on deep learning in non-Gaussian environments in this implementation method is as follows:

[0035] Step 1: Construct the training set;

[0036] Step 2: Construct a deep neural network model;

[0037] Step 3: Train the constructed deep neural network model based on the training set to obtain the trained deep neural network model;

[0038] Step 4: Input the received underwater acoustic pulse signal into the trained deep neural network model, and the trained deep neural network model outputs the reconstructed pulse signal.

[0039] This invention proposes a dual-path Transformer network model based on a time-frequency domain attention mechanism, aiming to achieve accurate reconstruction of unknown pulse signal waveforms in complex marine environments, especially under conditions where non-Gaussian interference and Gaussian white noise coexist. This invention integrates dilated convolution, attention mechanisms, and jumper structures with the Transformer network structure, effectively extracting local and global features of pulse signals in the time-frequency domain, ultimately achieving accurate modeling of pulse signals in complex marine environments. The key conclusions are as follows:

[0040] (1) Under different input signal-to-interference ratio (SIR) and signal-to-noise ratio (SNR) conditions, the method proposed in this invention outperforms other comparative algorithms in both key indicators, namely signal distortion ratio and SIR gain. This result demonstrates that the method of this invention has higher robustness and accuracy when processing pulse signals in complex noise environments.

[0041] (2) By introducing dilated convolution, attention mechanism, jumper structure and Transformer network structure, the method of the present invention significantly enhances the feature extraction and fusion capability of pulse signals in the time and frequency domain, thereby improving the reconstruction accuracy of pulse signals.

[0042] (3) Through the processing and analysis of the measured data, the method of the present invention further verifies its feasibility and effectiveness in practical applications. Compared with other comparative methods, the method of the present invention exhibits superior performance when processing measured signals in complex marine background environments, fully demonstrating its wide applicability and good generalization ability.

[0043] Specific Implementation Method Two: This implementation method differs from Specific Implementation Method One in that: in step one, the training set X is constructed; the specific process is as follows:

[0044] Physical scene model such as Figure 2 As shown.

[0045] The receiving hydrophone not only receives the desired pulse signal emitted by the transmitting transducer, but also receives marine environmental noise and non-Gaussian interference from sound sources such as marine life, ship navigation and seabed drilling.

[0046] Assume the marine environmental noise is Gaussian white noise and the interference is α-steady distribution noise;

[0047] The acoustic signal is emitted from the transmitting transducer, propagates, and reaches the receiving hydrophone. The time-frequency domain representation of the signal received by the receiving hydrophone is as follows:

[0048] X(t,f)=S(t,f)+N(t,f)+I(t,f) (1)

[0049] Among them, represents the time-frequency domain representation of the signal received by the hydrophone; represents the time-frequency domain representation of the desired pulse signal emitted by the transmitting transducer; represents the time-frequency domain representation of the interference; and represents the time-frequency domain representation of the marine environmental noise.

[0050] t = 1, ..., T, f = 1, ..., F, where T and F are the total number of time frames and frequency frames in the time-frequency domain representation of the input signal X, respectively.

[0051] The other steps and parameters are the same as in Specific Implementation Method 1.

[0052] Specific Implementation Method Three: This implementation method differs from Specific Implementation Method One or Two in that: in step two, a deep neural network model is constructed; the specific process is as follows:

[0053] The overall network structure is as follows Figure 3 As shown;

[0054] Deep neural network models include encoders, separators, and decoders;

[0055] The working process of a deep neural network model is as follows:

[0056] Representing the time-frequency domain of the signal received by the hydrophone As the input signal for the deep neural network model, the input signal has 2 channels, which are the real and imaginary parts of the short-time Fourier transform of the signal, respectively; T and F are the total number of time frames and frequency frames represented by the input signal in the time-frequency domain, respectively. It is a real number;

[0057] The input signal is processed by an encoder to extract signal features (Encoder_Out), and the signal features (Encoder_Out) are then processed by a separator to obtain a signal mask (Mask).

[0058] The product of the signal feature Encoder_Out and the signal mask Mask is used as the input to the decoder, and the decoder outputs the reconstructed signal.

[0059] Other steps and parameters are the same as in specific implementation method one or two.

[0060] Specific Implementation Method Four: This implementation method differs from Specific Implementation Methods One to Three in that the encoder sequentially includes a first dilation block and a ninth unit;

[0061] The first dilation block includes a first unit, a second unit, a third unit, a fourth unit, a fifth unit, a sixth unit, a seventh unit, an eighth unit, a first skip connection structure, a second skip connection structure, and a third skip connection structure.

[0062] Each of the first, second, third, fourth, fifth, sixth, seventh, eighth, and ninth units sequentially includes: a dilated convolutional layer and a ReLU activation function layer;

[0063] The encoder works as follows:

[0064] Input data The first unit is input, and the first unit outputs feature A;

[0065] Feature A is input into the second unit, and the second unit outputs feature B;

[0066] Feature A is input to the first skip connection structure, and the first skip connection structure outputs feature C.

[0067] Feature B is input into the third unit, and the third unit outputs feature D;

[0068] Feature B is input to the second skip connection structure, and the second skip connection structure outputs feature E.

[0069] Feature D is input into the fourth unit, and the fourth unit outputs feature F;

[0070] Feature D is input to the third skip connection structure, and the third skip connection structure outputs feature G;

[0071] Feature F is input into the fifth unit, and the fifth unit outputs feature H;

[0072] Feature H and the third skip connection structure output feature G as inputs to the sixth unit, and the sixth unit outputs feature I;

[0073] Feature I and the second skip connection structure output feature E as input to the seventh unit, and the seventh unit outputs feature J;

[0074] Feature J and the first skip connection structure output feature C are input to the eighth unit, and the eighth unit outputs feature K;

[0075] Feature K is input to the ninth unit, and the ninth unit outputs feature L, which is the encoder output feature Encoder_Out.

[0076] The expression is:

[0077] Encoder_Out=ReLU(Dilated_Conv(Dilation_Block(X))) (2)

[0078] Here, Dilation_Block() is the dilation block, Dilated_Conv is the dilated convolutional layer, and ReLU is the ReLU activation function layer.

[0079] The other steps and parameters are the same as those in specific implementation methods one through three.

[0080] Specific Implementation Method 5: This implementation method differs from Specific Implementation Methods 1 to 4 in that each of the first, second, and third skip connection structures includes a first convolutional layer, a first ReLU activation function layer, a first max pooling layer, a second convolutional layer, a first Tanh activation function layer, a second max pooling layer, a third convolutional layer, and a first Sigmoid activation function layer.

[0081] The working process of each Skip Connection structure in the first Skip Connection structure, the second Skip Connection structure, and the third Skip Connection structure is as follows:

[0082] Input feature A′ is sequentially input into the first convolutional layer and the first ReLU activation function layer, and the first ReLU activation function layer outputs feature B′;

[0083] The input feature A′ is sequentially fed into the first max pooling layer, the second convolutional layer, and the first Tanh activation function layer. The first Tanh activation function layer outputs feature C′.

[0084] The input feature A′ is sequentially fed into the second max pooling layer, the third convolutional layer, and the first Sigmoid activation function layer. The first Sigmoid activation function layer outputs feature D′.

[0085] Multiply features B′, C′, and D′ by the dot product to obtain feature E′;

[0086] Feature E′ serves as the output feature of the Skip Connection structure.

[0087] The other steps and parameters are the same as those in specific implementation methods one through four.

[0088] Specific Implementation Method Six: This implementation method differs from Specific Implementation Methods One to Five in that the separator includes a normalization layer (LayerNorm), a fourth convolutional layer, a first dual-path self-attention mechanism Transformer module (DPSATBlock), a second dual-path self-attention mechanism Transformer module (DPSAT Block), a fifth convolutional layer, a first PReLU activation function layer, a sixth convolutional layer, a second Tanh activation function layer, a seventh convolutional layer, a second Sigmoid activation function layer, an eighth convolutional layer, and a second ReLU activation function layer;

[0089] The working process of the separator is as follows:

[0090] The encoder output feature Encoder_Out is sequentially input into the normalization layer (LayerNorm) and the fourth convolutional layer, and the fourth convolutional layer outputs feature F′;

[0091] Feature F′ is input to the first dual-path self-attention mechanism Transformer module (DPSAT Block), and the first dual-path self-attention mechanism Transformer module (DPSAT Block) outputs feature G′;

[0092] Feature G′ is input to the second dual-path self-attention mechanism Transformer module (DPSAT Block), and the second dual-path self-attention mechanism Transformer module (DPSAT Block) outputs feature H′;

[0093] Feature H′ is sequentially input into the fifth convolutional layer and the first PReLU activation function layer, and the first PReLU activation function layer outputs feature I′;

[0094] Feature I′ is sequentially input into the sixth convolutional layer and the second Tanh activation function layer, and the second Tanh activation function layer outputs feature J′;

[0095] Feature I′ is sequentially input into the seventh convolutional layer and the second Sigmoid activation function layer, and the second Sigmoid activation function layer outputs feature K′;

[0096] The dot product of feature J′ and feature K′ yields feature L′;

[0097] Feature L′ is sequentially input into the eighth convolutional layer and the second ReLU activation function layer, and the second ReLU activation function layer outputs feature M′;

[0098] Feature M′ is the separator output signal mask.

[0099] The other steps and parameters are the same as those in one of the specific implementation methods one to five.

[0100] Specific Implementation Method Seven: This implementation method differs from Specific Implementation Methods One to Six in that each of the first and second dual-path self-attention mechanism Transformer modules (DPSAT Block) includes a time-domain sub-block, a frequency-domain sub-block, a fifteenth convolutional layer, and a second PReLU activation function layer;

[0101] The time-domain sub-block includes: a ninth convolutional layer, a third ReLU activation function layer, a first Transformer structure, a third max pooling layer, a tenth convolutional layer, a third Tanh activation function layer, an eleventh convolutional layer, and a third Sigmoid activation function layer.

[0102] The frequency-domain sub-block includes: a twelfth convolutional layer, a fourth ReLU activation function layer, a second Transformer structure, a fourth max pooling layer, a thirteenth convolutional layer, a fourth Sigmoid activation function layer, a fourteenth convolutional layer, and a fourth Tanh activation function layer;

[0103] The working process of each dual-path self-attention mechanism Transformer module (DPSATBlock) in the first dual-path self-attention mechanism Transformer module (DPSAT Block) and the second dual-path self-attention mechanism Transformer module (DPSAT Block) is as follows:

[0104] 1) Feature Z′ is sequentially input into the ninth convolutional layer and the third ReLU activation function layer, and the third ReLU activation function layer outputs feature N′;

[0105] The feature Z′ is sequentially input into the first Transformer structure and the third max pooling layer, and the third max pooling layer outputs the feature O′.

[0106] Feature O′ is sequentially input into the tenth convolutional layer and the third Tanh activation function layer, and the third Tanh activation function layer outputs feature P′;

[0107] Feature O′ is sequentially input into the eleventh convolutional layer and the third Sigmoid activation function layer, and the third Sigmoid activation function layer outputs feature Q′.

[0108] Multiply features N′, P′, and Q′ by the dot product to obtain feature R′; feature R′ is the output feature of the time-domain sub-block.

[0109] 2) Feature Z′ is sequentially input into the twelfth convolutional layer and the fourth ReLU activation function layer, and the fourth ReLU activation function layer outputs feature S′;

[0110] The feature Z′ is sequentially input into the second Transformer structure and the fourth max pooling layer, and the fourth max pooling layer outputs the feature T′.

[0111] Feature T′ is sequentially input into the thirteenth convolutional layer and the fourth Sigmoid activation function layer, and the fourth Sigmoid activation function layer outputs feature U′.

[0112] Feature T′ is sequentially input into the fourteenth convolutional layer and the fourth Tanh activation function layer, and the fourth Tanh activation function layer outputs feature V′.

[0113] Multiply the features S′, U′, and V′ by a dot product to obtain feature W′; feature W′ is the output feature of the frequency-domain sub-block.

[0114] 3) Input the output features R′ of the time-domain sub-block and W′ of the frequency-domain sub-block into the fifteenth convolutional layer and the second PReLU activation function layer in sequence. The output feature Y′ of the second PReLU activation function layer is used as the output feature of each dual-path self-attention mechanism Transformer block (DPSAT Block).

[0115] The other steps and parameters are the same as those in specific implementation methods one through six.

[0116] Specific Implementation Method Eight: This implementation method differs from Specific Implementation Methods One to Seven in that the working process expression of the separator is as follows:

[0117] The encoder output feature Encoder_Out is sequentially input into the normalization layer (LayerNorm) and the fourth convolutional layer. The fourth convolutional layer outputs the feature DPSAT_In; represented as:

[0118] DPSAT_In=Conv(Layer_Norm(Encoder_Out)) (3)

[0119] Where Layer_Norm represents the normalization layer (LayerNorm);

[0120] DPSAT_In is sequentially input into the R-layer dual-path self-attention mechanism Transformer module (DPSAT Block) to obtain the feature DPSAT_Out, where R=2; This is represented as:

[0121] DPSAT_Out = DPSAT_Block R (…(DPSAT_Block1(DPSAT_In))) (4)

[0122] Among them, DPSAT_Block R DPSAT_Block1 represents the Layer R dual-path self-attention mechanism Transformer module (DPSATBlock), and DPSAT_Block1 represents the Layer 1 dual-path self-attention mechanism Transformer module (DPSAT Block).

[0123] Each of the dual-path self-attention mechanism Transformer modules (DPSAT Block) includes a frequency-domain sub-block and a time-domain sub-block, such as... Figure 6 As shown;

[0124] The working process of frequency-domain sub-block is represented as follows:

[0125] DPSAT_F_ReLU=ReLU(Conv(DPSAT_Out r-1 (5)

[0126] DPSAT_F=Transformer(DPSAT_Out r-1 [:,f,:]),f=1,…,F (6)

[0127] DPSAT_F_Tanh=Tanh(Conv(Maxpool(DPSAT_F))) (7)

[0128] DPSAT_F_Sigmoid=Sigmoid(Conv(Maxpool(DPSAT_F))) (8)

[0129] DPSAT_F_sub=DPSAT_F_ReLU·DPSAT_F_Tanh·DPSAT_F_Sigmoid (9)

[0130] Where DPSAT_F_ReLU represents DPSAT_Out in the frequency-domain sub-block. r-1 The features output after sequentially inputting the convolutional layer and ReLU activation function layer; r = 1, 2; when r = 1, DPSAT_Out r-1 =DPSAT_In;

[0131] DPSAT_Out r-1 [:,f,:] represents DPSAT_Out r-1 Features of each frequency frame; DPSAT_F represents DPSAT_Out r-1 The output of each frequency frame feature after passing through the Transformer structure;

[0132] DPSAT_F_Tanh represents the output of DPSAT_F after passing through a max pooling layer, a convolutional layer, and a Tanh activation function layer in sequence; DPSAT_F_Sigmoid represents the output of DPSAT_F after passing through a max pooling layer, a convolutional layer, and a Sigmoid activation function layer in sequence.

[0133] DPSAT_F_sub represents the dot product of DPSAT_F_ReLU, DPSAT_F_Tanh, and DPSAT_F_Sigmoid;

[0134] The working process of a time-domain sub-block is represented as follows:

[0135] DPSAT_T_ReLU=ReLU(Conv(DPSAT_Out r-1 (10)

[0136] DPSAT_T=Transformer(DPSAT_Οut r-1 [:,:,t]),t=1,…,T (11)

[0137] DPSAT_T_Tanh=Tanh(Conv(Maxpool(DPSAT_T))) (12)

[0138] DPSAT_T_Sigmoid=Sigmoid(Conv(Maxpool(DPSAT_T))) (13)

[0139] DPSAT_T_sub=DPSAT_T_ReLU·DPSAT_T_Tanh·DPSAT_T_Sigmoid (14)

[0140] Wherein, DPSAT_T_ReLU represents DPSAT_Out in the time-domain sub-block. r-1 The features output after sequentially inputting the convolutional layer and ReLU activation function layer; r = 1, 2; when r = 1, DPSAT_Out r-1 =DPSAT_In;

[0141] DPSAT_Οut r-1 [:,:,t] represents DPSAT_Out r-1 Features of each time frame; DPSAT_T represents DPSAT_Out r-1 The output of each time frame feature after passing through the Transformer structure;

[0142] DPSAT_T_Tanh represents the output of DPSAT_T after passing through a max pooling layer, a convolutional layer, and a Tanh activation function layer in sequence; DPSAT_T_Sigmoid represents the output of DPSAT_T after passing through a max pooling layer, a convolutional layer, and a Sigmoid activation function layer in sequence.

[0143] DPSAT_T_sub represents the dot product of DPSAT_T_ReLU, DPSAT_T_Tanh, and DPSAT_T_Sigmoid;

[0144] DPSAT_Out r =PReLU(Conv[DPSAT_F_sub,DPSAT_T_sub]) (15)

[0145] DPSAT_Out r This represents the output feature of the r-th layer dual-path self-attention mechanism Transformer module (DPSATBlock); such as Figure 3 The final Mask is represented as

[0146] Separator_Feature=PReLU(Conv(DPSAT_Out)) (16)

[0147] Mask=ReLU(Conv(Tanh(Conv(Separator_Feature))·Sigmoid(Conv(Separator_Feature)))) (17)

[0148] Where Separator_Feature represents the output of DPSAT_Out after passing through a convolutional layer and a PReLU activation function layer in sequence; · represents dot product.

[0149] The other steps and parameters are the same as those in specific implementation methods one through seven.

[0150] Specific Implementation Method Nine: This implementation method differs from Specific Implementation Methods One to Eight in that the decoder sequentially includes a second dilation block and an eighteenth unit;

[0151] The second dilation block includes the tenth unit, the eleventh unit, the twelfth unit, the thirteenth unit, the fourteenth unit, the fifteenth unit, the sixteenth unit, the seventeenth unit, the fourth skip connection structure, the fifth skip connection structure, and the sixth skip connection structure;

[0152] Each of the tenth, eleventh, twelfth, thirteenth, fourteenth, fifteenth, sixteenth, and seventeenth units sequentially includes: a dilated convolutional layer and a ReLU activation function layer;

[0153] The eighteen units sequentially include: a dilated convolutional layer (Dilated Conv) and a Tanh activation function layer;

[0154] The decoder operates as follows:

[0155] Multiply the encoder output feature Encoder_Out with the separator output Mask to obtain the feature.

[0156] The output feature B of the second unit is multiplied by the output Mask of the separator to obtain the feature.

[0157] The output feature D of the third unit is multiplied by the output Mask of the separator to obtain the feature.

[0158] The output feature F of the fourth unit is multiplied by the output Mask of the separator to obtain the feature.

[0159] The output feature H of the fifth unit is multiplied by the output Mask of the separator to obtain the feature.

[0160] The output feature I of the sixth unit is multiplied by the output Mask of the separator to obtain the feature.

[0161] The output feature J of the seventh unit is multiplied by the output Mask of the separator to obtain the feature.

[0162] The output feature K of the eighth unit is multiplied by the output Mask of the separator to obtain the feature.

[0163] feature Input the tenth unit, the tenth unit outputs features

[0164] feature Input the fourth skip connection structure (Skip Connection), and output the feature δ.

[0165] feature and characteristics Input to Unit 11, Unit 11 outputs features

[0166] feature Given a fifth skip connection structure as input, the fifth skip connection structure outputs feature δ′.

[0167] feature and characteristics Input the twelfth unit, the twelfth unit outputs features

[0168] feature Input a sixth skip connection structure (Skip Connection), output the following features:

[0169] feature and characteristics Input the thirteenth unit, the thirteenth unit outputs features

[0170] feature and characteristics Input the fourteenth unit, the fourteenth unit outputs the features

[0171] feature feature The feature δ″ is input into the fifteenth unit, and the fifteenth unit outputs the feature.

[0172] feature feature The feature δ′ is input into the sixteenth unit, and the sixteenth unit outputs the feature.

[0173] feature feature The feature δ is input into the seventeenth unit, and the seventeenth unit outputs the feature.

[0174] feature Input the 18th unit, the 18th unit outputs features feature This is to provide the decoder with an estimate of the desired signal;

[0175] The working process of each Skip Connection structure in the fourth Skip Connection, fifth Skip Connection, and sixth Skip Connection structures is as follows:

[0176] Input feature A′ is sequentially input into the first convolutional layer and the first ReLU activation function layer, and the first ReLU activation function layer outputs feature B′;

[0177] The input feature A′ is sequentially fed into the first max pooling layer, the second convolutional layer, and the first Tanh activation function layer. The first Tanh activation function layer outputs feature C′.

[0178] The input feature A′ is sequentially fed into the second max pooling layer, the third convolutional layer, and the first Sigmoid activation function layer. The first Sigmoid activation function layer outputs feature D′.

[0179] Multiply features B′, C′, and D′ by the dot product to obtain feature E′;

[0180] Feature E′ serves as the output feature of the Skip Connection structure.

[0181] The other steps and parameters are the same as those in specific implementation methods one through eight.

[0182] Specific Implementation Method 10: This implementation method differs from one of the specific implementation methods 1 to 9 in that: in step 3, the constructed deep neural network model is trained based on the training set to obtain a trained deep neural network model;

[0183] The specific process is as follows:

[0184] The time-frequency domain signal representation X of the hydrophone received in step one is normalized to obtain the normalized data; the expression is:

[0185]

[0186] The normalized data is used as the input to the deep neural network model constructed in step two, and the expected signal is used as the output of the deep neural network model.

[0187] Set the loss function as follows:

[0188]

[0189] Where MSE() represents the mean square error and S represents the desired signal. L(θ) represents the expected signal estimate of the output of the deep neural network model, L(θ) represents the loss function, and θ represents the parameters of the deep neural network model.

[0190]

[0191] Among them, g θ (·) is the mapping function between the input and output of the neural network;

[0192] The network uses the Adaptive Moment Estimation (Adam) method to update the parameters in the network;

[0193]

[0194] Continue until convergence, and you will obtain a trained deep neural network model.

[0195] The other steps and parameters are the same as those in specific implementation methods one through nine.

[0196] Encoder and Decoder

[0197] The structure of the encoder and decoder is as follows: Figure 4 As shown. The encoder and decoder have similar structures, consisting of a series of dilated convolutional layers (DCVs) superimposed with ReLU activation functions. In this invention, the number of convolutional layer filters is 64. Non-Gaussian interference exhibits complex spectral characteristics and dynamic changes over time in the time-frequency domain. Capturing global information helps reveal the inherent patterns of signals in time series and frequency distribution. Dilated convolutional layers (DCVs) effectively increase the receptive field without changing the output feature map size, thus better capturing global information in the input data. The use of dilated convolutional layers (DCVs) helps to identify and distinguish between pulse signals and non-Gaussian interference in the time-frequency domain. However, the superposition of a series of convolutional layers causes the neural network to tend to fit low-frequency features first and then high-frequency features, while also introducing the gradient vanishing problem. This invention adopts... Figure 5The skip connection structure shown performs dynamic feature selection on the skip connection input features. The network uses this skip connection structure to connect the high-frequency and low-frequency features extracted by the encoder and decoder, establishing a direct connection between the encoder and decoder as a whole. Through this multi-scale feature fusion mechanism, the network can more comprehensively capture local details and global structure in the input data, while effectively mitigating the gradient vanishing problem. In summary, the network fully utilizes the high and low-frequency information of the input signal, effectively recognizing and distinguishing non-Gaussian interference from impulse signals while more effectively recovering the detailed structure of impulse signals.

[0198] like Figures 3-4 As shown, the input data can be represented as The encoder produces a feature map, i.e., the encoder output Encoder_Out. Encoder_Out is multiplied by the separator's output Mask and then passed through a decoder to obtain an estimate of the desired signal.

[0199] Separator

[0200] In the separator section, the input DPSAT_In of the DPSAT Block is passed through R layers of DPSAT Blocks to further extract the features of the input signal, thereby obtaining a mask. In this invention, R = 2. The structure of the r-th layer of the DPSAT Block is as follows: Figure 6 As shown.

[0201] The DPSAT Block comprises a time-domain sub-block and a frequency-domain sub-block, simultaneously capturing temporal and frequency information from the model's input features in both the time and frequency dimensions. The DPSAT Block employs a Transformer structure to capture global information in the input signal, effectively distinguishing the distribution differences between non-Gaussian interference and impulse signals in the time-frequency domain, ultimately suppressing non-Gaussian interference while preserving impulse signal information. Subsequently, the DPSAT Block uses max-pooling layers nested with Tanh and Sigmoid activation functions in both the time-domain and frequency-domain sub-blocks to achieve dynamic feature selection in the time-frequency domain. This feature selection is then multiplied by the input signal after ReLU activation to further extract the time-frequency domain features of the input signal. This is an improvement in the attention mechanism of this invention. This mechanism increases the weight of the desired signal region and reduces the weight of background noise and non-Gaussian interference regions, enhancing the network's ability to select important impulse features and effectively suppressing noise and non-Gaussian interference while extracting desired signal features. Finally, the extracted time-frequency domain features are passed through a convolutional layer (Conv) and the PReLU activation function to obtain the output DPSAT_Out of the r-th layer DPSAT Block. r .

[0202] like Figure 3 Encoder_Out is processed by a normalization layer (Layer Norm) and a convolutional layer (Conv) to obtain DPSAT_In. The final DPSAT_Out is obtained from DPSAT_In via an R layer DPSAT Block. The structure of each DPSAT Block is as follows: Figure 6 As shown. Each DPSAT Block layer mainly includes time domain word blocks and frequency domain word blocks. The structures of time-domain sub-blocks and frequency-domain sub-blocks are similar.

[0203] The Transformer structure in the DPSAT Block also includes the SelfAttention attention mechanism. The detailed structures of the Transformer structure and the attention mechanism are as follows: Figures 7-8 As shown.

[0204] In the Transformer architecture, Transformer_In is transformed into Transformer_Out by combining an attention mechanism with a feedforward neural network, residual connections, and layer normalization. The attention mechanism effectively combines information from different locations in the input data by processing multiple attention distributions in parallel, thereby improving the model's feature extraction capabilities. In the SelfAttention mechanism, Self_Attention_In is transformed through h linear transformations to obtain parallel query matrix Q, key matrix K, and value matrix V, where h = 2 in this invention. The output of each head is calculated through parallel attention, the results of each head are merged, and the final Self_Attention_Out is obtained through a Linear layer.

[0205] train

[0206] During training, the parameters in the network are updated using the Adaptive Moment Estimation (Adam) method. The training iterations are 100, and the update process is shown in the following formula.

[0207]

[0208] Where m represents the first-order momentum of the gradient, v represents the second-order momentum of the gradient, β1 and β2 are the exponential decay rates of the first-order and second-order momentum, respectively, lr is the learning rate, and ε is a minimum value to prevent the denominator of equation (24) from being 0 during parameter updates. In this invention, β1 = 0.9, β2 = 0.999, lr = 5e-5, and ε = 1e-8 are taken. As the network parameters are updated, the network output value gradually approaches the label and finally obtains the reconstructed signal.

[0209] Sample set settings

[0210] The network input signal contains pulse signals, Gaussian white noise, and α-stabilized interference. Each signal lasts for 1 second, with a sampling rate of 10 kHz. During training, there are 360 ​​training data sets and 360 test data sets. The input signal-to-noise ratio (SNR) ranges from 0 dB to 10 dB, with a basic interval of 2 dB. The input signal-to-noise interference (SIR) ranges from -10 dB to 0 dB, with a basic interval of 2 dB. In both training and test data, the CW signal frequency is randomly distributed in the range [1000, 4500] Hz, the initial and cutoff frequencies of the LFM pulse signal are randomly distributed in the range [1000, 4500] Hz, the pulse width is randomly distributed in the range [5, 200] ms, and the pulse start time occurs at any point within the range [0, 1] s.

[0211] Evaluation criteria

[0212] To evaluate the network performance, the signal-to-interference-plus-noise ratio (SIR) gain and signal-to-distortion ratio (SDR) are used to assess the network performance, thereby verifying the effectiveness of the network of this invention.

[0213] ①. Signal-to-interference-plus-noise ratio gain

[0214] The signal-to-interference-plus-noise ratio (SINR) is defined as the ratio of the average power of the desired signal to the average power of interference and noise. SINR gain is the difference between the SINR of the reconstructed signal and the SINR of the original signal; SINR gain reflects the signal reconstruction capability of a deep neural network.

[0215]

[0216] SINR gain =SINR out -SINR in (26)

[0217] Among them, P s P represents the desired average signal power, which in this invention is the average power of the pulse signal. in SINR represents the average power of interference and noise. out SINR represents the signal-to-interference-plus-noise ratio (SINR) of the reconstructed signal. in The signal-to-interference-plus-noise ratio (SIR) of the signal before reconstruction is shown.

[0218] ②. Signal distortion ratio

[0219] An important indicator for evaluating the degree of signal distortion after reconstruction is the signal distortion ratio, which is defined as the ratio of the square of the L2 norm of the desired signal to the square of the L2 norm of the difference between the desired signal and the reconstructed signal.

[0220]

[0221] The beneficial effects of the present invention are verified using the following embodiments:

[0222] Simulation Analysis

[0223] This section utilizes a dual-path Transformer network based on a time-frequency domain attention mechanism to process the input signal, thereby verifying the feasibility and generalization of the network in the reconstruction task of pulse signals under non-Gaussian interference environments. The simulation conditions are as follows: the pulse width of the pulse signal is 100ms, the frequency of the CW signal is 2.5kHz, the start frequency of the LFM pulse signal is 1kHz, the end frequency is 1.45kHz, and the start frequency of the HFM pulse signal is 1000Hz, the end frequency is 2000Hz. The input signal signal-to-noise ratio is set to 10dB, and the signal-to-interference ratio is -10dB. The characteristic exponent of the α-stable distributed interference is α = 1.8, the skewness exponent is β = 0, the scaling parameter is γ = 1, and the position parameter is a = 0. Figures 9-11 These are the time-domain waveforms of the input signal, the desired signal, and the reconstructed signal when the desired signal is a CW signal, an LFM pulse signal, or an HFM pulse signal, respectively.

[0224] like Figures 9-10 As shown, before processing by the signal reconstruction network, the desired signal in the time domain is almost completely submerged in background noise and interference. After the input signal is processed by a dual-path Transformer network based on a time-frequency domain attention mechanism, the reconstructed signal effectively retains the characteristics of the original pulse signal while most of the interference is eliminated. The network of this invention effectively removes α-stable distribution interference and Gaussian white noise, achieving a significant effect in canceling interference and effectively enhancing the target pulse signal.

[0225] pass Figure 11 It can be seen that when the type of the desired signal is different from that of the training data, the dual-path Transformer network based on the time-frequency domain attention mechanism can still restore the desired signal that has been almost submerged in the background noise, achieving a reconstruction effect with almost no distortion. This shows that the network of the present invention has a certain degree of generalization. The experimental results prove the wide application potential of the network in the pulse signal reconstruction task.

[0226] To systematically analyze the signal reconstruction performance of the network of this invention, ablation experiments were conducted to verify the roles of the jumper structure, expanded convolution, Transformer network structure, and attention mechanism in the method of this invention. Simulation conditions were the same as above, and four comparative models were introduced. Comparative Model 1: The attention mechanism in the DPSAT Block module was removed; Comparative Model 2: The Transformer network structure in the network of this invention was removed; Comparative Model 3: The jumper structure in the encoder and decoder modules of the network of this invention was replaced with a normal jumper structure; Comparative Model 4: The expanded convolutional layer in the encoder and decoder modules of the network of this invention was replaced with a normal convolutional layer. Tables 1-2 show the signal-to-interference-plus-noise ratio (SIR) gain and signal distortion ratio of the network of this invention and the comparative models under different input SIR and SNR conditions.

[0227] Table 1. Signal-to-Interference-Ratio (SIR) Gain and Signal Distortion Ratio of the Network of the Present Invention and the Comparative Model under Different SIRs.

[0228]

[0229] Table 2. Signal-to-Interference-Ratio Gain and Signal Distortion Ratio of the Network of the Present Invention and the Comparative Model under Different Signal-to-Noise Ratios.

[0230]

[0231]

[0232] The ablation experiments demonstrate that the network model designed in this invention exhibits significant advantages over the four comparative models in pulse signal reconstruction, particularly in its superior modeling ability for transient signals. Dilated convolutions and the Transformer structure increase the receptive field, enabling the network to capture dependencies over longer distances, which is crucial for reconstructing subtle structures and transient changes in the signal. The introduction of attention mechanisms and jumper structures further enhances the network's ability to select important features. The attention mechanism dynamically adjusts the weights of different feature regions, allowing the network to focus more on key information beneficial to signal reconstruction, thereby effectively suppressing noise and interference. The jumper structure enables dynamic selection of input features and the fusion of global and local features. In the ablation experiments, removing the four structures resulted in varying degrees of decrease in the network's signal-to-interference-plus-noise ratio (SNR) gain and signal-to-distortion ratio (SCR), fully demonstrating the important role of the network in improving the accuracy of pulse signal reconstruction. Notably, the network of this invention outperforms the comparative models under different input SNR and input SNR conditions, indicating that the network maintains high stability and robustness regardless of changes in the signal environment. In summary, the network of this invention outperforms the comparison model in terms of signal-to-interference-plus-noise ratio gain and signal-to-distortion ratio under different signal-to-noise ratios and signal-to-interference-plus-noise ratios, which is sufficient to demonstrate the effectiveness of the dilated convolutional layer, attention mechanism, Transformer structure and jumper structure.

[0233] To further examine the effectiveness of the proposed method, the performance of the proposed algorithm is compared and analyzed with that of three deep neural network denoising algorithms: UNet, CBDNet, and RIDNet. Curves showing the relationship between signal-to-interference-plus-noise ratio (SINR) gain and input SINR, as well as the relationship between signal distortion ratio and input SINR, are plotted. Figures 12-15 As shown.

[0234] Figures 12-15The results show that the input signal-to-interference ratio (SIR) and signal-to-noise ratio (SNR) have a certain impact on the signal reconstruction performance of the network of this invention. As the input SIR and SNR increase, the SIR gain decreases, while the signal distortion ratio increases significantly with increasing input SIR. Analysis indicates that as the input SIR and SNR increase, the performance of the dual-path Transformer network based on the time-frequency domain attention mechanism is less affected by interference and noise, and the signal distortion decreases. Therefore, the input SIR and SNR should not be too low to ensure good signal reconstruction results. In summary, the method of this invention exhibits good stability and robustness over a wide range of input SIR and SNR, effectively reducing interference and noise components in the input signal while preserving the pulse signal characteristics. The SIR gain and signal distortion ratio of the method of this invention are superior to the other three comparative methods. This is because the method of this invention fully combines the local and global features of the input signal, enabling the network to fully consider the signal's variation characteristics in local regions and the overall trend in global features during signal reconstruction, thereby achieving a more comprehensive and accurate modeling of the pulse signal.

[0235] To examine the feasibility of the method of this invention in real-world physical scenarios, this section utilizes a dual-path Transformer network based on a time-frequency domain attention mechanism to process measured data, further verifying the effectiveness and feasibility of the method. In this experiment, the CW signal has a period of 1s, a pulse width of 10ms, a center frequency of 3kHz, and a sampling rate of 10kHz. Figure 16 The time-domain waveform of the measured signal. Figure 17 (a)-(d) show the reconstructed time-domain waveforms of the signal obtained by processing the measured data using four methods: the method of this invention, UNet, CBDNet, and RIDNet. Table 3 shows the statistical results of the signal-to-interference-plus-noise ratio gain and signal distortion ratio of the network of this invention and the comparison network when processing the measured data.

[0236] Table 3. Statistical results of signal-to-interference-plus-noise ratio gain and signal distortion ratio of the proposed network and the comparative network when processing measured data.

[0237]

[0238] Depend on Figure 16-17 It is known that the pulse signals in the measured data are not only significantly interfered with by non-Gaussian distributed impulse noise, but the background noise also exhibits non-uniform characteristics in the frequency domain, specifically, the noise intensity in the low-frequency band is significantly higher than that in the high-frequency band. Under this complex environment, the method of this invention achieves good signal reconstruction results on the experimental data, further verifying the generalization ability of the method and demonstrating that the dual-path Transformer network based on the time-frequency domain attention mechanism can be applied to the processing of measured data. Figures 16-17It can be seen that UNet, CBDNet, and RIDNet can all effectively reconstruct pulse signals when processing experimental data, but the method of this invention performs better in accurately distinguishing the target signal from background interference. The dual-path Transformer network based on the time-frequency domain attention mechanism, through its unique time-frequency domain attention mechanism, almost completely eliminates the interference and noise components in the reconstructed signal. Table 3 shows that the dual-path Transformer network based on the time-frequency domain attention mechanism has better performance in processing measured signals, further verifying the feasibility and effectiveness of the method of this invention.

[0239] This invention may have other embodiments. Without departing from the spirit and essence of this invention, those skilled in the art can make various corresponding changes and modifications according to this invention, but these corresponding changes and modifications should all fall within the protection scope of the appended claims.

Claims

1. A deep learning-based pulse signal reconstruction method for non-Gaussian environments, characterized by: The specific process of the method is as follows: Step 1: Construct the training set; Step 2: Construct a deep neural network model; Step 3: Train the constructed deep neural network model based on the training set to obtain the trained deep neural network model; Step 4: Input the received underwater acoustic pulse signal into the trained deep neural network model, and the trained deep neural network model outputs the reconstructed pulse signal; Step two involves constructing a deep neural network model; the specific process is as follows: Deep neural network models include encoders, splitters, and decoders; The working process of a deep neural network model is as follows: Representing the time-frequency domain of the signal received by the hydrophone As the input signal for the deep neural network model, the input signal has 2 channels, which are the real and imaginary parts of the short-time Fourier transform of the signal, respectively; T and F are the total number of time frames and frequency frames represented by the input signal in the time-frequency domain, respectively. It is a real number; The input signal is processed by an encoder to extract signal features (Encoder_Out), and the signal features (Encoder_Out) are then processed by a separator to obtain a signal mask (Mask). The product of the signal feature Encoder_Out and the signal mask Mask is used as the input to the decoder, and the decoder outputs the reconstructed signal. The encoder comprises, in sequence, a first expansion module and a ninth unit; The first expansion module includes a first unit, a second unit, a third unit, a fourth unit, a fifth unit, a sixth unit, a seventh unit, an eighth unit, a first jump connection structure, a second jump connection structure, and a third jump connection structure; Each of the first, second, third, fourth, fifth, sixth, seventh, eighth, and ninth units sequentially includes: a dilated convolutional layer and a ReLU activation function layer; The encoder operates as follows: Input data The first unit is input, and the first unit outputs feature A; Feature A is input into the second unit, and the second unit outputs feature B; Feature A is input to the first skip connection structure, and the first skip connection structure outputs feature C; Feature B is input into the third unit, and the third unit outputs feature D; Feature B is input to the second skip connection structure, and the second skip connection structure outputs feature E. Feature D is input into the fourth unit, and the fourth unit outputs feature F; Feature D is input to the third hop connection structure, and the third hop connection structure outputs feature G; Feature F is input into the fifth unit, and the fifth unit outputs feature H; Feature H and the third jump connection structure output feature G as input to the sixth unit, and the sixth unit outputs feature I; Feature I and the second hop connection structure output feature E, which is input to the seventh unit; the seventh unit outputs feature J. Feature J and the first jump connection structure output feature C, which is input to the eighth unit. The eighth unit outputs feature K. Feature K is input to the ninth unit, and the ninth unit outputs feature L, which is the encoder output feature Encoder_Out. The expression is: Encoder_Out = ReLU(Dilated_Conv(Dilation_Block(X)))(2) where Dilation_Block() is the dilation block, Dilated_Conv is the dilated convolutional layer, and ReLU is the ReLU activation function layer; The separator includes a normalization layer, a fourth convolutional layer, R dual-path self-attention mechanism Transformer modules, a fifth convolutional layer, a first PReLU activation function layer, a sixth convolutional layer, a second Tanh activation function layer, a seventh convolutional layer, a second Sigmoid activation function layer, an eighth convolutional layer, and a second ReLU activation function layer. The working process of the separator is as follows: The encoder output feature Encoder_Out is sequentially input into the normalization layer and the fourth convolutional layer, and the fourth convolutional layer outputs feature F′. Feature F′ is sequentially input into R dual-path self-attention mechanism Transformer modules, and the R dual-path self-attention mechanism Transformer modules output feature H′; Feature H′ is sequentially input into the fifth convolutional layer and the first PReLU activation function layer, and the first PReLU activation function layer outputs feature I′; Feature I′ is sequentially input into the sixth convolutional layer and the second Tanh activation function layer, and the second Tanh activation function layer outputs feature J′; Feature I′ is sequentially input into the seventh convolutional layer and the second Sigmoid activation function layer, and the second Sigmoid activation function layer outputs feature K′; The dot product of feature J′ and feature K′ yields feature L′; Feature L′ is sequentially input into the eighth convolutional layer and the second ReLU activation function layer, and the second ReLU activation function layer outputs feature M′; Feature M′ is the separator output signal mask; The decoder includes, in sequence, a second expansion module and an eighteenth unit; The second expansion module includes a tenth unit, an eleventh unit, a twelfth unit, a thirteenth unit, a fourteenth unit, a fifteenth unit, a sixteenth unit, a seventeenth unit, a fourth jump connection structure, a fifth jump connection structure, and a sixth jump connection structure; Each of the tenth, eleventh, twelfth, thirteenth, fourteenth, fifteenth, sixteenth, and seventeenth units sequentially includes: a dilated convolutional layer and a ReLU activation function layer; The eighteen units sequentially include: dilated convolutional layers and Tanh activation function layers; The decoder operates as follows: Multiply the encoder output feature Encoder_Out with the separator output Mask to obtain the feature. The output feature B of the second unit is multiplied by the output Mask of the separator to obtain the feature. The output feature D of the third unit is multiplied by the output Mask of the separator to obtain the feature. The output feature F of the fourth unit is multiplied by the output Mask of the separator to obtain the feature. The output feature H of the fifth unit is multiplied by the output Mask of the separator to obtain the feature. The output feature I of the sixth unit is multiplied by the output Mask of the separator to obtain the feature. The output feature J of the seventh unit is multiplied by the output Mask of the separator to obtain the feature. The output feature K of the eighth unit is multiplied by the output Mask of the separator to obtain the feature. feature Input the tenth unit, the tenth unit outputs features feature Input the fourth hop connection structure, and the fourth hop connection structure outputs the feature δ; feature and characteristics Input to Unit 11, Unit 11 outputs features feature Input the fifth hop connection structure, and the fifth hop connection structure outputs the feature δ′; feature and characteristics Input the twelfth unit, the twelfth unit outputs features feature Input the sixth hop connection structure, and the sixth hop connection structure outputs the feature δ″. feature and characteristics Input the thirteenth unit, the thirteenth unit outputs features feature and characteristics Input the fourteenth unit, the fourteenth unit outputs the features feature feature The feature δ″ is input into the fifteenth unit, and the fifteenth unit outputs the feature. feature feature The feature δ′ is input into the sixteenth unit, and the sixteenth unit outputs the feature. feature feature The feature δ is input into the seventeenth unit, and the seventeenth unit outputs the feature. feature Input the 18th unit, the 18th unit outputs features feature This is to provide the decoder with an estimate of the desired signal; The working process of each of the fourth, fifth, and sixth jump connection structures is as follows: Input feature A′ is sequentially input into the first convolutional layer and the first ReLU activation function layer, and the first ReLU activation function layer outputs feature B′; The input feature A′ is sequentially fed into the first max pooling layer, the second convolutional layer, and the first Tanh activation function layer. The first Tanh activation function layer outputs feature C′. The input feature A′ is sequentially fed into the second max pooling layer, the third convolutional layer, and the first Sigmoid activation function layer. The first Sigmoid activation function layer outputs feature D′. Multiply features B′, C′, and D′ by the dot product to obtain feature E′; Feature E′ serves as the output feature of the skip connection structure.

2. The pulse signal reconstruction method based on deep learning in a non-Gaussian environment according to claim 1, characterized in that: The construction of the training set in step one is as follows: The receiving hydrophone not only receives the desired pulse signal emitted by the transmitting transducer, but also receives marine environmental noise and interference at the same time; Assume the marine environmental noise is Gaussian white noise and the interference is α-steady distribution noise; The acoustic signal is emitted from the transmitting transducer, propagates, and reaches the receiving hydrophone. The time-frequency domain representation of the signal received by the receiving hydrophone is X(t,f)=S(t,f)+N(t,f)+I(t,f) (1) where X is the time-frequency domain representation of the signal received by the receiving hydrophone; S is the time-frequency domain representation of the desired pulse signal emitted by the transmitting transducer; I is the time-frequency domain representation of the interference; and N is the time-frequency domain representation of the marine environmental noise. t = 1, ..., T, f = 1, ..., F, where T and F are the total number of time frames and frequency frames in the time-frequency domain representation of the input signal X, respectively.

3. The pulse signal reconstruction method based on deep learning in a non-Gaussian environment according to claim 2, characterized in that: Each of the first, second, and third skip connection structures includes a first convolutional layer, a first ReLU activation function layer, a first max pooling layer, a second convolutional layer, a first Tanh activation function layer, a second max pooling layer, a third convolutional layer, and a first Sigmoid activation function layer. The working process of each of the first, second, and third jump connection structures is as follows: Input feature A′ is sequentially input into the first convolutional layer and the first ReLU activation function layer, and the first ReLU activation function layer outputs feature B′; The input feature A′ is sequentially fed into the first max pooling layer, the second convolutional layer, and the first Tanh activation function layer. The first Tanh activation function layer outputs feature C′. The input feature A′ is sequentially fed into the second max pooling layer, the third convolutional layer, and the first Sigmoid activation function layer. The first Sigmoid activation function layer outputs feature D′. Multiply features B′, C′, and D′ by the dot product to obtain feature E′; Feature E′ serves as the output feature of the skip connection structure.

4. The pulse signal reconstruction method based on deep learning in a non-Gaussian environment according to claim 3, characterized in that: Each of the R dual-path self-attention mechanism Transformer modules includes a time domain sub-block and a frequency domain sub-block, a fifteenth convolutional layer, and a second PReLU activation function layer. The temporal sub-block includes: a ninth convolutional layer, a third ReLU activation function layer, a first Transformer structure, a third max pooling layer, a tenth convolutional layer, a third Tanh activation function layer, an eleventh convolutional layer, and a third Sigmoid activation function layer. The frequency domain sub-block includes: a twelfth convolutional layer, a fourth ReLU activation function layer, a second Transformer structure, a fourth max pooling layer, a thirteenth convolutional layer, a fourth Sigmoid activation function layer, a fourteenth convolutional layer, and a fourth Tanh activation function layer; The working process of each of the R dual-path self-attention mechanism Transformer modules is as follows: 1) Feature Z′ is sequentially input into the ninth convolutional layer and the third ReLU activation function layer, and the third ReLU activation function layer outputs feature N′; The feature Z′ is sequentially input into the first Transformer structure and the third max pooling layer, and the third max pooling layer outputs the feature O′. Feature O′ is sequentially input into the tenth convolutional layer and the third Tanh activation function layer, and the third Tanh activation function layer outputs feature P′; Feature O′ is sequentially input into the eleventh convolutional layer and the third Sigmoid activation function layer, and the third Sigmoid activation function layer outputs feature Q′. Multiply the features N′, P′, and Q′ by the dot product to obtain feature R′; feature R′ is the output feature of the time domain sub-block. 2) Feature Z′ is sequentially input into the twelfth convolutional layer and the fourth ReLU activation function layer, and the fourth ReLU activation function layer outputs feature S′; The feature Z′ is sequentially input into the second Transformer structure and the fourth max pooling layer, and the fourth max pooling layer outputs the feature T′. Feature T′ is sequentially input into the thirteenth convolutional layer and the fourth Sigmoid activation function layer, and the fourth Sigmoid activation function layer outputs feature U′. Feature T′ is sequentially input into the fourteenth convolutional layer and the fourth Tanh activation function layer, and the fourth Tanh activation function layer outputs feature V′. Multiply the features S′, U′, and V′ by a dot product to obtain feature W′; feature W′ is the output feature of the frequency domain sub-block. 3) Input the time domain sub-block output feature R′ and the frequency domain sub-block output feature W′ into the fifteenth convolutional layer and the second PReLU activation function layer in sequence. The second PReLU activation function layer outputs feature Y′, and feature Y′ is used as the output feature of each dual-path self-attention mechanism Transformer module.

5. The pulse signal reconstruction method based on deep learning in a non-Gaussian environment according to claim 4, characterized in that: The working process expression of the separator is as follows: The encoder output feature Encoder_Out is sequentially input into the normalization layer and the fourth convolutional layer, and the fourth convolutional layer outputs the feature DPSAT_In; represented as: DPSAT_In=Conv(Layer_Norm(Encoder_Out))(3) where, Layer_Norm represents the normalization layer; DPSAT_In is sequentially input into the R-layer dual-path self-attention mechanism Transformer module to obtain the feature DPSAT_Out, where R=2; This is represented as: DPSAT_Out = DPSAT_Block R (…(DPSAT_Block1(DPSAT_In)))(4) where DPSAT_Block R DPSAT_Block1 represents the Transformer module with dual-path self-attention mechanism at layer R, and DPSAT_Block1 represents the Transformer module with dual-path self-attention mechanism at layer 1. Each dual-path self-attention mechanism Transformer module includes a frequency domain sub-block and a time domain sub-block; The working process of the frequency domain sub-block is represented as follows: DPSAT_F_ReLU=ReLU(Conv(DPSAT_Out r-1 ))(5) DPSAT_F=Transformer(DPSAT_Out r-1 [:,f,:]),f=1,…,F(6) DPSAT_F_Tanh=Tanh(Conv(Maxpool(DPSAT_F)))(7) DPSAT_F_Sigmoid=Sigmoid(Conv(Maxpool(DPSAT_F)))(8) DPSAT_F_sub=DPSAT_F_ReLU·DPSAT_F_Tanh·DPSAT_F_Sigmoid(9) where DPSAT_F_ReLU represents DPSAT_Out in the frequency domain sub-block r-1 The features output after sequentially inputting into the convolutional layer and the ReLU activation function layer; r = 1, 2; DPSAT_Out r-1 [:,f,:] represents DPSAT_Out r-1 Features of each frequency frame; DPSAT_F represents DPSAT_Out r-1 The output of each frequency frame feature after passing through the Transformer structure; DPSAT_F_Tanh represents the output of DPSAT_F after passing through a max pooling layer, a convolutional layer, and a Tanh activation function layer in sequence; DPSAT_F_Sigmoid represents the output of DPSAT_F after passing through a max pooling layer, a convolutional layer, and a Sigmoid activation function layer in sequence. DPSAT_F_sub represents the dot product of DPSAT_F_ReLU, DPSAT_F_Tanh, and DPSAT_F_Sigmoid; The working process of the time domain sub-block is represented as follows: DPSAT_T_ReLU=ReLU(Conv(DPSAT_Out r-1 ))(10) DPSAT_T=Transformer(DPSAT_Οut r-1 [:,:,t]),t=1,…,T(11) DPSAT_T_Tanh=Tanh(Conv(Maxpool(DPSAT_T)))(12) DPSAT_T_Sigmoid=Sigmoid(Conv(Maxpool(DPSAT_T)))(13) DPSAT_T_sub=DPSAT_T_ReLU·DPSAT_T_Tanh·DPSAT_T_Sigmoid(14) where DPSAT_T_ReLU represents DPSAT_Out in the time domain subblock r-1 The features output after sequentially inputting into a convolutional layer and a ReLU activation function layer; DPSAT_Οut r-1 [:,:,t] represents DPSAT_Out r-1 Features of each time frame; DPSAT_T represents DPSAT_Out r-1 The output of each time frame feature after passing through the Transformer structure; DPSAT_T_Tanh represents the output of DPSAT_T after passing through a max pooling layer, a convolutional layer, and a Tanh activation function layer in sequence; DPSAT_T_Sigmoid represents the output of DPSAT_T after passing through a max pooling layer, a convolutional layer, and a Sigmoid activation function layer in sequence. DPSAT_T_sub represents the dot product of DPSAT_T_ReLU, DPSAT_T_Tanh, and DPSAT_T_Sigmoid; DPSAT_Out r =PReLU(Conv[DPSAT_F_sub,DPSAT_T_sub])(15) DPSAT_Out r This represents the output feature of the r-th layer dual-path self-attention mechanism Transformer module; r = 1, 2; The final Mask is represented as Separator_Feature = PReLU(Conv(DPSAT_Out))(16) Mask = ReLU(Conv(Tanh(Conv(Separator_Feature))●Sigmoid(Conv(Separator_Feature))))(17) where Separator_Feature represents the output of DPSAT_Out after passing through a convolutional layer and a PReLU activation function layer in sequence; ● represents dot product.

6. The pulse signal reconstruction method based on deep learning in a non-Gaussian environment according to claim 5, characterized in that: In step three, the constructed deep neural network model is trained based on the training set to obtain a trained deep neural network model; the specific process is as follows: The time-frequency domain signal representation X of the hydrophone received in step one is normalized to obtain the normalized data; the expression is: The normalized data is used as the input to the deep neural network model constructed in step two, and the expected signal is used as the output of the deep neural network model. Set the loss function as follows: Where MSE() represents the mean square error and S represents the desired signal. L(θ) represents the expected signal estimate of the output of the deep neural network model, L(θ) represents the loss function, and θ represents the parameters of the deep neural network model. Among them, g θ (·) is the mapping function between the input and output of the neural network; The network uses adaptive moment estimation to update the parameters in the network; Continue until convergence, and you will obtain a trained deep neural network model.