Underwater sound signal blind source separation method and system based on double-path parallel gating network

By adopting a dual-path parallel gated network and a band splitting strategy in underwater blind source separation, combined with a bidirectional long and short-term memory network, the problems of unstable separation performance and high computational complexity in traditional methods are solved, and a more efficient and robust underwater blind source separation effect is achieved.

CN120089155APending Publication Date: 2025-06-03HARBIN ENG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510249790.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The traditional underwater blind source separation method faces problems such as unstable separation performance and high computational complexity when processing complex mixed signals.

Method used

The blind source separation method of water acoustic signals based on dual-path parallel gated network is adopted, combining the split band strategy, parallel gated network and bidirectional long and short-term memory network, and the feature dependence of time series and frequency band dimensions is captured through dual-path modeling.

Benefits of technology

It significantly improves the accuracy and efficiency of underwater blind source separation, reduces the computational complexity, enhances the robustness of the system, and improves the separation accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120089155A_ABST
    Figure CN120089155A_ABST
Patent Text Reader

Abstract

The invention discloses an underwater sound signal blind source separation method and system based on a double-path parallel gating network, and the method comprises the steps: collecting a data set of an original underwater sound signal, and converting the data set into a spectrogram; using the spectrogram to train a BSM-DPGN model; and inputting an underwater sound signal to be separated into the BSM-DPGN model to complete the blind source separation of the underwater sound signal. According to the method, the problems of multiband feature extraction, time sequence dependent modeling, signal reconstruction and the like in underwater blind source separation are effectively solved, the application performance of the model in an actual underwater environment is improved, and a new technical scheme is provided for the field of underwater acoustic signal processing. According to the underwater acoustic signal blind source separation method based on the double-path parallel gating network, the ship noise separation task can be effectively completed, and the separation precision and robustness are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of underwater acoustic signal processing, and particularly to an underwater acoustic signal blind source separation method and system based on a dual-path parallel gating network. Background Art

[0002] The background of underwater blind source separation research mainly stems from the increasing human activities in the ocean, and the radiated noise of ships has become a key basis for underwater target detection. However, underwater detection signals are often complex mixtures of multiple noise sources, including marine organisms, sea waves, ships, submarines, etc. These signals have the characteristics of non-stationarity, non-linearity and multi-source aliasing, and are affected by the multi-path effect and noise band overlap of the ocean environment, resulting in increased difficulty in signal separation. Traditional blind source separation methods usually face problems such as unstable separation performance and high computational complexity when dealing with underwater signals.

[0003] In recent years, deep learning techniques have been widely applied in blind source separation tasks. In particular, the recurrent neural network (RNN) has become a research hotspot due to its excellent performance in processing time series data.

[0004] In recent years, scholars at home and abroad have conducted in-depth analysis and research on blind source separation. Among the existing literature, the most famous and effective methods mainly include: 1. Joint optimization of masks and deep recurrent neural networks for monaural source separation: In 2015, HUANG P-S, KIM M, HASEGAWA-JOHNSON M A, et al. Joint optimization of masks and deep recurrent neural networks for monaural source separation [J]. IEEE / ACM Transactions on Audio, Speech, and Language Processing, 2015, 23(12): 2136-2147.. explored the joint optimization of masking functions and deep recurrent neural networks for monaural source separation tasks, including monaural speech separation, monaural singing voice separation, and speech denoising. The joint optimization of deep recurrent neural networks and an additional masking layer enforces reconstruction constraints. In addition, the discriminant criteria for training neural networks were also explored to further improve the separation performance. 2. Monaural speech separation based on deep recurrent neural networks using recurrent temporal restricted Boltzmann machines: In 2017, SAMUIS, CHAKRABARTI I, GHOSH S K. Deep recurrent neural network based monaural speech separation using recurrent temporal restricted boltzmann machines [C]. Stockholm: INTERSPEECH, 2017: 3622-3626. proposed a single-channel speech separation framework using recurrent temporal restricted Boltzmann machines (RTRBM), which jointly models all sources in the mixed signal as the target of a deep recurrent neural network (DRNN). The proposed method outperforms the method based on non-negative matrix factorization (NMF) and traditional speech enhancement methods based on DNN and DRNN.3. Time-Frequency Mask-Aware Bidirectional LSTM: A Deep Learning Approach for Underwater Acoustic Signal Separation. In 2022, Chen, J.; Liu, C.; Xie, J.; An, J.; Huang, N. Sensors 2022, 22, 5598. This method uses the bidirectional long short-term memory (Bi-LSTM) method to explore the features of the time-frequency (T-F) mask and proposes a T-F mask-aware Bi-LSTM for signal separation. Utilizing the sparsity of the T-F image, the designed Bi-LSTM network can extract discriminative features for separation, further improving the separation performance. 4. Underwater Acoustic Nonlinear Blind Ship Noise Separation Using Recurrent Attention Neural Networks. In 2024, Ruiping Song, Xiao Feng, Junfeng Wang, Haixin Sun, Mingzhang Zhou, Hamada Esmaiel. Remote Sensing (IF 4.2) Pub Date: 2024-02-09, DOI: 10.3390 / rs16040653. This method considers the softening effect of underwater bubbles in the separation of underwater nonlinear blind ship radiated noise, combines the advantages of RNN and Transformer, and proposes an end-to-end recurrent attention neural network, improving the separation accuracy. Summary of the Invention

[0005] To solve the technical problems in the above background, the present invention introduces a sub-band strategy and a dual-path parallel gating network, which significantly improves the effect and efficiency of underwater blind source separation, while reducing the computational complexity and enhancing the robustness of the system.

[0006] To achieve the above object, the present invention provides a method for blind source separation of underwater acoustic signals based on a dual-path parallel gating network, characterized in that the steps include:

[0007] Collect a dataset of original underwater acoustic signals and convert it into a spectrogram;

[0008] Use the spectrogram to train the BSM-DPGN model;

[0009] Input the underwater acoustic signal to be separated into the BSM-DPGN model to complete the blind source separation of the underwater acoustic signal.

[0010] Preferably, the method for obtaining the spectrogram includes:

[0011] Perform a non-linear transformation on the original underwater acoustic signal:

[0012] Convert the non-linearly transformed underwater acoustic signal into a spectrogram.

[0013] Preferably, the BSM-DPGN model includes: a band splitting module, a dual-path parallel gated network, and a mask generation module;

[0014] The band splitting module is used to split the obtained spectrogram into several sub-band spectrograms according to different frequency bands and then input them into the dual-path parallel gated network;

[0015] The dual-path parallel gated network is used to capture the feature dependencies between the signal sub-bands in the spectrogram and mine the deep feature correlations;

[0016] The mask generation module generates masks for each sub-band, combines all the masks together and multiplies them with the spectrogram to obtain the final separation result.

[0017] Preferably, the step of performing a non-linear transformation on the original underwater acoustic signal includes: using a non-linear model based on the softening effect of underwater bubbles:

[0018]

[0019]

[0020] where p(x,t) represents the sound pressure; v(x,t) represents the change in the volume of the bubble; V(x,t) represents the current bubble volume; c 0l and ρ 0l respectively represent the sound speed and the medium density in seawater; N g represents the number of bubbles per unit volume; δ represents the viscous damping coefficient in seawater; ω 0g represents the resonance angular frequency of the bubble; a, b, η represent non-linear coefficients;

[0021] The step of converting the non-linearly transformed underwater acoustic signal into a spectrogram includes: performing frame splitting, windowing, and short-time Fourier transform on the non-linearly transformed underwater acoustic signal;

[0022] Among them, the frame splitting operation includes:

[0023] Divide the original signal into multiple overlapping frames. Assuming the original signal is y(t), the signal after frame splitting is {y 1 ,y2 , … y K}, where the frame signal y k (t) is expressed as:

[0024] y k (t) = y(t + kM), 0 ≤ t < N, k = 1, 2,..., Z

[0025] where Z represents the total number of frames; N represents the length of each frame; M represents the frame shift;

[0026] The windowing operation includes:

[0027] x k (t) = y k (t) · w(t)

[0028] where x k (t) represents the windowed signal; w(t) is the window function;

[0029] The Fourier operation includes:

[0030]

[0031] where X k (f) represents the frame signal spectrogram; F{·} represents the Fourier transform, and f is the discrete frequency index.

[0032] Preferably, the sub - band module defines a high - resolution division for the low - frequency part and a low - resolution division for the high - frequency part according to the frequency - band distribution characteristics of the underwater acoustic signal; the input mixed - signal frequency band is segmented into K sub - band spectra with predefined bandwidths where G represents the frequency range of the i - th sub - band; i

[0033] Using the sub - band module, high - resolution sub - bands are obtained; the sub - band spectra are respectively passed to their corresponding layer normalization modules and fully - connected layers, and finally are transformed into fixed - dimension features after standardization. The process is expressed by the formula:

[0034] Z i = FC(norm(B i ))

[0035] where norm(·) represents layer normalization; FC(·) represents the fully - connected layer; Z i represents the sub - band feature.

[0036] Preferably, the dual - path parallel gating network includes:

[0037] H = HIE(Padding(X))​

[0038] G = σ(W g [X, H] + b g )

[0039]

[0040] Among them, padding(·) represents the operation of padding the front end of the processed signal with a zero vector of size 0 along the length dimension; HIE(·) represents a linear layer with a weight matrix, which aggregates all relevant historical information of each time step in parallel by sliding along the sequence length dimension; G and H represent intermediate variables involved in the gating mechanism; represents the element-wise product; σ(·) and tanh(·) represent the sigmoid and tanh activation functions; out represents the output of the PGN.

[0041] Preferably, the BSM-DPGN model is trained and optimized using the Adam algorithm:

[0042]

[0043] Among them, represents the first moment estimate with bias correction; represents the second moment estimate with bias correction; ε represents preventing division by zero errors during implementation; ζ represents the learning rate.

[0044] The present invention also provides an underwater acoustic signal blind source separation system based on a dual-path parallel gating network. The system is used to implement the above method and includes: an acquisition module, a construction module, and a separation module;

[0045] The acquisition module is used to acquire a dataset of the original underwater acoustic signal and convert it into a spectrogram;

[0046] The construction module is used to train the BSM-DPGN model using the spectrogram;

[0047] The separation module is used to input the underwater acoustic signal to be separated into the BSM-DPGN model to complete the blind source separation of the underwater acoustic signal.

[0048] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0049] (1) The present invention adopts an underwater acoustic signal blind source separation method based on a dual-path parallel gating network, combines the subband strategy, the parallel gating network (PGN), and the bidirectional long short-term memory network (BLSTM), and effectively captures the feature dependencies in the time series and frequency band dimensions through dual-path modeling, significantly improving the accuracy and efficiency of underwater blind source separation;

[0050] (2) The banding strategy proposed by the present invention can reasonably divide the frequency bands of the low-frequency part and the high-frequency part according to the frequency band distribution characteristics of the underwater acoustic signals. By independently modeling each sub-band signal, the process of extracting spectral features is optimized, the modeling ability of the dependence relationship between different frequency bands is enhanced, thereby improving the robustness of the model in complex underwater environments;

[0051] (3) The present invention uses a parallel gating network (PGN) for time series modeling, optimizing the learning process of temporal features. Through PGN, the model achieves a faster calculation speed while maintaining the same theoretical complexity, and effectively enhances the modeling ability of time series information. Especially in the underwater blind source separation task, the model's ability to capture time dependencies is improved;

[0052] (4) The present invention applies a bidirectional long short-term memory network (BLSTM) for modeling in the frequency band dimension, effectively capturing the long-term dependence relationship between frequency bands, and enhancing the model's expressive ability through residual connections, thereby improving the separation effect, reducing information loss during the separation process, and significantly enhancing the separation quality;

[0053] (5) The present invention generates an accurate time-frequency mask through a time-frequency mask generation module based on a multi-layer perceptron (MLP), and realizes signal reconstruction through element-wise multiplication. This method can efficiently separate the target signal from noise, reduce the spectral reconstruction error existing in traditional blind source separation methods, and further improve the accuracy of the underwater blind source separation task.

[0054] In summary, through these innovative points, the present invention effectively solves problems such as multi-band feature extraction, temporal dependence modeling, and signal reconstruction in underwater blind source separation, improves the application performance of the model in actual underwater environments, and provides a new technical solution for the field of underwater acoustic signal processing. A method for blind source separation of underwater acoustic signals based on a dual-path parallel gating network proposed by the present invention can effectively complete the ship noise separation task and significantly improve the separation accuracy and robustness. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] In order to more clearly illustrate the technical solutions of the present invention, the accompanying drawings required for use in the embodiments are briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0056] Figure 1 It is a schematic diagram of the working process of the BSM-DPGN model according to an embodiment of the present invention;

[0057] Figure 2 It is a structural diagram of a banded underwater acoustic signal blind source separation model according to an embodiment of the present invention;

[0058] Figure 3 It is a schematic diagram of the parallel gating network according to an embodiment of the present invention;

[0059] Figure 4 It is a structural diagram of the BSM-DPGN model according to an embodiment of the present invention;

[0060] Figure 5 It is a waveform example diagram of the separated signal according to an embodiment of the present invention; among them, (a) represents the target waveform diagram of a passenger ship; (b) represents the target waveform diagram of a motorboat; (c) represents the waveform diagram of the aliased signal; (d) represents the target waveform diagram of the passenger ship after separation;

[0061] Figure 6 It is a comparative curve diagram of the separation effects of the BSM-DPGN model according to an embodiment of the present invention before and after improvement on the ShipsEar dataset; among them, (a) represents the target waveform diagram of a passenger ship; (b) represents the target waveform diagram of a motorboat; (c) represents the waveform diagram of the aliased signal; (d) represents the target waveform diagram of the passenger ship after separation by different models;

[0062] Figure 7 It is a comparative curve diagram of the separation effects of the BSM-DPGN model according to an embodiment of the present invention under different sub-band strategies; among them, (a) represents the target waveform diagram of a passenger ship; (b) represents the target waveform diagram of a motorboat; (c) represents the waveform diagram of the aliased signal; (d) represents the target waveform diagram of the passenger ship after separation by different models. Detailed implementation manners

[0063] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0064] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0065] Embodiment 1

[0066] This embodiment provides an underwater acoustic signal blind source separation method based on a dual-path parallel gating network. The steps include:

[0067] S1. Collect a dataset of original underwater acoustic signals and convert it into a spectrogram.

[0068] Obtain a dataset of original underwater acoustic signals, perform a non-linear transformation on the original underwater acoustic signals, and convert the non-linearly transformed underwater acoustic signals into spectrograms.

[0069] S101. Perform a non - linear transformation on the original underwater acoustic signal.

[0070] To simulate the non - linear effect during the underwater propagation of the underwater acoustic signal, a non - linear model based on the softening effect of underwater bubbles is adopted. This model is derived from the wave equation and the Rayleigh - Plesset equation.

[0071]

[0072] Among them, p(x, t) represents the sound pressure, which varies with spatial coordinates and time; v(x, t) represents the change in the bubble volume, v(x, t)=V(x, t)-v 0g , V(x, t) represents the current bubble volume, and v 0g represents the initial bubble volume, and the calculation formula is expressed as where R 0g represents the initial radius of the bubble; c 0l and ρ 0l respectively represent the sound speed and the medium density in seawater; N g represents the number of bubbles per unit volume; δ represents the viscous damping coefficient in seawater; ω 0g represents the resonance angular frequency of the bubble; b = 1 / (6v 0g ) and η = 4πR 0g / ρ 0l represent non - linear coefficients, where γ g represents the specific heat ratio of the gas.

[0073] Let i and j represent the spatial and temporal indices respectively. The first - order and second - order partial derivatives in time and space can be converted into discrete formats, which are expressed as follows:

[0074]

[0075] Among them, p i represents p(x, t = i), and similarly, p j represents p(x, t = j); x represents the spatial coordinate; t represents the time coordinate.

[0076] In the time domain:

[0077]

[0078] Among them, h and τ respectively represent the spatial step size and the time step size; v represents the change value of the bubble volume.

[0079] Finally, the sound pressure at any spatial coordinate and time can be calculated by the following formula:

[0080]

[0081] The bubble volume change at the same point satisfies:

[0082]

[0083] S102. Convert the non-linearly transformed underwater acoustic signal into a spectrogram.

[0084] After performing frame segmentation, windowing, and short-time Fourier transform on the non-linearly transformed underwater acoustic signal, a spectrogram is obtained.

[0085] Among them, the frame segmentation operation includes:

[0086] Divide the original signal into multiple overlapping frames. Assuming the original signal is y(t), the framed signal is {y 1 , y 2 , … y K}, where the frame signal y k (t) is expressed as:

[0087] y k (t) = y(t + kM), 0 ≤ t < N, k = 1, 2, …, Z

[0088] Among them, Z represents the total number of frames; N represents the length of each frame; M represents the frame shift;

[0089] The framed signal after frame segmentation prevents signal information loss through windowing. The windowing process is expressed as:

[0090] x k (t) = y k (t) · w(t)

[0091] Among them, x k (t) represents the windowed signal; w(t) is the window function;

[0092] Perform short-time Fourier transform (STFT) on each windowed frame signal to obtain the frame signal spectrogram X k (f), that is:

[0093]

[0094] Among them, X k (f) represents the frame signal spectrogram; F{·} represents the Fourier transform, and f is the discrete frequency index.

[0095] S2. Use the spectrogram to train the BSM-DPGN model.

[0096] The BSM-DPGN model is trained using a dataset containing the original underwater acoustic signal spectrogram. The constructed BSM-DPGN model includes: a band splitting module, a dual-path parallel gating network, and a mask generation module.

[0097] Specifically, the full-band spectrum is split into K sub-band spectrums by the band splitting module according to the different bands, and then input into the dual-path parallel gating network; the process is as follows:

[0098] According to the frequency band distribution characteristics of the underwater acoustic signal, the high-resolution division of the low-frequency part and the low-resolution division of the high-frequency part are defined; the input mixed signal frequency band Split into K segments with predefined bandwidth The subband spectrum in G i represents the frequency range of the i-th sub-band;

[0099] The band splitting module is used to obtain high-resolution sub-bands; the sub-band spectrum is passed to the corresponding layer normalization module and the fully connected layer respectively, and finally converted into fixed-dimensional features after standardization. The process is expressed by the formula:

[0100] Z i =FC(norm(B i ))

[0101] Where norm(·) represents layer normalization; FC(·) represents the fully connected layer; Z i Represents subband characteristics.

[0102] After that, the dual-path parallel gating network first divides the signal into time steps according to the time dimension for parallel processing, and uses the historical information extraction layer (HIE) to extract information from all time steps in parallel mode through linear operations. Then, a single gating is used to simultaneously control the selection and fusion of information, avoiding complex gating operations and reducing computational overhead. Next, the frequency band RNN is used to capture the inter-band feature dependencies of the K sub-bands of each frame between frequency bands. Specifically, the layer normalization module is applied to the input of the module, and then the BLSTM layer and the FC layer are applied to perform the actual modeling. A residual connection is added between the input and output of the FC layer. The dual-path parallel gating network is as follows:

[0103] The parallel gating network can be expressed as:

[0104] H=HIE(Padding(X))

[0105] G=σ(W g [X,H]+b g )

[0106]

[0107] Among them, padding(·) represents the operation of padding the front end of the processed signal with a zero vector of size 0 along the length dimension; HIE(·) represents a linear layer with a weight matrix, which aggregates all relevant historical information of each time step in parallel by sliding along the sequence length dimension; G and H represent intermediate variables involved in the gating mechanism; represents the element-wise product; σ(·) and tanh(·) represent the sigmoid and tanh activation functions; out represents the output of the PGN.

[0108] Band dimension modeling is used to capture the inter-band feature dependencies of K sub-bands in each frame. First, the layer normalization module is applied to the input of the module, and then the BLSTM layer and the FC layer are applied to perform the actual modeling. A residual connection is added between the input and output of the FC layer. Multiple such RNNs can be stacked to create a deeper architecture.

[0109] The band RNN can be expressed by the formula:

[0110]

[0111] Among them, Norm(·) represents layer normalization; BiLSTM(·) represents a bidirectional long short-term memory network; FC(·) represents a fully connected layer.

[0112] Finally, the mask generation module uses a multi-layer perceptron (MLP) to generate masks for each sub-band, and after merging all the masks together, multiplies them with the spectrogram to obtain the final separation result;

[0113] The mask generation module calculates the mean absolute error loss in the frequency domain and the mean absolute error loss in the time domain between the clean target and the separated target. Among them, the loss in the frequency domain calculates the differences between the real and imaginary parts of the source signal and the estimated signal, while the loss in the time domain calculates the differences between the time-domain waveforms of the source signal and the estimated signal after the inverse short-time Fourier transform (iSTFT). The total loss function L obj is as follows:

[0114]

[0115] In the formula, represents the complex-valued spectrogram of the clean target, and the subscripts l and i are the real and imaginary parts respectively; iSTFT is the inverse operator of STFT.

[0116] The BSM-DPGN model of this embodiment is trained and optimized using the Adam algorithm:

[0117]

[0118] Among them, Represents the first moment estimation of bias correction; Represents the second moment estimation of bias correction; ε represents preventing division by zero error during implementation; ζ represents the learning rate.

[0119] S3. Input the underwater acoustic signal to be separated into the BSM-DPGN model to complete the blind source separation of the underwater acoustic signal.

[0120] The working process of the BSM-DPGN model in this embodiment is as Figure 1 shown.

[0121] Embodiment 2

[0122] Next, in combination with this embodiment, it will be described in detail how the BSM-DPGN model of the present invention is constructed.

[0123] First, construct a basic blind source separation model for underwater acoustic signals. Utilize the time series processing ability of the recurrent neural network and the gradient optimization advantage of residual connection, divide it into multiple stacked recurrent units for feature extraction and transmission, and achieve information fusion between the recurrent units of each layer through residual connection to construct the baseline blind source separation model for underwater acoustic signals. Add a subband module to the above-mentioned blind source separation model for underwater acoustic signals, divide a spectrogram into a series of subband spectrograms with a set of predefined bandwidths to improve the frequency resolution. Use each subband spectrogram as the input of the baseline blind source separation model for underwater acoustic signals, and finally fuse the outputs to achieve underwater blind source separation. Figure 2 It is the structural diagram of the subband blind source separation model for underwater acoustic signals proposed by the present invention.

[0124] After that, based on the subband blind source separation model for underwater acoustic signals, construct a blind source separation model for underwater acoustic signals based on a dual-path parallel gating network (i.e., the proposed BSM-DPGN model).

[0125] In order to improve the performance of underwater blind source separation and solve the problems of high signal complexity and severe noise interference in the underwater environment, a signal processing strategy based on a parallel gating network (PGN) is proposed. Through the parallel computing advantage of PGN, features can be processed and updated simultaneously in multiple audio channels, thus accelerating the signal processing process and avoiding the computational bottleneck in traditional methods. Through PGN, the features in each audio channel are selectively updated to dynamically adjust the signal transmission ratio to extract richer feature information. PGN controls the intensity of information flow through a gating mechanism, selectively transmits, updates, or suppresses the key information in the signal, thereby effectively reducing noise interference, enhancing the effectiveness of the signal, improving the separation effect and the robustness of the model. This strategy can better capture the correlation between different sound sources and improve the accuracy and generalization ability of underwater blind source separation. Figure 3 It is the schematic diagram of the parallel gating network. Figure 4It is the structural diagram of the underwater acoustic signal blind source separation model based on the dual-path parallel gating network proposed by the present invention.

[0126] The sub-band data is processed by the parallel gating network and the band RNN, and finally all sub-band masks are generated and merged by the mask generation module, and finally the target signal spectrogram is generated. Figure 5 It is the waveform example diagram of the separated signal of the present invention.

[0127] Loss function. The target task of this embodiment is to separate the mixed audio, and the mean absolute error (MAE) loss function can be used to optimize the model. Specifically, the loss function consists of the MAE loss in the frequency domain and the MAE loss in the time domain. The loss in the frequency domain calculates the difference between the source signal and the estimated signal in the real and imaginary parts, while the loss in the time domain calculates the difference between the source signal and the estimated signal in the time domain waveform after the inverse short-time Fourier transform (iSTFT). The total loss function is as follows:

[0128]

[0129] In the formula, represents the complex-valued spectrogram of the clean target, and the subscripts l and i are the real and imaginary parts respectively; iSTFT is the inverse operator of STFT.

[0130] Embodiment III

[0131] To verify the advancement of the present invention, this embodiment is specifically set as a control for illustration.

[0132] The underwater public dataset is used as the input of the underwater acoustic signal blind source separation model based on the dual-path parallel gating network, and a comparative experiment is carried out to evaluate the separation effect of the model. At the same time, an ablation experiment is carried out to verify the effectiveness of the module, and a research on an underwater acoustic signal blind source separation method based on the dual-path parallel gating network is completed.

[0133] Specifically, in the experimental stage of this embodiment, the underwater blind source separation task is completed. The parallel gating network and the band RNN of the model will be used as the feature encoding part of this model. After the original underwater audio data is processed by the short-time Fourier transform and the sub-band module, the sub-band spectrum is processed by the parallel gating network, the band RNN and the mask generation module to generate the underwater blind source separation result.

[0134] To verify the separation effect of the present invention in the underwater blind source separation task, the signal-to-distortion ratio (SDR) is used as the evaluation index in the experiment. SDR is used to measure the signal-to-noise ratio between the separated signal and the original signal. It comprehensively considers the difference between the source signal and the distorted signal and can effectively evaluate the performance of blind source separation. The calculation method of SDR is:

[0135]

[0136] Among them, s is the original signal; y is the separated signal; the lengths of both s and y are L.

[0137] The higher the SDR, the better the separation effect, the more accurate the signal recovery, and the smaller the distortion. SDR is one of the standard indicators for evaluating the quality of blind source separation. It can reflect the performance of the model in actual signal recovery, especially suitable for underwater signal separation tasks, and can comprehensively evaluate the separation ability and robustness of the model.

[0138] In this embodiment, the ShipsEar dataset is used for model learning and performance verification. This database consists of 90 recordings, with a sampling frequency of 52734 Hz and a duration ranging from 15 seconds to 15 minutes. These recordings include the sounds of 11 types of ship engines and natural environmental noises, which can be combined into 5 categories. The present invention selects two categories, passenger ships and motorboats, from the dataset, selects the underwater acoustic source signals considered to be clean among them, cuts them into 5 s and downsamples them to 22050 Hz, obtaining 686 passenger ship audio and 159 motorboat audio data. Then, for each passenger ship audio, a random motorboat audio is selected, the powers of the two audio data are calculated, and the motorboat audio is scaled to have the same power as the passenger ship audio and mixed at a signal-to-noise ratio of 0 dB to obtain the 2-mix dataset. The proportion of training data, validation data, and test data in the 2-mix dataset is allocated as 7:1:2. The initial learning rate is set to 1e-3 and decays by 0.98 every two epochs. Stop when no best validation is found in 10 consecutive epochs. To prove the performance gain of underwater blind source separation brought by the improved part of the present invention and at the same time verify the robustness of the model, this experiment will conduct ablation experiments on the improved performance of the model on the ShipsEar dataset. Table 1 shows the SDR comparison results before and after improvement on the original unimproved blind source separation model when training on the ShipsEar dataset. In the description of this experiment and the following figures, BSM (Band Split Module) refers to the blind source separation model for underwater acoustic signals with band splitting, DPGN (Dual-path Parallel Gated Network) represents adding a dual-path parallel gated network to the baseline model, and BSM-DPGN refers to the blind source separation model for underwater acoustic signals based on the dual-path parallel gated network.

[0139] Table 1

[0140]

[0141] As can be seen from Table 1, on the ShipsEar dataset, compared with the model before improvement, the signal-to-noise ratio of the BSM model has increased from 4.721 to 5.015, an increase of 0.294, that is, the blind source separation performance index has achieved a relatively significant improvement. After adding the parallel gating network, compared with the model before improvement, the signal-to-noise ratio of the BSM-DPGN model has increased from 4.721 to 5.408, an increase of 0.687. Based on the performance of the BSM-DPGN model in the blind source separation task on the ShipsEar dataset, it is proved that the sub-band module and the dual-channel parallel gating network of the present invention are effective and robust in improving the blind source separation task, that is, the blind source separation method of underwater acoustic signals based on the dual-path parallel gating network has better blind source separation performance. Figure 6 It is a comparison curve graph of the separation effects of the BSM-DPGN model before and after improvement on the ShipsEar dataset, where Figure 6 (a) is the audio waveform graph of the passenger ship target, Figure 6 (b) is the audio waveform graph of the motorboat target, Figure 6 (c) is the aliased audio waveform graph, Figure 6 (d) is the audio waveform graph of the passenger ship after separation by different models.

[0142] In order to verify the influence of different sub-band strategies on the separation performance of the model, in this embodiment, the separation performance of the model brought by different sub-band strategies will be verified on the ShipsEar dataset. Table 2 gives the comparison results of the model separation performance on the ShipsEar dataset under different sub-band strategies. Sub-band strategy V1 divides the entire spectrogram evenly with a bandwidth of 5 kHz (the remaining part is merged into the last sub-band). Sub-band strategy V2 divides the entire spectrogram evenly with a bandwidth of 3 kHz (the remaining part is merged into the last sub-band). Sub-band strategy V3 divides the entire spectrogram evenly with a bandwidth of 1 kHz (the remaining part is merged into the last sub-band). Sub-band strategy V4 divides the frequency band below 4 kHz into bandwidths of 250 Hz, divides the frequency band between 4 kHz and 8 kHz into bandwidths of 500 Hz, divides the frequency band between 8 kHz and 11 kHz into bandwidths of 1 kHz, and regards the rest as one sub-band. Sub-band strategy V5 divides the frequency band below 1 kHz into bandwidths of 100 Hz, divides the frequency band between 1 kHz and 4 kHz into bandwidths of 250 Hz, divides the frequency band between 4 kHz and 8 kHz into bandwidths of 500 Hz, divides the frequency band between 8 kHz and 11 kHz into bandwidths of 1 kHz, and regards the rest as one sub-band. Figure 7 It is a comparison curve graph of the separation effects under different sub-band strategies, where Figure 7 (a) is the audio waveform graph of the passenger ship target, Figure 7 (b) is the audio waveform graph of the motorboat target, Figure 7 (c) is the aliased audio waveform graph,Figure 7 (d) is the audio waveform diagram of the passenger ship after different models are separated. Combining Table 2 and Figure 5 it can be seen that for the ShipsEar dataset, the separation performance of the model reaches the best when the banding strategy is V5. From the data analysis in Table 2, it can be known that as the banding is refined, the signal-to-noise ratio gradually rises from 4.770 to 5.408, indicating that the fine-grained segmentation scheme can improve the performance of BSM-DPGN.

[0143] Table 2

[0144]

[0145] To prove the performance of the underwater acoustic signal blind source separation model based on the parallel gating network proposed in the present invention in a noisy environment, different intensities of Gaussian white noise are added to the mixed audio data in this embodiment. This embodiment will conduct experimental comparisons on different noise intensities on the ShipsEar dataset. Table 3 shows the comparison experimental results of different noise intensities in the present invention. From the data analysis in Table 3, it can be known that when the noise intensity is 0 dB, the SDR can reach 4.366, indicating that the BSM-DPGN model also has good performance in the case of high noise interference. As the noise intensity decreases, at a noise intensity of 15 dB, the SDR reaches 5.379, only differing from the SDR without noise by 0.029. The overall data of this embodiment shows that the underwater acoustic signal blind source separation method based on the dual-path parallel gating network proposed in the present invention has good noise resistance.

[0146] Table 3

[0147]

[0148] To prove that the BSM-DPGN model proposed in the present invention has better segmentation performance compared with other object segmentation methods, this embodiment conducts a comparative experiment on the BSM-DPGN model and other advanced method models on the ShipsEar dataset. Table 4 gives the overall performance comparison results of the BSM-DPGN model and the advanced models.

[0149] Table 4

[0150] Model Signal-to-Noise Ratio (SDR) BSM-DPGN 5.408 RANN 4.684 IAM-BiLSTM 4.360

[0151] Observing the data in Table 4, it can be known that the model in this embodiment shows competitive performance on the ShipsEar dataset. Generally speaking, the signal-to-noise ratio (SDR) trained by the model in this embodiment on the ShipsEar data is relatively good, reaching 5.408. This indicates that the model in this embodiment can separate the mixed audio more accurately.

[0152] In summary, the underwater acoustic signal blind source separation method based on the dual-path parallel gating network proposed by the present invention can effectively complete the underwater blind source separation task, and has good segmentation accuracy and robustness.

[0153] Embodiment 4

[0154] This embodiment also provides an underwater acoustic signal blind source separation system based on the dual-path parallel gating network, including: an acquisition module, a construction module, and a separation module; the acquisition module is used to acquire a dataset of the original underwater acoustic signal and convert it into a spectrogram; the construction module is used to train the BSM-DPGN model by using the spectrogram; the separation module is used to input the underwater acoustic signal to be separated into the BSM-DPGN model to complete the blind source separation of the underwater acoustic signal.

[0155] The above-described embodiments are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.

Claims

1. A method for blind source separation of underwater acoustic signals based on a dual-path parallel gating network, characterized in that the steps include: Collecting data sets of raw underwater acoustic signals and converting them into spectrograms; Using the spectrogram, training a BSM-DPGN model; The underwater acoustic signal to be separated is input into the BSM-DPGN model to complete the blind source separation of the underwater acoustic signal.

2. The method for blind source separation of underwater acoustic signals based on a dual-path parallel gating network according to claim 1 is characterized in that: The method for obtaining the spectrum diagram includes: Perform nonlinear transformation on the original underwater acoustic signal: Convert the underwater acoustic signal after nonlinear transformation into a spectrogram.

3. The method for blind source separation of underwater acoustic signals based on a dual-path parallel gating network according to claim 1 is characterized in that: The BSM-DPGN model includes: a band division module, a dual-path parallel gating network and a mask generation module; The band splitting module is used to split the acquired spectrum into a plurality of sub-band spectrums according to different bands, and then input the sub-band spectrums into the dual-path parallel gating network; The dual-path parallel gating network is used to capture the feature dependencies between signal sub-bands in the spectrum graph and mine deep feature associations; The mask generation module generates a mask for each sub-band, combines all masks together and multiplies them with the spectrum map to obtain the final separation result.

4. The method for blind source separation of underwater acoustic signals based on a dual-path parallel gating network according to claim 2 is characterized in that: The steps of performing nonlinear transformation on the original underwater acoustic signal include: using a nonlinear model based on the softening effect of bubbles in water: Where p(x, t) represents the sound pressure; v(x, t) represents the change in bubble volume; V(x, t) represents the current bubble volume; c 0l and ρ 0l Respectively represent the speed of sound and medium density in seawater; N g represents the number of bubbles per unit volume; δ represents the viscous damping coefficient in seawater; ω 0g represents the resonant angular frequency of the bubble; a, b, η represent the nonlinear coefficients; The step of converting the underwater acoustic signal after nonlinear transformation into a spectrum diagram comprises: performing framing, windowing, and short-time Fourier transformation on the underwater acoustic signal after nonlinear transformation; The framing operations include: The original signal is divided into multiple overlapping frames. The original signal is set to y(t). Then the signal after framing is {y1, y2, …y K }, where the frame signal y k (t) is expressed as: y k (t)=y(t+kM),0≤t<N,k=1,2,...,Z Where Z represents the total number of frames; N represents the length of each frame; M represents the frame shift; The windowing operation includes: x k (t)=y k (t)·w(t) Among them, x k (t) represents the windowed signal; w(t) is the window function; Fourier operations include: Among them, X k (f) represents the frame signal spectrum diagram; F{·} represents Fourier transform, and f is the discrete frequency index.

5. The method for blind source separation of underwater acoustic signals based on a dual-path parallel gating network according to claim 3 is characterized in that: The band division module defines the high-resolution division of the low-frequency part and the low-resolution division of the high-frequency part according to the frequency band distribution characteristics of the underwater acoustic signal; Split into K segments with predefined bandwidth The subband spectrum in G i represents the frequency range of the i-th sub-band; The band splitting module is used to obtain high-resolution sub-bands; the sub-band spectra are respectively passed to the corresponding layer normalization module and the fully connected layer, and finally converted into fixed-dimensional features after standardization. The process is expressed by the formula: From i =FC(norm(B i )) Where norm(·) represents layer normalization; FC(·) represents the fully connected layer; Z i Represents subband characteristics.

6. The method for blind source separation of underwater acoustic signals based on a dual-path parallel gating network according to claim 3 is characterized in that: The dual-path parallel gating network comprises: H=HIE(Padding(X)) G=σ(W g [X,H]+b g ) Where padding(·) represents the operation of padding the front end of the signal processing along the length dimension with a zero-filled vector of size 0; HIE(·) represents a linear layer with a weight matrix that aggregates all relevant historical information of each time step in parallel by sliding along the sequence length dimension; G and H represent the intermediate variables involved in the gating mechanism; represents element-wise product; σ(·) and tanh(·) represent sigmoid and tanh activation functions; out represents the output of PGN.

7. The method for blind source separation of underwater acoustic signals based on a dual-path parallel gating network according to claim 3 is characterized in that: The BSM-DPGN model uses the Adam algorithm for training optimization: in, represents the bias-corrected first-moment estimate; represents the bias-corrected second-order moment estimate; ε represents the prevention of division by zero errors in the implementation process; ζ represents the learning rate.

8. A blind source separation system for underwater acoustic signals based on a dual-path parallel gating network, the system being used to implement the method according to any one of claims 1 to 7, characterized in that: include: Acquisition module, construction module and separation module; The acquisition module is used to collect data sets of original underwater acoustic signals and convert them into spectrograms; The building module is used to train the BSM-DPGN model using the spectrogram; The separation module is used to input the underwater acoustic signal to be separated into the BSM-DPGN model to complete the blind source separation of the underwater acoustic signal.