Underwater acoustic signal recognition method based on cross-attention feature fusion
By using cross attention feature fusion and TF-transformer extraction of time-frequency features in water acoustic signal recognition technology, the problem of signal distortion in low signal-to-noise ratio environment is solved, and the recognition performance and robustness are significantly improved.
Patent Information
- Application Number
- CN202510248707.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-05-20
- Estimated Expiration
- 2045-03-04
AI Technical Summary
Existing water acoustic signal recognition technology can easily lead to signal distortion and loss of features in low signal-to-noise ratio environments, affecting the recognition performance.
The water acoustic signal recognition method based on cross attention feature fusion is adopted. By simultaneously processing amplitude and phase information in the noise reduction front end, and dynamically fusion amplitude and phase features are extracted in the feature fusion module, time-frequency features are extracted in the recognition back end in combination with TF-transformer.
It significantly improves the signal-to-noise ratio, reduces information loss, enhances the system's robustness and recognition performance, and improves the average recognition rate, accuracy, recall rate and F1 score.
Smart Images

Figure CN119758348B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of underwater acoustic signal processing, and particularly relates to an underwater acoustic signal recognition method based on cross-attention feature fusion. Background Art
[0002] The passive sonar system has the functions of remotely monitoring, positioning, tracking, and identifying targets such as ships in water. In particular, by receiving the radiated noise signals emitted by the target ship itself, the passive sonar can achieve target recognition. The underwater acoustic target recognition technology discriminates target individuals by analyzing the characteristics of underwater acoustic signals, which is one of the key technologies for realizing the intelligence of other underwater equipment and an important research direction in the field of underwater acoustic engineering. In the actual marine environment, when collecting underwater target signals, the hydrophone will inevitably be interfered by underwater environmental noises, such as the sounds of wind, rain, and marine organisms, resulting in a relatively low signal-to-noise ratio when the signal reaches the detection system. This poses a huge challenge to the fine analysis and processing of target signals, thereby causing a significant decline in the recognition performance of underwater acoustic signals.
[0003] In recent years, researchers have optimized the accuracy of underwater acoustic signal recognition from multiple perspectives, including improving the signal-to-noise ratio and optimizing the network model. A common strategy is to integrate an audio noise reduction front-end into the target recognition system, aiming to preprocess the low-signal-to-noise ratio audio signal to improve the signal-to-noise ratio and the quality of the target audio, thereby enhancing the accuracy of target recognition. The noise reduction front-end algorithms usually include Wiener filtering, spectral subtraction, and deep learning-based noise reduction algorithms, which process the signal in the time domain and time-frequency domain through different methods to reduce the influence of noise. However, the application of the noise reduction front-end does not always improve the recognition performance. In fact, the processed underwater acoustic audio signal after noise reduction may be distorted, which usually stems from over-suppressing the noise of the target signal or the artifacts introduced during the noise reduction process. Especially in a low-signal-to-noise ratio environment, the noise reduction algorithm often causes partial loss of the spectral characteristics of the target signal, thereby affecting the accuracy of the subsequent recognition system.
[0004] To reduce the distortion effect brought by the noise reduction model, researchers usually first preprocess the data of the noise reduction front-end and then let the backend recognition model directly process the noise-reduced audio. However, this method highly depends on the performance of the noise reduction model. To address the distortion problem caused by the noise reduction front-end, researchers have developed a joint training framework. The noise reduction front-end and the recognition backend are jointly trained to improve the robustness and accuracy of target recognition by optimizing the synergistic effect of feature extraction and recognition processes.
[0005] This method ensures that the front-end and back-end models can be jointly optimized, reducing error accumulation and improving the discriminative ability of features, thus providing better performance in a noisy environment. However, in the above method, only amplitude spectrum noise reduction and recognition are considered, without considering phase information, which affects the accuracy and robustness of recognition.
[0006] Therefore, by simultaneously considering amplitude and phase information in the noise reduction front-end, we can achieve precise signal reconstruction and optimized extraction of underwater acoustic target signal features, thereby better improving the signal-to-noise ratio. Self-attention and cross-attention can fully capture global and local information, facilitating the fusion of noise amplitude information with the amplitude and phase information output by the noise reduction front-end. Since the TF-transformer demonstrates excellent performance in speech recognition, we consider using it for feature extraction in back-end recognition to better capture audio features. Therefore, inspired by the above method, we propose a joint training framework based on the fusion of amplitude and phase features to jointly train the noise reduction and recognition tasks, aiming to improve the accuracy and robustness of underwater target recognition. Summary of the Invention
[0007] In view of this, the objective of the present invention is to provide an underwater acoustic signal recognition method based on cross-attention feature fusion, which can optimize the problem of underwater audio distortion caused by the noise reduction front-end.
[0008] An underwater acoustic signal recognition method based on cross-attention feature fusion includes:
[0009] Converting the underwater acoustic signal into spectrogram data D;
[0010] Inputting the spectrogram data D into the noise reduction front-end module; the noise reduction front-end module includes an encoder part, an information exchange part, and a decoder part; each part is divided into an amplitude branch and a phase branch;
[0011] The spectrogram data D is processed in the amplitude branch of the encoder part to obtain amplitude mask prediction data A1; the spectrogram data D is processed in the phase branch of the encoder part to obtain phase mask prediction data P1;
[0012] In the information exchange part, information sharing is performed on the amplitude mask prediction data A1 and the phase mask prediction data P1, and each branch receives the key information of the other party, thereby obtaining feature A4 corresponding to the amplitude branch and feature P4 in the phase branch;
[0013] In the decoder part, the amplitude branch decodes the feature A4 to obtain feature A7; the phase branch decodes the feature P4 to obtain feature P5;
[0014] The feature A7 and the feature P5 are sent to the feature fusion module based on the cross-attention mechanism for processing to obtain the fusion feature C5;
[0015] The fused feature C5 is sent to the recognition backend module for processing to obtain the category of the target.
[0016] Preferably, when converting the underwater acoustic signal into spectrogram data D, the underwater acoustic signal is segmented into multiple time periods; the segmented underwater acoustic signal is processed using Fourier transform to obtain the spectrogram. ; where the variable T represents the time step, F represents the number of frequency bands, that is, the number of frequency bands or frequency points divided along the frequency axis of the spectrogram, and each frequency band represents the energy or amplitude information of the signal within a specific frequency range; the variable 2 represents the real and imaginary parts of the signal.
[0017] Preferably, in the encoder, the amplitude branch consists of two convolutional layers and two activation functions. The spectrogram data D1 first passes through the first convolutional layer and the first ReLu activation function, and then passes through the second convolutional layer and the second ReLu activation function to obtain the feature A1; the phase branch consists of two convolutional layers, and the spectrogram data D2 passes through the two convolutional layers in sequence to obtain the feature P1.
[0018] Preferably, in the amplitude branch of the information exchange module, the feature A1 passes through three convolutional layers in sequence, and after each convolutional layer, there is a batch normalization layer BN and a ReLu activation function, and finally the feature A2 is obtained; in the phase branch, the feature P1 first undergoes linear normalization processing, and then passes through two convolutional layers in sequence to obtain the feature P2;
[0019] The feature A2 passes through a convolutional layer and a Tanh activation function to obtain the feature A3, and then is multiplied element-wise with the feature P2 to obtain the feature P4; after the feature P2 passes through a convolutional layer and a Tanh activation function to obtain the feature P3, it is then multiplied element-wise with the feature A2 to obtain the feature A4.
[0020] Preferably, in the decoder part, the feature A4 passes through a convolutional layer, a reshaping layer, and three fully connected layers FC to obtain the feature A5; the activation function connected after the convolutional layer is the Sigmoid activation function, the number of units in the first two fully connected layers FC is 600, the activation function connected after each fully connected layer FC is the ReLu activation function, the number of units in the last fully connected layer FC is the number of frequency bands, and the activation function is Sigmoid; the feature P4 passes through a convolutional layer and amplitude regularization to obtain the feature P5, completing the phase mask prediction; the feature A5 is multiplied element-wise with the input spectrogram data D to obtain the feature A6, and the feature A6 is then multiplied element-wise with the feature P5 to obtain the feature A7 to complete the amplitude mask prediction.
[0021] Preferably, the feature fusion module adopts a dual-branch structure, independently processing the amplitude and phase features respectively, specifically including:
[0022] Feature A7 and feature P5 are respectively corresponding to obtain feature A8 and feature P6 through Mel-spectrum conversion and self-attention mechanism in sequence. Feature A8 and feature P6 are respectively corresponding to obtain feature A9 and feature P7 through two fully connected layers FC. Feature A10 and feature P8 are respectively obtained by adding them to feature A8 and feature P6 in a skip connection manner. Feature A10 is input as the Q parameter of cross-attention, and feature P8 is input as the K parameter and V parameter of cross-attention. Feature C1 is obtained through cross-attention. Feature C2 is obtained by adding feature C1 to feature A10 in a skip connection manner. Then, feature C3 is obtained through a convolutional layer and the subsequent activation function ReLu. Feature C4 is obtained by successively passing through a fully connected layer FC and its activation function Sigmoid, a convolutional layer and the subsequent connected activation function ReLu. Feature C5 is obtained by adding feature C4 to feature C3 in a skip connection manner, completing feature fusion.
[0023] Preferably, in the recognition backend module, feature C5 first processes the input feature through two convolutional layers to obtain feature C6; feature C6 is further deeply processed through 3 feature extraction modules in sequence to obtain feature C7;
[0024] Among them, each of the feature extraction modules is composed of a TF-transformer module, three convolutional layers, and a TF-transformer module connected in sequence, and is used to extract the features of underwater acoustic signals. Among them, the three convolutional layers are used to extract the local correlation of features, with sizes of 5×5, 1×25, and 5×5 respectively, and a normalization layer, ReLu activation function, and Dropout layer are attached after each convolutional layer;
[0025] The TF-transformer is responsible for capturing the temporal and frequency dependencies in the signal; the TF-transformer module consists of two parts, the F-transformer module and the T-transformer module, which act on the frequency dimension and the time dimension of the signal respectively; in the F-transformer module, the feature C6 is processed by the multi-head attention mechanism to obtain the feature F1, the feature F2 is obtained by skip connection, the feature F3 is obtained through the fully connected layer FC, the feature F4 is obtained through the LSTM and the fully connected layer FC, the feature F5 is obtained by skip connection, the feature F6 is obtained through the fully connected layer FC, and the feature F7 is obtained by skip connection. The feature F7 passes through the T-transformer module to obtain the feature T7; the feature C7 aggregates the information of the feature map in the time dimension and the frequency dimension through average pooling to obtain the feature C8 and the feature C9 respectively. The feature C8 and the feature C9 are added to obtain the feature C10, which passes through the first fully connected layer FC, followed by the ReLu activation function processing, and finally passes through the second fully connected layer FC processing to obtain the final target category.
[0026] Preferably, it also includes the training of the noise reduction front-end module; among them, the loss function for training the noise reduction front-end module is the loss of amplitude prediction and the loss of phase prediction weighted sum, specifically:
[0027]
[0028]
[0029] Among them, represents the feature output by the amplitude branch of the decoder of the noise reduction front-end module, represents the modulo operation; represents the original spectrogram data, represents the modulo operation first and then the power-law compression spectrogram operation symbol represents the dot product operation.
[0030] Preferably, it also includes the training of the recognition backend module. Among them, the loss function for training the recognition backend module is as follows:
[0031]
[0032] Among them, is the number of samples, is the th sample th true value of the category, is the output of the recognition backend module for the The prediction probability of the th sample for each of the
[0033] Preferably, the noise reduction front-end module, the feature fusion module, and the recognition back-end module are jointly trained, and the loss function of the joint training is the weighted sum of the loss functions of the noise reduction front-end module and the recognition back-end module.
[0034] The present invention has the following beneficial effects:
[0035] The present invention provides an underwater acoustic signal recognition method based on cross-attention feature fusion. In the noise reduction module, the present invention uses a dual-branch structure to process amplitude and phase information respectively, and adopts an information interaction mode, enabling the amplitude or phase processing process to utilize the information of the other path as a reference, thereby improving the feature representation ability; the feature fusion module is used to dynamically fuse the amplitude and phase information obtained by the noise reduction module, solving the problem of only using amplitude information while ignoring phase features, enhancing the robustness of the system while significantly reducing information loss; in the recognition module, the present invention designs a feature extraction module, using TF-transformer to process the time and frequency dimensions respectively, and modeling the time-frequency distribution characteristics of speech by capturing the global dependence relationship between the two. Description of the Drawings
[0036] Figure 1 is the joint training framework based on cross-attention fusion;
[0037] Figure 2 is the noise reduction front-end module;
[0038] Figure 3 is the feature fusion module based on cross-attention;
[0039] Figure 4 is the recognition back-end module;
[0040] Figure 5 is the feature extraction structure;
[0041] Figure 6 is the TF-transform structure;
[0042] Figure 7 is the feature distribution before feature fusion;
[0043] Figure 8 is the feature distribution after feature fusion;
[0044] Figure 9 is the comparison of recognition rate, precision, recall rate, and F1 score among different models under different datasets. Detailed Embodiments
[0045] The present invention will be described in detail below with reference to the accompanying drawings and by way of examples.
[0046] As Figure 1 shown, the proposed joint training model consists of a noise reduction front end, a feature fusion module, and a recognition back end. First, the audio data of the underwater acoustic signal is converted into spectrogram data by STFT and input into the noise reduction front end module: , where T represents the time step. This means that when converting the sound signal into a spectrogram, the signal is divided into a certain number of time segments or frames. Each time step usually corresponds to a specific short time window within the signal. Within this window, the signal is processed using the Fourier transform (or a similar transform method) to analyze the frequency components within that time segment. The variable F represents the number of frequency bands, that is, the number of frequency bands or frequency points into which the spectrogram is divided along the frequency axis. Each frequency band represents the energy or amplitude information of the signal within a specific frequency range. The variable 2 represents the real and imaginary parts of the signal.
[0047] As Figure 2 shown, the noise reduction front end module includes an encoder, an information exchange module, and a decoder. Each part is divided into two branches, an amplitude branch (upper layer) and a phase branch (lower layer), which respectively complete the prediction of the amplitude mask and the phase mask; the spectrogram data is also divided into two branches, which are respectively defined as spectrogram data D1 and spectrogram data D2. In the encoder, the amplitude branch consists of two convolutional layers and two activation functions. The spectrogram data D1 first passes through the first convolutional layer Conv(7×1) and the ReLu activation function, and then passes through the second convolutional layer Conv(1×7) and the ReLu activation function to obtain the feature A1. The two convolutional layers do not change the length in the time domain and frequency domain, but only change the number of channels. The phase branch consists of two convolutional layers. The spectrogram data D2 sequentially passes through the convolutional layers Conv(5×3) and Conv(5×3) to obtain the feature P1. Similar to the amplitude encoding layer, the two convolutional layers do not change the length in the time domain and frequency domain, but only change the number of channels.
[0048] In the information exchange module, the amplitude branch first models the energy distribution of the underwater audio signal to help the model understand which parts are the main energy concentration areas of the signal, which is beneficial to restoring the high-frequency components in the low-frequency power. A1 passes through three convolutional layers in sequence, and each convolutional layer has BN (Batch Normalization) and the ReLu activation function to obtain A2. Using three convolutional layers can handle the correlation of frequency domain features. In the phase branch, P1 passes through two convolutional layers in sequence to obtain P2. Before the phase information is transmitted into the convolutional layer, Layer Normalization (LN) is first performed. Different from BN, LN normalizes all features of a single data sample. LN can reduce the internal covariate shift during the training of deep neural networks, thereby accelerating the training process and making the model training more stable. To promote information sharing between the amplitude branch and the phase branch, we add a Conv and the activation function Tanh. Specifically, A2 passes through Conv and the activation function Tanh to obtain A3, and then A3 is multiplied pointwise with P2 to obtain P4. P2 passes through Conv and the activation function Tanh to obtain P3, and then P3 is multiplied pointwise with A2 to obtain A4. Through this design, each branch can receive the key information of the other party, thereby obtaining a more comprehensive and detailed input feature representation. This complementary information exchange mechanism not only enhances the feature learning ability of the network but also improves the overall signal processing efficiency and accuracy.
[0049] In the decoding layer, A4 passes through a convolutional layer, the shaping layer Reshape, and three fully connected layers (FC) in sequence to obtain A5. The activation function of the convolutional layer is Sigmoid, the number of units in the first two FCs is 600, the activation function is ReLu, and the number of units in the last FC is F, and the activation function is Sigmoid. P4 passes through a convolutional layer and amplitude regularization to obtain P5, completing the phase mask prediction. A5 is multiplied pointwise with D to obtain A6, and A6 is then multiplied pointwise with P5 to obtain A7 to complete the amplitude mask prediction.
[0050] As Figure 4As shown in the figure, the present invention designs a feature fusion module. The module adopts a dual-branch structure to independently process amplitude and phase features respectively. Subsequently, feature fusion is achieved through a cross-attention mechanism. The cross-attention mechanism can extract complementary information from amplitude features and phase features, significantly enhancing key features. At the same time, the fused mel spectrogram can retain the inherent details of the signal and exhibit high feature cross-correlation and excellent non-linear feature extraction ability. Therefore, feature A7 and feature P5 are respectively transformed into feature A8 and feature P6 through mel spectrogram transformation and self-attention mechanism in sequence. Feature A8 and feature P6 are respectively transformed into feature A9 and feature P7 through two fully connected layers (FC). Feature A9 and feature P7 are added to feature A8 and feature P6 respectively through skip connection to obtain feature A10 and feature P8. We use feature A10 as the Q parameter input of cross-attention, feature P8 as the K and V parameter inputs of cross-attention. After cross-attention, we get C1, which is added to feature A10 through skip connection to obtain feature C2. Feature C2 passes through a convolutional layer and its activation function ReLu to obtain feature C3. Feature C3 passes through an FC and its activation function sigmoid, a convolutional layer and its activation function ReLu in sequence to obtain feature C4. Feature C4 is added to feature C3 through skip connection to obtain feature C5, completing feature fusion.
[0051] As Figure 5 shown, in the recognition backend, C5 first undergoes preliminary processing of the input features through two convolutional layers to obtain C6. The kernel size of each convolutional layer is 3*3, and both are attached with ReLu activation function and Dropout layer to enhance feature expression ability and prevent overfitting. After completing the preliminary processing, C6 is further deeply processed through 3 feature extraction modules to obtain C7. Each feature extraction module consists of a TF-transformer, three convolutional layers, and a TF-transformer, and is used to extract the features of underwater acoustic signals. Among them, the three convolutional layers can capture the local correlation of features, with sizes of 5*5, 1*25, and 5*5 respectively, and each convolutional layer is attached with a normalization layer, ReLu activation function, and Dropout layer; the TF-transformer is responsible for capturing complex temporal and frequency dependencies in the signal, while the convolutional layer is used to extract the local correlation of features. The structure of the TF-transformer is as Figure 6As shown, it includes two parts, the F-transformer and the T-transformer, which act on the frequency dimension and the time dimension respectively. In the F-transformer, C6 passes through the multi-head attention mechanism to obtain F1, obtains F2 through the skip connection method, obtains F3 through the FC, obtains F4 through the LSTM and the FC, obtains F5 through the skip connection method, obtains F6 through the FC, and obtains F7 through the skip connection method. Similarly, F7 passes through the T-transformer to obtain T7. C7 output by the feature extraction module passes through average pooling to aggregate the information of the feature map in the time dimension and the frequency dimension respectively to obtain C8 and C9, passes through the addition operation C10, passes through the first FC (with a size of 1024), and attaches the ReLu activation function. The size of the second FC is 8, which is equal to the number of target classes and is used to implement the final classification task.
[0052] The loss function of the noise reduction front-end model is MSE, and the loss value is mainly composed of two parts and , refers to the loss of amplitude prediction, is the loss of phase prediction.
[0053]
[0054]
[0055]
[0056] Among them, represents the features output by the amplitude branch of the decoder of the noise reduction front-end module, means taking the modulus operation on the feature A7 and then performing power-law compression on the spectrum; represents the original spectrogram data, means first for taking the modulus operation and then performing power-law compression on the spectrum. By applying power-law compression to change the dynamic range of the spectrogram, the weak part of the signal becomes more obvious relative to the strong part, which helps to analyze the subtle changes of the prominent signal.
[0057] In the recognition back-end, we selected the cross-entropy loss as the loss function. The cross-entropy loss is widely used in classification tasks due to its intuitiveness and effectiveness. This loss function guides the model to optimize parameters by quantifying the difference between the model prediction distribution and the true distribution. Specifically, the cross-entropy loss will encourage the model to assign a higher probability to the correct class, and at the same time effectively suppress the probability of the wrong class, thereby improving the classification performance.
[0058]
[0059] Among them, denotes the number of samples, denotes the th sample's true value of the th category, denotes the th sample's
[0060] Integrate the noise reduction front end, feature fusion, and recognition back end into one, and use the joint training method to optimize the noise reduction and recognition objectives simultaneously. The optimization method of joint training can not only improve the model performance, but also enhance its robustness and generalization ability. For the loss function of joint training :
[0061]
[0062] where are the relative weight parameters for these loss functions.
[0063] The recognition method of the present invention can significantly improve the feature expression ability of the model by introducing a feature fusion module, and effectively alleviate the feature mismatch problem between the noise reduction front end and the recognition back end. Feature fusion enhances the model's robustness to noise and improves the overall recognition performance by integrating amplitude and phase information. The feature clustering situations with and without feature fusion are as shown in Figure 7 and Figure 8 .
[0064] Using the method of fusing the amplitude features of noisy audio and the amplitude features of denoised audio, although it improves the noise robustness of underwater target recognition, it loses the phase information and limits the recognition performance. Therefore, we fully exploit the potential of phase information by fusing the amplitude features of denoised audio and the estimated phase features. The model is tested on noisy datasets and clean datasets. This model has improved by 6.99%, 6.61%, 7.04%, and 6.91% respectively in terms of average recognition rate, accuracy, recall rate, and F1 score. The specific results are as shown in Figure 9 .
[0065] The present invention adds a feature extraction module containing TF-transformer in the recognition. Among them, TF-transformer is responsible for capturing the complex temporal and frequency dependencies in the signal, while the convolutional layer is used to extract the local correlations of features.
[0066] In summary, the above are only the preferred embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for underwater acoustic signal recognition based on cross-attention feature fusion, characterized in that: include: Convert the underwater acoustic signal into spectrum data D; Inputting the spectrum data D into the noise reduction front-end module; The noise reduction front-end module includes an encoder part, an information exchange part, and a decoder part; each part is divided into an amplitude branch and a phase branch; The spectrum data D is processed in the amplitude branch of the encoder part to obtain the amplitude mask prediction data A1; the spectrum data D is processed in the phase branch of the encoder part to obtain the phase mask prediction data P1; In the information exchange part, the amplitude mask prediction data A1 and the phase mask prediction data P1 are shared, and each branch receives the key information of the other party, thereby obtaining the corresponding feature A4 in the amplitude branch and the feature P4 in the phase branch; In the decoder part, the amplitude branch decodes feature A4 to obtain feature A7; the phase branch decodes feature P4 to obtain feature P5; The feature A7 and the feature P5 are sent to a feature fusion module based on a cross attention mechanism for processing to obtain a fused feature C5; The fused feature C5 is sent to the recognition backend module for processing to identify the category of the target.
2. The method for underwater acoustic signal recognition based on cross-attention feature fusion as claimed in claim 1, characterized in that: When converting the underwater acoustic signal into the spectrum data D, the underwater acoustic signal is divided into multiple time periods; the segmented underwater acoustic signal is processed using Fourier transform to obtain the spectrum data D. ; Wherein, variable T represents the time step, F represents the number of frequency bands, that is, the number of frequency bands or frequency points into which the spectrum graph is divided along the frequency axis, and each frequency band represents the energy or amplitude information of the signal within the set frequency range; variable 2 represents the real and imaginary parts of the signal.
3. A method for underwater acoustic signal recognition based on cross-attention feature fusion as described in claim 1 or 2, characterized in that: In the encoder, the amplitude branch consists of two convolutional layers and two activation functions. The spectrum graph data D1 is first processed by the first convolutional layer and the first ReLu activation function, and then processed by the second convolutional layer and the second ReLu activation function to obtain feature A1; the phase branch consists of two convolutional layers, and the spectrum graph data D2 is processed by two convolutional layers in turn to obtain feature P1.
4. The method for underwater acoustic signal recognition based on cross-attention feature fusion as claimed in claim 3, characterized in that: In the amplitude branch of the information exchange part, feature A1 is processed by three convolutional layers in sequence, each of which is followed by a batch normalization layer BN and a ReLu activation function, and finally feature A2 is obtained; in the phase branch, feature P1 is first processed by linear normalization, and then processed by two convolutional layers in sequence to obtain feature P2; Feature A2 is processed through a convolution layer and Tanh activation function to obtain feature A3, which is then multiplied with feature P2 to obtain feature P4. After feature P2 is processed by a convolution layer and Tanh activation function to obtain feature P3, it is then dot-multiplied with feature A2 to obtain feature A4.
5. The method for underwater acoustic signal recognition based on cross-attention feature fusion as claimed in claim 4, characterized in that: In the decoder part, feature A4 passes through the convolution layer, the shaping layer and three fully connected layers FC in sequence to obtain feature A5; the activation function connected after the convolution layer is the Sigmoid activation function, the number of units of the first two fully connected layers FC is 600, the activation function connected after each fully connected layer FC is the ReLu activation function, the number of units of the last fully connected layer FC is the number of frequency bands, and the activation function is Sigmoid; feature P4 passes through the convolution layer and amplitude regularization in sequence to obtain feature P5, and completes the phase mask prediction; feature A5 is dot-multiplied with the input spectrum graph data D to obtain feature A6, and feature A6 is dot-multiplied with feature P5 to obtain feature A7 to complete the amplitude mask prediction.
6. The method for underwater acoustic signal recognition based on cross-attention feature fusion as claimed in claim 5, characterized in that: The feature fusion module adopts a dual-branch structure to independently process the amplitude and phase features, specifically including: Feature A7 and feature P5 are transformed into feature A8 and feature P6 respectively after Mel spectrum conversion and self-attention mechanism. Feature A8 and feature P6 are transformed into feature A9 and feature P7 respectively after two fully connected layers FC. They are added with feature A8 and feature P6 through skip connection to obtain feature A10 and feature P8 respectively. Feature A10 is used as the Q parameter input of cross attention, and feature P8 is used as the K parameter and V parameter input of cross attention. After cross attention, feature C1 is obtained. Feature C1 is added with feature A10 through skip connection to obtain feature C2. Feature C3 is obtained after a convolution layer and the subsequent activation function ReLu. Feature C3 is successively transformed into feature C4 after a fully connected layer FC and its activation function Sigmoid, convolution layer and the subsequent activation function ReLu. Feature C4 is added with feature C3 through skip connection to obtain feature C5, thus completing feature fusion.
7. The method for underwater acoustic signal recognition based on cross-attention feature fusion as claimed in claim 6, characterized in that: In the recognition backend module, feature C5 is first processed by two convolutional layers to obtain feature C6; feature C6 is further processed by three feature extraction modules in turn to obtain feature C7; Each of the feature extraction modules is composed of a TF-transformer module, three convolutional layers and a TF-transformer module connected in sequence, and is used to extract the features of underwater acoustic signals; wherein the three convolutional layers are used to extract the local correlation of features, and the sizes are 5×5, 1×25, and 5×5 respectively, and each convolutional layer is followed by a normalization layer, a ReLu activation function and a Dropout layer; The TF-transformer module is responsible for capturing the time and frequency dependencies in the signal; the TF-transformer module consists of two parts, the F-transformer module and the T-transformer module, which act on the signal frequency dimension and time dimension respectively; in the F-transformer module, feature C6 is processed by the multi-head attention mechanism to obtain feature F1, feature F2 is obtained by skip connection, feature F3 is obtained by the fully connected layer FC, feature F4 is obtained by LSTM and the fully connected layer FC, feature F5 is obtained by skip connection, feature F6 is obtained by the fully connected layer FC, and feature F7 is obtained by skip connection, and feature F7 is obtained by the T-transformer module to obtain feature T7; feature C7 is averaged and pooled to aggregate the information of the feature map in the time dimension and frequency dimension to obtain feature C8 and feature C9 respectively, feature C8 and feature C9 are obtained by addition operation to obtain feature C10, which is processed by the first fully connected layer FC and attached to the subsequent ReLu activation function, and finally processed by the second fully connected layer FC to obtain the final target category.
8. The method for underwater acoustic signal recognition based on cross-attention feature fusion as claimed in claim 7, characterized in that: It also includes training of the noise reduction front-end module; wherein the loss function used for training the noise reduction front-end module is the loss of amplitude prediction and the loss of phase prediction Weighted summation, specifically: in, Characterizes the amplitude branch output of the decoder of the noise reduction front-end module, Express Modulo operation; represents the raw spectrogram data, Indicates first Modulo operation and then power law compression spectrum operation symbol Represents the dot product operation.
9. The method for underwater acoustic signal recognition based on cross-attention feature fusion as claimed in claim 8, characterized in that: It also includes the training of the recognition backend module, where the loss function used to train the recognition backend module is as follows: in, is the number of samples, For the The sample The true value of each category, To identify the backend module The sample The predicted probability of each category, C represents the number of target categories.
10. The underwater acoustic signal recognition method based on cross-attention feature fusion according to claim 9, characterized in that: The denoising front-end module, feature fusion module and recognition back-end module are jointly trained, and the loss function of the joint training is the weighted sum of the loss functions of the denoising front-end module and the recognition back-end module.
Citation Information
Patent Citations
Short burst underwater acoustic communication signal modulation identification method based on deep learning
CN112615804A
Underwater sound target identification method fusing attention mechanism and deep residual shrinkage network
CN117310668A