Underwater sound biological target identification method based on attention mechanism

By adopting a convolutional neural network model based on attention mechanism in water acoustic target recognition, the characteristic information of water acoustic signals is extracted, and the problem of identification deviation under complex noise interference is solved, and higher recognition accuracy and robustness are achieved.

CN120126489AActive Publication Date: 2025-06-10OCEAN UNIV OF CHINA
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510421665.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-06-10
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

In the case of complex noise interference, existing water acoustic target recognition methods have problems of identification deviation and inaccurate feature extraction.

Method used

The convolutional neural network model based on attention mechanism is adopted to extract the characteristic information of water acoustic signals in different scales through the channel attention mechanism to distinguish key sounds and noises, thereby improving the recognition accuracy.

Benefits of technology

Under complex noise interference, the accuracy of water acoustic biological target recognition and the robustness of the model are improved, and the critical area extraction capability of audio features is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126489A_ABST
    Figure CN120126489A_ABST
Patent Text Reader

Abstract

The invention discloses an underwater sound biological target recognition method based on an attention mechanism, and belongs to the technical field of underwater target recognition. The invention provides an attention mechanism assisted convolutional neural network model for underwater acoustic signal target recognition, and underwater acoustic signal feature information of different scale spaces is extracted through a multi-head attention mechanism to improve the target recognition precision under the condition of noise interference; and the most critical audio features for the classification task are extracted, so that the classification performance is improved. According to the method, under the condition that the audio sample contains complex noise, the key area features of the audio sample can be extracted, the robustness and generalization ability of the model are improved, the high accuracy of a target recognition task is achieved, and technical support is provided for ocean target recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of underwater acoustic biological target recognition, and particularly relates to an underwater acoustic biological target recognition method based on an attention mechanism. Background Art

[0002] Sound waves are generally considered to be the best information carriers for perceiving and recognizing underwater targets, and are the only form of energy known to humans that can propagate long distances in water. Underwater acoustic biological target recognition technology has important application values in the fields of marine ecological protection, fishery management, marine biodiversity research, etc., and has become a hot issue of concern to research scholars. However, the high-noise and strong-interference environment of the ocean, the complex underwater acoustic propagation channel, and the lack of data resources have brought great challenges to underwater acoustic target recognition.

[0003] Early underwater acoustic target recognition was mainly achieved by machine learning algorithms, including Hidden Markov Model (HMM), Support Vector Machine Model (SVM), Gaussian Mixture Model (GMM), and Decision Tree (DT), etc. Shameer K. Mohammed et al. proposed an underwater target classifier based on GTCC and HMM with 20 states in 2018 and evaluated it on measured data. However, the parameter estimation process of HMM is relatively complex and requires a large number of data samples, so it is very likely that the requirements cannot be met in actual target recognition tasks. Parada et al. used unpredictability measures and MUSIC algorithms to extract features in 2014 to improve the recognition performance of GMM. However, for large-scale underwater biological audio datasets, the training time of GMM may be very long, and in real-time classification scenarios, the real-time requirements may not be met. In 2015, Moura et al. applied One-Class Support Vector Machine (SVM) to passive sonar system detection to solve the classification problem of sparse negative samples. However, when dealing with large-scale underwater biological audio datasets, the training speed of SVM will significantly slow down. The structure of a decision tree is similar to a flowchart and can naturally process audio data with multiple features and classify by sequentially judging different features. However, decision trees are prone to overfitting, and the local optimality of feature principles will affect the classification effect.

[0004] In recent years, the theoretical research of deep learning has developed rapidly, providing new ideas for underwater biological target recognition. In 2019, Yang et al. proposed a new end-to-end neural network model for underwater acoustic target recognition - the Auditory Perception Inspired Deep Neural Network, which is used to achieve the decomposition, feature extraction and classification of ship radiated noise for underwater acoustic target recognition tasks. In 2020, in order to reduce the network complexity, Miao et al. designed a deep convolutional neural network structure suitable for underwater acoustic signal classification, called the TF Feature Network, and evaluated the network performance on two underwater acoustic target recognition datasets. In 2021, Panetta et al. proposed an underwater image enhancement network based on generative adversarial networks, which improves the performance of the tracker for underwater data by converting the vision of the underwater area into an enhanced / cleared underwater domain. In 2021, Liu et al. proposed a convolutional recurrent neural network model, which includes three steps to process the recognition of underwater targets: extracting three channels from the Mel-spectrum log as input; using data augmentation to expand the data in the time domain and time-frequency domain; using the proposed model for automatic feature learning to identify the target category. Xiao et al. proposed an attention-based neural network for target recognition in multi-source interference pressure spectrograms in 2021, using the attention module to detect the internal working principle of the neural network. In 2023, Hauer et al. proposed the ORCA-SPY framework for fully automatic biogenic source simulation, classification and localization of passive killer whale acoustic monitoring, and embedded the framework into PAMGuard. Even in a noisy environment, ORCA-SPY can be used to detect and track target signals. In 2023, Xue et al. added a channel attention mechanism based on the residual neural network. This channel attention mechanism can enhance stable spectral features and eliminate unstable signals caused by Doppler frequency shift. In the same year, Tang et al. proposed a new method to convert the Mel spectrogram into three-dimensional data and introduced an efficient three-dimensional spectrogram network that processes the time and frequency dimensions separately. Feng et al. proposed an underwater acoustic target recognition model based on Transformer, and verified through experiments that the proposed model has good accuracy performance.

[0005] In summary, the above methods have achieved good performance in underwater acoustic target recognition, but they do not distinguish between key sounds and noise during feature extraction. Therefore, there is a large deviation for data affected by complex noise in the actual ocean. Summary of the Invention

[0006] The purpose of the present invention is to provide an underwater acoustic biological target recognition method based on the attention mechanism to make up for the deficiencies of the prior art.

[0007] The real marine environment is complex and changeable. Therefore, the acoustic signal data obtained from the marine environment is interfered by complex environmental noise, and there are problems such as uneven resolution, too weak target signals, and uneven intensity. When the existing target recognition methods extract features, they do not distinguish the key information of the acoustic signal from the irrelevant noise. Although the results obtained during the training of the model have achieved good recognition performance, there are still large deviations in target recognition under actual complex noise interference. The present invention proposes a convolutional neural network model assisted by an attention mechanism for underwater acoustic signal target recognition. Through the channel attention mechanism, the feature information of underwater acoustic signals in different scale spaces is extracted to improve the target recognition accuracy under noise interference; the audio features that are most critical for the classification task are extracted, thereby improving the classification performance.

[0008] To achieve the above object, the present invention adopts the following technical solutions: An underwater acoustic biological target recognition method based on an attention mechanism, comprising the following steps: S1: Obtain underwater biological audio data and perform preprocessing; S2: Extract the Mel spectrum features of the audio data and perform data augmentation using a data augmentation method; S3: Divide the processed data set into a training set, a validation set, and a test set; S4: Construct a neural network model based on the attention mechanism - the ACNN model. The ACNN model includes an input layer, a convolutional block, an adaptive average pooling layer, a fully connected layer, a regularization layer, and an output layer; the convolutional block contains four layers, and each layer includes an attention module, a convolutional layer, a ReLU activation layer, and a batch normalization layer; S5: ACNN model verification and parameter update: Evaluate the model performance on the validation set, calculate the loss function and accuracy, and adjust the ACNN model parameters and architecture according to the validation set results to optimize the performance; S6: Based on the ACNN model optimized by training, perform testing, and compare the predicted labels with the true labels to test the recognition accuracy of the ACNN model.

[0009] Further, the S1 includes: S1-1: Obtain underwater biological audio data, and the format of the audio data is the wav format; then read the audio file to obtain the audio signal and the sampling rate; S1-2: Use spectral subtraction to remove noise, estimate the noise power spectrum, subtract it from the power spectrum of the audio signal, and restore the pure signal; the spectral subtraction formula is: ; where, is the power spectrum of the noisy audio, is the noise power spectrum, To adjust parameters, is the estimated pure signal power; S1-3: Use the fixed-length segmentation method to segment the long audio into short segments. Segment the audio at a fixed length (0.02 seconds). For audio with a sampling rate of the segmentation points are ×2 sampling points; S1-4: Use the Librosa library to implement sampling rate conversion and unify the sampling rate to ensure data consistency.

[0010] Furthermore, the specific steps of S2 are as follows: S2-1: First standardize the sampling rate of the audio data file so that all arrays have the same size, then adjust the size of the audio samples to have the same length, and then perform random augmentation on the original audio signal; S2-2: Convert the augmented audio to a Mel spectrogram. The transformation formula is: ; ; where represents the original frequency, represents the converted Mel frequency; The calculation formula of the Mel filter is as follows: ; where represents the filter number, and 、 correspond to the start point, middle point, and end point of the th filter respectively; S2-3: When performing audio feature extraction, each audio segment can be represented by the corresponding Mel spectrogram, and a certain frequency band corresponding to each frame is a feature; S2-4: After extracting the Mel spectrogram features of the audio samples, use data augmentation techniques to process them, using frequency shift and time masking processing.

[0011] Furthermore, in S3, the leave-one-out method is used for dataset partitioning, that is, when the number of samples in each category is very limited, one sample is left as the validation set, and the rest are used for training.

[0012] Further, in S4, the convolutional block contains four layers: the convolutional layer (Conv1) in the first layer is set to have 2 input channels, 8 output channels, and a 3×3 convolutional kernel, where the height and width are the spatial dimensions of the convolutional layer output, corresponding to the time step t and frequency dimension f of the audio feature map; the convolutional layer (Conv2) in the second layer is set to have 8 input channels and 16 output channels; the convolutional layer (Conv3) in the third layer is set to have 16 input channels and 32 output channels; the convolutional layer (Conv4) in the fourth layer is set to have 32 input channels and 64 output channels.

[0013] Further, in S4, the data processing process includes the following steps: S4-1: Input the data processed by S1 to S3 into the ACNN model through the input layer. The shape of the input data is: batch size × number of channels × time step × feature dimension; S4-2: The data reaches the convolutional block: in each layer, the data first passes through the attention module, then gradually extracts audio features through the convolutional layer, introduces non-linearity through the ReLU activation layer, and finally normalizes the output through the batch normalization layer; S4-3: After the operations in the convolutional block are completed, the data enters the adaptive average pooling layer to reduce the spatial dimension of the feature map; S4-4: After flattening the feature map of the pooling layer, map it to the output of the number of categories through a fully connected layer; S4-5: After the fully connected layer, use the regularization layer to prevent overfitting; S4-6: Finally, output through the output layer.

[0014] Further, in the attention module, the audio feature data First, it is input through the input layer. The input shape is batch size × number of channels × time step × feature dimension. The input features are preliminarily processed through the attention vector layer to generate an attention scoring matrix ; then, it passes through the activation dense layer and uses The activation function to weight the attention scoring matrix To generate an attention weight vector ; after the activation dense layer, there is a Gaussian layer. The Gaussian layer multiplies the attention weight vector With the Gaussian kernel function to generate a Gaussian layer output vector , and then, the attention module masks the output sequence of the Gaussian layer and the input sequence through the merge and fusion layer to generate the final output vector .

[0015] Even further, the attention module includes: (1) The soft mask passes through is implemented by an activated dense layer, and each input data sample is a vector of length and size of the attention weight vector is expressed as: ; ; where is the attention scoring matrix of size , learned by CNN and function, representing the correlation between each input feature in and the attention event; ; (2) Define the Gaussian layer, multiply it with the above Gaussian kernel function, and represent it with matrix operations: ; ; where is the output vector of the Gaussian layer of size , and is the standard deviation of the Gaussian function; (3) The output vector of the attention module is a vector of size f, expressed by Hadamar multiplication (denoted as ⊙) as: .

[0016] Furthermore, in S5, the cross-entropy loss is used as the loss function. There are N samples, each sample has C classes, and the predicted output of the model is , and the true label is . Then the calculation formula of the cross-entropy loss function L is: ; where is the true label (0 or 1) of the -th sample belonging to the -th class, is the probability that the model predicts the -th sample belonging to the -th class, and the calculation process is: Forward propagate the input data through the model to obtain the predicted output of each sample; Apply the softmax activation to convert the logits output by the model into a probability distribution: ; where is the -th sample output by the model and the ​The logits value of the class; then for each sample , calculate the cross-entropy loss between its predicted probability distribution and the true label distribution: ; Average the losses of all samples to obtain the final loss value: ; Calculate the gradient of the loss function with respect to the model parameters through backpropagation, and use the optimizer to update the model parameters according to the gradient.

[0017] Compared with the prior art, the present invention has at least the following beneficial effects: The present invention can extract the key region features of audio samples in the case where the audio samples contain complex noises, improve the robustness and generalization ability of the model, achieve a high accuracy rate in the target recognition task, and provide technical support for marine target recognition. Description of the Drawings

[0018] Figure 1 is the audio classification process based on underwater acoustic target signals.

[0019] Figure 2 is the implementation flowchart of the present invention.

[0020] Figure 3 is the waveform diagram of an audio sample of a certain marine organism.

[0021] Figure 4 is the Mel spectrogram of an audio sample of a certain marine organism.

[0022] Figure 5 is the architecture of the underwater biological audio classification model based on the attention mechanism.

[0023] Figure 6 is the confusion matrix obtained after the model is trained. Detailed Embodiments

[0024] To make the purpose, technical solutions and advantages of the present disclosure clearer and more understandable, the following further describes the present disclosure in detail with reference to specific embodiments and the accompanying drawings.

[0025] Embodiment 1: The framework of the existing mainstream underwater acoustic target recognition system is as Figure 1 shown, and the whole workflow includes target signal acquisition, database creation, data preprocessing, feature extraction, and target recognition.

[0026] This embodiment provides an underwater acoustic biological target recognition method based on the attention mechanism, as Figure 2 shown, including the following steps: Step 1: Dataset preparation Collect underwater biological audio data, preprocess the audio files, perform denoising, segmentation, sampling rate conversion, etc. Take an audio dataset containing 32 marine biological sound signal files as an example, read the audio files and draw the waveform diagrams of the audio files and save them. Figure 3 It is an example waveform diagram of an audio sample of a certain marine organism.

[0027] First, obtain underwater biological audio data, that is, marine biological sound signal files, from a public dataset, covering various marine organisms such as whales and dolphins. The format of the audio data is the wav format; then use the audio processing library Librosa (Python) to read the audio files and obtain the audio signals and sampling rates.

[0028] Due to the complex marine environment, underwater biological audio data is usually interfered by background noises such as ship noises and water flow sounds, which are manifested as random or periodic signals, will mask the target biological sounds, and reduce the recognition accuracy. Therefore, denoising processing is carried out. The denoising method uses spectral subtraction to estimate the noise power spectrum and subtract it from the power spectrum of the audio signal to restore the pure signal. The spectral subtraction formula is: ; Among them, is the power spectrum of the noisy audio, is the noise power spectrum, is the adjustment parameter, is the estimated pure signal power.

[0029] To facilitate the processing and feature extraction of the model, use the fixed-length segmentation method to segment the long audio into short segments, segment the audio at a fixed length (0.02 seconds). For the audio with a sampling rate of , the segmentation points are ×2 sampling points.

[0030] Since the sampling rates of audio from different sources are different, the Librosa library is used to implement sampling rate conversion to unify the sampling rate to ensure data consistency and facilitate model input.

[0031] Step 2: Feature extraction Extract the Mel-spectrogram features of the audio files and perform data augmentation using data augmentation methods. During the audio preprocessing process, first standardize the sampling rate of the loaded audio data files so that all arrays have the same size. After that, use the method of silent padding or truncating the length to adjust the sizes of all audio samples to have the same length, and then randomly augment the original audio signal by applying time offsets to move the audio left or right by a random amount.

[0032] Convert the augmented audio into a Mel spectrogram. The Mel scale is a linear transformation of Hertz. For a signal in Mel scale units, the human perception of signals with the same frequency difference is almost the same. A commonly used transformation formula is: ; ; where represents the original frequency, represents the converted Mel frequency. The Mel spectrogram is a spectrogram in Mel scale. The original signal is preprocessed, framed and windowed, and finally obtained by multiplying the spectrum with several Mel filters. The Mel filter bank is a set of triangular filters with equal height, and the starting point of each filter is at the midpoint of the previous filter. The specific calculation formula of the Mel filter is as follows: ; where represents the filter number, and 、 correspond to the starting point, midpoint and ending point of the th filter respectively. When performing audio feature extraction, each audio segment can be represented by the corresponding Mel spectrogram, and a certain frequency band corresponding to each frame is a feature. After extracting the Mel spectrogram features of the audio samples, use data augmentation technology to process them. This technology uses two methods: frequency shift (randomly masking a series of continuous frequencies by adding horizontal bars on the spectrogram) and time masking (randomly occluding a time range from the spectrogram using vertical lines). The Mel spectrogram of a certain audio sample of a certain marine organism obtained is as Figure 4 shown.

[0033] The Mel spectrogram features are designed based on the perceptual features of the human auditory system. It simulates the sensitivity of the human ear to sounds of different frequencies, making the features more in line with human perception of sounds. The Mel spectrogram converts the audio signal from the time domain to the frequency domain and uses the Mel filter bank to reduce the dimensionality of the data, which helps to reduce the computational complexity and improve the processing speed. The Mel spectrogram can capture the global features of the audio signal and has a certain robustness to noise and changes in the audio signal, so that the model can maintain good performance in a complex noise environment.

[0034] Step 3: Dataset division Divide the dataset into a training set, a validation set and a test set. The present invention uses the leave-one-out method for dataset division, that is, when the number of samples in each category is very limited, one sample is left as the validation set and the rest are used for training. This method is useful when the number of categories is large but the number of samples in each category is small.

[0035] Step 4: Construct a neural network model based on the attention mechanism - the ACNN model As Figure 5 shown, the ACNN model consists of an input layer, a convolutional block, an adaptive average pooling layer, a fully connected layer, a regularization layer, and an output layer. The data passes through the input layer, the convolutional block, the adaptive average pooling layer, the fully connected layer, and the regularization layer in sequence, and finally outputs through the output layer. The data processed in Step 1, Step 2, and Step 3 is input into the model through the input layer. The shape and size of the input data are: batch size × number of channels × time step × feature dimension. The data reaches the convolutional block. The convolutional block contains four layers. Each layer contains an attention module, a convolutional layer, a ReLU activation layer, and a batch normalization layer (BatchNorm2D). In the first layer, the data first passes through the attention module to focus on important features, then gradually extracts audio features through the convolutional layer, then introduces non-linearity through the ReLU activation layer, and finally normalizes the output through the batch normalization layer. The convolutional layer (Conv1) in this layer is set to have 2 input channels, 8 output channels, and a 3×3 convolutional kernel, where the height and width are the spatial dimensions of the output of the convolutional layer, corresponding to the time step t and the frequency dimension f of the audio feature map. Secondly, the data reaches the second layer of the convolutional block. It still first passes through the attention module, and then passes through the convolutional layer, the activation layer, and the batch normalization layer in sequence. The convolutional layer (Conv2) in the second layer is set to have 8 input channels and 16 output channels. After the second layer outputs, the data reaches the third layer of the convolutional block. The convolutional layer (Conv3) in this layer is set to have 16 input channels and 32 output channels. Finally, the data reaches the fourth layer of the convolutional block. The convolutional layer (Conv4) in the fourth layer is set to have 32 input channels and 64 output channels. After the operations of the data in the convolutional block are completed, it enters the adaptive average pooling layer to reduce the spatial dimension of the feature map. After flattening the feature map of the pooling layer, it is mapped to the output of the number of categories through a fully connected layer. After the fully connected layer, a regularization (Dropout) layer is used to prevent overfitting. Finally, the number of categories is output through the output layer.

[0036] This model architecture gradually extracts audio features through multiple convolutional layers and enhances the attention to important features through the attention module. Finally, classification is performed through the fully connected layer and the regularization (Dropout) layer to improve the generalization ability of the model and prevent overfitting. This architecture is suitable for processing time-series data of audio signals, such as target recognition tasks. The present invention will specifically describe the specific implementation of the attention module.

[0037] In Figure 5 the ACNN model architecture of is the batch size, c is the number of channels, t is the time step, f is the size of the frequency dimension, = 16, the number of task categories = 32, and the attention module is located in the convolutional block of the model and belongs to the first module of each layer in the convolutional block. In this module, the audio feature data First, it is input through the input layer. The input shape is (batch size × number of channels × time steps × feature dimension). The input features are preliminarily processed through the attention vector layer to generate an attention scoring matrix , and this matrix is used to measure the importance of each frequency component in the input features. Then, it passes through the activation dense layer and uses The activation function to weight the attention scoring matrix to generate an attention weight vector , which is used to control the sensitivity of the frequency components. To prevent over-concentration, a Gaussian layer follows the activation dense layer. The Gaussian layer multiplies the attention weight vector with the Gaussian kernel function to generate the output vector of the Gaussian layer . The role of the Gaussian kernel function is to prevent the attention from over-concentrating on a single frequency point, thereby improving the robustness of the model to frequency fluctuations. For underwater organisms, in addition to continuous broadband line spectra, it also reduces the influence of Doppler frequency shift, measurement errors, and other factors on the random variation of features. Then, the attention module masks the output sequence of the Gaussian layer and the input sequence through the merge and fusion layer to generate the final output vector . The attention layer is like a mask that only retains the features related to the target.

[0038] To construct the attention module using the mask mechanism of soft attention, the most basic soft mask can be implemented through The activation dense layer. Each input data sample is a vector with a length of , and the attention weight vector with a size of is expressed as: ; ; where is the attention scoring matrix with a size of , is learned by CNN and function, representing the correlation between each input feature in and the attention event. The stronger the correlation, the higher the score. Then, the definition of the Gaussian layer is carried out, multiplying the above with the Gaussian kernel function, and expressed in matrix operations: ; ; where is the output vector of the Gaussian layer with a size of , that is, the attention weight layer, is the standard deviation of the Gaussian function, The larger it is, the more dispersed the attention is, which helps to prevent the attention from being overly concentrated on a single frequency point, thus improving the robustness to frequency fluctuations. Finally, the output vector is a vector of size f, expressed through Hadamard multiplication (denoted as ⊙) as: .

[0039] In this model, to reduce the overfitting phenomenon, Dropout regularization is adopted. A Dropout layer is added after the fully connected layers of the model. During the training process, a certain proportion of neurons are randomly inactivated, enabling the model to use different subsets of neurons in each training iteration, thereby enhancing the generalization ability of the model.

[0040] Meanwhile, to prevent overlearning and avoid local minima, the present invention uses the Adam optimizer to dynamically adjust the learning rate. During the model training process, the Adam optimizer is used to update the parameters of the model. The Adam optimizer automatically adjusts the learning rate of each parameter by calculating the first-order moment estimate and the second-order moment estimate of the gradient. This enables the model to better adapt to the distribution characteristics of the data when dealing with classification tasks, improving the training efficiency and performance of the model.

[0041] Step 5. Model verification and parameter update Evaluate the model performance on the validation set, calculate the loss function and accuracy. According to the results on the validation set, adjust the model parameters and architecture to optimize the performance.

[0042] In the classification task, the cross-entropy loss is used as the loss function. There are N samples, each sample has C classes, and the predicted output of the model is , and the true label is . Then the calculation formula for the cross-entropy loss function L is: ; where is the true label (0 or 1) of the th sample belonging to the th class, is the probability that the model predicts the th sample belonging to the th class. The calculation process is as follows: Pass the input data through the model for forward propagation to obtain the predicted output of each sample; Apply the softmax activation to convert the logits output by the model into a probability distribution: ; where is the output of the model for the th sample for the The logit values of the classes; then for each sample , calculate the cross-entropy loss between its predicted probability distribution and the true label distribution: ; Average the losses of all samples to obtain the final loss value: ; Calculate the gradient of the loss function with respect to the model parameters through backpropagation, and use the optimizer to update the model parameters according to the gradient.

[0043] Step 6. Model testing Test the model on the test set. After processing a certain type of marine biological audio as described in the steps, use it as the test set input to the model for testing, and compare the predicted labels with the true labels to test the classification accuracy of the model.

[0044] Step 7. Performance evaluation, use metrics such as accuracy, F1-score, and confusion matrix to evaluate the model performance.

[0045] Accuracy reflects the overall proportion of correct classifications. The F1-score is the harmonic mean of precision and recall. The confusion matrix shows the predicted and actual results for each class, thereby discovering the recognition problems of the model for certain classes. The accuracy ( ) calculation formula is: ; Precision ) calculation formula is: ; Recall ) calculation formula is: ; The F1-score calculation formula is: ; Evaluate the performance of the model in the underwater acoustic biological target recognition task more comprehensively through the calculation of these metrics.

[0046] Example 2:

[0047] Since the cost function of target classification may be non-convex and may get stuck at local minima due to environmental noise, reverberation, and multipath interference in the underwater acoustic channel, the present invention adopts the Adam optimizer in the proposed model to dynamically adjust the learning rate to prevent overfitting and avoid local minima. Due to the sensitivity to environmental changes and signal distortion, deep neural network methods often suffer from overfitting. Therefore, dropout regularization is adopted to reduce the network parameters before the output.

[0048] In this embodiment, the processed dataset is trained on the CAMPPlus, ERes2Net, ResNetSE, CNN models, and the proposed ACNN model in this paper, and the classification accuracies obtained are shown in the following table:

[0049] Table Classification task accuracies and F1 scores of different deep neural networks. The F1 score is the harmonic mean of the precision and recall of the model .

[0050] Table 1 shows the classification task accuracies obtained by training our dataset on other models and the model proposed in this paper. The dataset used in this invention contains a total of 32 biological species, and each underwater creature has approximately 5 audio samples.

[0051] Figure 6 The classification confusion matrix obtained by training the model proposed in this invention can show the accuracy between the predicted label and the true label, indicating that the model proposed in this invention has good performance and high accuracy, and can be widely applied to the underwater biological target recognition task in complex noise scenarios.

[0052] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the disclosure of this invention. It should be understood that the above are only specific embodiments of the disclosure of this invention and are not used to limit the disclosure of this invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the disclosure of this invention shall be included within the protection scope of the disclosure of this invention.

Claims

1. A method for underwater acoustic biological target recognition based on attention mechanism, characterized in that: The following steps are involved: S1: Acquire underwater biological audio data and perform preprocessing; S2: Extract the Mel-spectrogram features of the audio data and use the data enhancement method to perform data enhancement; S3: Divide the processed data set into training set, validation set and test set; S4: Construct an attention-based neural network model - ACNN model, which includes an input layer, a convolution block, an adaptive average pooling layer, a fully connected layer, a regularization layer and an output layer; the convolution block contains four layers, each of which includes an attention module, a convolution layer, a ReLU activation layer, and a batch normalization layer; S5: ACNN model verification and parameter update: Evaluate model performance on the validation set, calculate loss function and accuracy, and adjust ACNN model parameters and architecture to optimize performance based on the validation set results; S6: Test the ACNN model after training and optimization, and compare the predicted labels with the actual labels to test the recognition accuracy of the ACNN model.

2. The underwater acoustic biological target recognition method according to claim 1, characterized in that: The S1 includes: S1-1: Acquire underwater biological audio data in wav format; then read the audio file to obtain the audio signal and sampling rate; S1-2: Use spectral subtraction to denoise. The spectral subtraction formula is: ; in, is the power spectrum of the noise frequency, is the noise power spectrum, To adjust the parameters, is the estimated pure signal power; S1-3: Use the fixed-length segmentation method to segment long audio into short segments. The audio is split at ×2 sampling points; S1-4: Use the Librosa library to implement sampling rate conversion and unify the sampling rate to ensure data consistency.

3. The underwater acoustic biological target recognition method according to claim 1, characterized in that: The S2 includes: S2-1: First standardize the sampling rate of the audio data file, and then randomly augment the original audio signal; S2-2: Convert the augmented audio into a Mel-spectrogram. The transformation formula is: ; ; in Represents the original frequency, Represents the converted Mel frequency; the calculation formula of the Mel filter is as follows: ; in Represents the filter number, and , The corresponding The starting point, middle point and end point of each filter; S2-3: When extracting audio features, each audio segment can be represented by a corresponding Mel spectrum, and a frequency band corresponding to each frame is a feature; S2-4: After extracting the Mel spectrum features of the audio sample, it is processed using data augmentation technology, using frequency shifting and time masking.

4. The underwater acoustic biological target recognition method according to claim 1, characterized in that: In S3, the leave-one-out method is used to divide the data set, that is, when the samples of each category are very limited, one sample is reserved as a validation set and the rest are used for training.

5. The underwater acoustic biological target recognition method according to claim 1, characterized in that: In the S4, the convolution block includes four layers: the convolution layer in the first layer is set to 2 input channels, 8 output channels, and a 3×3 convolution kernel, where the height and width are the spatial dimensions of the convolution layer output, corresponding to the time step t and frequency dimension f of the audio feature map; the convolution layer in the second layer is set to 8 input channels and 16 output channels; the convolution layer in the third layer is set to 16 input channels and 32 output channels; the convolution layer in the fourth layer is set to 32 input channels and 64 output channels.

6. The underwater acoustic biological target recognition method according to claim 1, characterized in that: In S4, the data processing process includes the following steps: S4-1: Input the data processed by S1 to S3 into the ACNN model through the input layer. The shape of the input data is: batch size × number of channels × time step × feature dimension; S4-2: Data reaches the convolutional block: In each layer, the data first passes through the attention module, then gradually extracts audio features through the convolution layer, then introduces nonlinearity through the ReLU activation layer, and finally standardizes the output through the batch normalization layer; S4-3: After the operation in the convolution block, the data enters the adaptive average pooling layer to reduce the spatial dimension of the feature map; S4-4: After flattening the feature map of the pooling layer, it is mapped to the output of the number of categories through a fully connected layer; S4-5: After the fully connected layer, a regularization layer is used to prevent overfitting; S4-6: Finally, output through the output layer.

7. The underwater acoustic biological target recognition method according to claim 1, characterized in that: In the attention module, the audio feature data First, the input is input through the input layer, and the input shape size is batch size × number of channels × time step × feature dimension. The input features are preliminarily processed through the attention vector layer to generate the attention score matrix ; After activating the dense layer, use Activation function attention score matrix Weighted, generate attention weight vector ; After activating the dense layer, a Gaussian layer is set, which converts the attention weight vector Multiply with the Gaussian kernel function to generate the Gaussian layer output vector Then, the attention module merges the output sequence of the Gaussian layer with the input sequence mask through the fusion layer to generate the final output vector .

8. The underwater acoustic biological target recognition method according to claim 7, characterized in that: The attention module comprises: (1) Soft mask pass The dense layer activated is implemented so that each input data sample is a length of A vector of size The attention weight vector It is expressed as: ; in The size is The attention score matrix, By CNN and Function to learn, representation The correlation between each input feature and the event of interest; (2) Define the Gaussian layer, multiply the above with the Gaussian kernel function, and express it using matrix operations: ; ; in Is the size of The output vector of the Gaussian layer is is the standard deviation of the Gaussian function; (3) Output vector of the attention module is a vector of size f, expressed by Hadamar multiplication as: 。 9. The underwater acoustic biological target recognition method according to claim 1, characterized in that: In S5, cross entropy loss is used as the loss function. There are N samples, each sample has C categories, and the prediction output of the model is , the true label is , then the calculation formula of the cross entropy loss function L is: ; in It is The samples belong to The true label of the class, The model predicts The samples belong to The probability of a class is calculated as follows: Propagate the input data forward through the model to obtain the predicted output for each sample ; Apply softmax activation to convert the logits output by the model into a probability distribution: ; in The output of the model is Sample No. The logits value of the class; then for each sample , calculate the cross entropy loss between its predicted probability distribution and the true label distribution: ; Average the losses of all samples to get the final loss value: ; The gradient of the loss function to the model parameters is calculated through back propagation, and the optimizer is used to update the model parameters according to the gradient.

Citation Information

Patent Citations

  • Underwater sound target identification method and system based on Mel-cepstrum and attention residual network

    CN116310770A

  • Underwater organism distribution characteristic analysis method based on underwater sound signal detection

    CN116973922A

  • Ship radiation noise target identification method fusing attention mechanism

    CN118262741A