An Underwater Acoustic Target Recognition Method Based on an Adaptive Multi-Feature Fusion Model
Through the adaptive multi-feature fusion model, combined with LSTM, 1D-CNN and 2D-CNN networks, the multi-dimensional features of the water acoustic signal are extracted and adaptively weighted fusion is carried out, which solves the problem of low recognition accuracy caused by single feature extraction and achieves higher water acoustic target recognition accuracy.
Patent Information
- Application Number
- CN202211618499.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-15
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-12-15
AI Technical Summary
Most existing water acoustic target recognition methods are based on a single time domain or frequency domain signal to extract water acoustic features, resulting in low recognition accuracy and failure to effectively mine time and frequency complementary information in the time domain and frequency domain.
Adaptive multi-feature fusion model is adopted, and the depth timing, depth space and depth frequency domain characteristics of the water acoustic signals are extracted through LSTM, 1D-CNN and 2D-CNN networks, and adaptive weighted fusion is used to enhance the discriminant characteristics.
The accuracy of water acoustic target recognition is significantly improved, and the recognition accuracy is improved through multi-dimensional feature extraction and adaptive weighting fusion, especially its performance on the ShipEar dataset.
Smart Images

Figure CN115909040B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of underwater acoustic target recognition, and in particular relates to an underwater acoustic target recognition method based on an adaptive multi-feature fusion model. Background Art
[0002] Hydroacoustic target recognition is one of the most important research directions in hydroacoustic signal processing. It is of great significance to the national economy and national defense and military, so it has become a hot topic in the field of hydroacoustics. The application of hydroacoustic signals for underwater detection, communication, life-saving and marine development is currently the most effective means. In underwater early warning defense and military offensive activities, sonar needs to distinguish the authenticity of the target through the received noise signal, and identify the type of each target when multiple targets are detected at the same time. According to the above two judgment results, it is decided what kind of action to take against the target, such as attack or avoidance.
[0003] The core of underwater target recognition lies in the processing of underwater acoustic signals. The sound source and propagation environment of underwater acoustic signals lead to the complexity of their signals. The noise sources are different, the radiated noise is very different, and the marine environment is complex, diverse, and time-varying. Therefore, the signals received by passive sonar are also very different. How to extract features that can be used to identify targets is a key issue in passive recognition of underwater acoustic targets and the primary issue for the automation of target recognition. This also makes the problem of underwater acoustic target recognition more challenging than ordinary speech recognition. The current methods are mainly divided into two types based on the feature extraction method. The first is to extract the features of underwater acoustic signals based on audio signals in the time domain. Among them, the typical method is to combine a one-dimensional convolutional neural network (1D-CNN) with LSTM, and use audio (MeI-scale FreguencyCeptraI Coefficients, MFCC) features as input to identify underwater acoustic targets. The second is to extract the features of underwater acoustic signals based on a two-dimensional spectrogram in the frequency domain. Among them, the typical method is to convert the underwater acoustic signal into a two-dimensional spectrogram first, and then input it into a two-dimensional convolutional neural network (2D-CNN) for recognition. The experimental results of measured data show that converting underwater acoustic signals into two-dimensional time-frequency spectra can effectively reduce the impact of noise, and thus can effectively improve classification and recognition performance. However, most of these methods extract the characteristics of underwater acoustic signals based on time-domain audio signals or frequency-domain spectrograms, and the perspectives considered are relatively single, and they do not start from the perspectives of both time and frequency domains to mine the time-frequency complementary information corresponding to the time-domain audio signals and the two-dimensional frequency-domain spectrograms, which is helpful for improving the accuracy of underwater acoustic target recognition.
[0004] In summary, at present, most of the underwater acoustic target recognition methods based on deep learning extract underwater acoustic features based on single time-domain or frequency-domain signals. However, considering only the time-domain audio signals or the two-dimensional time-frequency spectrograms in the frequency domain will miss some time-frequency information, resulting in insufficient recognition accuracy. Therefore, high-precision underwater acoustic target recognition methods have always been a hot issue for researchers in this field. Summary of the Invention
[0005] The present invention specifically proposes an underwater acoustic target recognition method based on an adaptive multi-feature fusion model to solve the problem that most of the existing underwater acoustic target recognition methods extract underwater acoustic features based on single time-domain or frequency-domain signals, and considering only the time-domain audio signals or the two-dimensional time-frequency spectrograms in the frequency domain will miss some time-frequency information, resulting in low recognition accuracy.
[0006] To achieve the above object, the specific technical solution of the present invention is as follows: The present invention provides an underwater acoustic target recognition method based on an adaptive multi-feature fusion model, including the following steps: An underwater acoustic target recognition method based on an adaptive multi-feature fusion model, including the following steps:
[0007] (1) Data preparation: Cut the original audio data to obtain a data set;
[0008] (2) Data preprocessing: Extract MFCC features for each audio and generate a two-dimensional time-frequency spectrogram
[0009] (3) Multi-dimensional feature extraction: including deep time-series feature extraction, deep spatial feature extraction, and deep frequency-domain feature extraction;
[0010] (4) Construction of an adaptive multi-feature fusion model:
[0011] 4.1. Input processing: Initially splice the features extracted by the three networks as the input;
[0012] 4.2: Adaptive weighting: Input the spliced feature set into the channel attention layer for adaptive weighting. The channel attention layer includes 3 modules, namely Squeeze, Excitation, and Scale. Among them, Squeeze uses global average pooling operation to process the global spatial information of each channel. Excitation normalizes each feature channel to generate weights; Scale weights the previously obtained normalized weights by multiplying them with the features of each channel.
[0013] 4.3. Output processing: Input the weighted information into the fully connected layer for underwater acoustic target recognition.
[0014] Further, in the above step (3), an LSTM network is trained based on the MFCC feature data of the underwater acoustic signal, and the output of the dropout layer is extracted as the deep time-series feature set of the underwater acoustic signal; a 1D-CNN network is trained based on the MFCC features of the underwater acoustic signal, and the output of the Fully-connected layer1 is extracted as the deep spatial feature set of the underwater acoustic signal; a 2D-CNN network is trained based on the two-dimensional time-frequency spectrogram generated from the original speech signal, and the output of the Global max pool1 layer is extracted as the deep frequency-domain feature set of the underwater acoustic signal.
[0015] Further, the constructed LSTM has a total of 4 layers, including an input layer, an LSTM layer, a dropout layer, and a fully-connected layer. The input layer is a time-series vector with a length of 1 and a dimension of 40; the number of hidden units in the LSTM layer is set to 128; a dropout layer is introduced, and the dropout rate is set to 0.2; the fully-connected layer contains 5 nodes, respectively representing the probabilities of the predicted samples being different underwater acoustic targets. Finally, the output of the dropout layer is extracted as the deep time-series feature set of the underwater acoustic signal.
[0016] Further, the above 1D-CNN network has a total of 9 layers, including 1 input layer, 2 convolutional layers, 2 pooling layers, 2 dropout layers, and 2 fully-connected layers. The input layer receives MFCC features of size 40×1, so the input size is set to 40×1; the 2 convolutional layers extract the spatial features of the underwater acoustic signal, 1 max pooling layer and 1 global max pooling layer are used for feature information compression, 2 dropout layers prevent overfitting of the model, and the 2 fully-connected layers are connected to output the probabilities of the predicted samples belonging to different underwater acoustic targets. Finally, the output of the Fully-connected layer1 is extracted as the deep spatial feature set of the underwater acoustic signal.
[0017] Further, the above 2D-CNN network has a total of 10 layers, including an input layer, three convolutional layers, three pooling layers, two dropout layers, and a fully-connected layer. The input layer receives a time-frequency spectrogram of size 224×224 with three RGB channels, so the input size is set to 224×224×3; the 3 convolutional layers extract the image features, 2 max pooling layers and one global max pooling layer are used for feature information compression, 2 dropout layers prevent overfitting of the model, and the one fully-connected layer is connected to output the probabilities of the predicted samples belonging to different underwater acoustic targets. Finally, the output of the Global max pool1 layer is extracted as the deep frequency-domain feature set of the underwater acoustic signal.
[0018] Compared with the prior art, the advantages of the present invention are as follows:
[0019] 1. The method of the present invention proposes a multi-dimensional feature extraction network structure. Starting from both the time domain and frequency domain perspectives simultaneously, in view of the characteristics that the MFCC feature data of underwater acoustic signals have both time-sequence continuity and spatial continuity, an LSTM and a 1D-CNN are respectively used to further extract the deep time-sequence features and deep spatial features of the MFCC features of underwater acoustic signals. In view of the characteristic that the two-dimensional time-frequency spectrogram generated based on audio contains rich frequency domain information, a 2D-CNN is used to further extract the deep frequency domain features of the two-dimensional time-frequency spectrogram of underwater acoustic signals. Thus, the time-frequency complementary information corresponding to the time-domain audio signal and the two-dimensional spectrogram in the frequency domain is mined, effectively improving the recognition accuracy.
[0020] 2. The method of the present invention proposes a feature weighted fusion strategy based on a channel attention mechanism. This strategy uses the attention mechanism to adaptively weight and fuse the three features extracted by the multi-dimensional feature extraction module, and adaptively weights the dependence of each channel to improve the representation ability of the network, that is, assigns more weights to effective features, solves the problem of inaccurate weights during the assignment of feature maps, provides more discriminative features for subsequent target recognition, and thus can effectively improve the recognition accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 is a flowchart of the present invention;
[0022] Figure 2 is a multi-feature adaptive fusion network proposed in an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0023] The present invention will be further described below with reference to the drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and cannot be used to limit the protection scope of the present invention.
[0024] The present invention provides an underwater acoustic target recognition method based on an adaptive multi-feature fusion model. Refer to Figure 1 It can be seen that the present invention includes the following steps:
[0025] Step 1: Input the original audio data in WAV format and cut the WAV format data.
[0026] Step 2: Preprocess each cut audio, extract MFCC features and generate a two-dimensional time-frequency spectrogram.
[0027] Step 3: Multi-dimensional feature extraction: including deep time-sequence feature extraction, deep spatial feature extraction and deep frequency domain feature extraction.
[0028] Train an LSTM network based on the MFCC feature data of underwater acoustic signals and extract the output of the dropout layer as the deep time-sequence feature set of underwater acoustic signals.
[0029] Train a 1D-CNN network based on the MFCC features of underwater acoustic signals and extract the output of the Fully-connected layer1 as the deep space feature set of underwater acoustic signals.
[0030] Train a 2D-CNN network based on the two-dimensional time-frequency spectrogram generated from the original speech signal and extract the output of the Global maxpool1 layer as the deep frequency domain feature set of underwater acoustic signals.
[0031] Step 4: Construction of an adaptive multi-feature fusion model: Refer to Figure 2 As can be seen from the entire construction process, the features extracted by the three networks in Step 3 are concatenated, and different weights are assigned to the extracted multi-dimensional features based on the feature weighting fusion strategy of the channel attention mechanism. The weighted information is input into two fully-connected layers for underwater acoustic target recognition.
[0032] Embodiment: An underwater acoustic target recognition method based on an adaptive multi-feature fusion model, the method comprising the following steps:
[0033] 1: Data preparation. Specifically:
[0034] The ShipEar dataset is used to evaluate the performance of the proposed method. This dataset was collected between 2012 and 2013 at the coastal area of Spain. The recordings were made with an autonomous acoustic digitalHyd SR-1 recorder manufactured by MarSensingLda (Faro, Portugal). This dataset contains a total of 90 audio files, with durations ranging from 15s to 10min. The audio categories include 11 types of ships and environmental noise. According to the source paper of the dataset, this dataset can be further divided into five categories: A, B, C, D, and E, where A, B, C, and D represent four major categories of ship types, and E is environmental noise. There are only 90 original audio data and the number of audio files in different categories varies greatly, which may lead to the problem of underfitting of the model. To solve this problem, the original audio data is cut into 3s segments to expand the dataset.
[0035] 2: Data preprocessing. The following preprocessing is performed on each cut audio: Extract MFCC features and generate a two-dimensional time-frequency spectrogram.
[0036] 2.1: Extract MFCC features: The extracted MFCC features have a dimension of (40, 309); the column vectors of the features are compressed by mean, and the final dimension of the MFCC features is (40, 1).
[0037] 2.2: Generate a two-dimensional time-frequency spectrogram: By performing a Fourier transform on the original audio, a two-dimensional time-frequency spectrogram is obtained. The size of the two-dimensional time-frequency spectrogram is 569×435, with three RGB channels. For the network, an overly large input image size will lead to an increase in the amount of computation, while an overly small cropped size will result in serious information loss. Cropping the image size to 224×224 is a better choice. Therefore, in the present invention, the generated two-dimensional time-frequency spectrogram is reshaped into 224×224×3.
[0038] 3: Multidimensional feature extraction. Specifically:
[0039] 3.1: Deep temporal feature extraction. The MFCC features of audio have the characteristic of temporal continuity. Therefore, in the present invention, based on the MFCC features of underwater acoustic signals, an LSTM network is used to further extract deep temporal features for recognition. The constructed LSTM has a total of 4 layers, including an input layer, an LSTM layer, a dropout layer, and a fully connected layer. The input layer is a temporal vector with a length of 1 and a dimension of 40; the number of hidden units in the LSTM layer is set to 128; in order to prevent the LSTM from overfitting on the training set, a dropout layer is introduced to reduce the computational amount during the training process of the model, and the dropout rate is set to 0.2; the fully connected layer as the output contains 5 nodes, respectively representing the probabilities of the predicted samples being different underwater acoustic targets. Finally, the output of the dropout layer is extracted as the deep temporal feature set of the underwater acoustic signal.
[0040] 3.2: Deep spatial feature extraction. The MFCC feature data has both the characteristics of spatial continuity and temporal continuity. Therefore, in the present invention, 1DCNN is simultaneously used to process the MFCC features of underwater acoustic signals, and the spatial characteristics of 1D-CNN are utilized to further extract the deep spatial features of underwater acoustic signals for recognition. The designed 1D-CNN network has a total of 9 layers, including 1 input layer, 2 convolutional layers, 2 pooling layers, 2 dropout layers, and 2 fully connected layers. The input layer receives MFCC features with a size of 40×1, so the input size is set to 40×1; the 2 convolutional layers extract the spatial features of the underwater acoustic signals, 1 max pooling layer and 1 global max pooling layer are used for feature information compression, and the 2 dropout layers prevent the model from overfitting by randomly selecting some neurons and temporarily discarding them. The 2 fully connected layers are connected to output the probabilities of the predicted samples belonging to different underwater acoustic targets. Finally, the output of Fully-connected layer1 is extracted as the deep spatial feature set of the underwater acoustic signal.
[0041] 3.3: Deep frequency domain feature extraction. The two-dimensional time-frequency spectrogram generated based on the original speech signal contains rich frequency domain information, which can be used as the basis for classification. Therefore, this invention uses 2DCNN to further extract deep frequency domain features from the two-dimensional time-frequency spectrogram. The designed 2D-CNN network has a total of 10 layers, including an input layer, three convolutional layers, three pooling layers, two dropout layers, and a fully connected layer. The input layer accepts a time-frequency spectrogram with a size of 224×224 and three RGB channels, so the input size is set to 224×224×3; the three convolutional layers extract image features, and the two max pooling layers and one global max pooling layer are used for feature information compression. The two dropout layers prevent model overfitting by randomly selecting some neurons and temporarily discarding them. Connecting a fully connected layer outputs the probabilities that the predicted samples belong to different underwater acoustic targets. Finally, the output of the Global max pool1 layer is extracted as the deep frequency domain feature set of the underwater acoustic signal.
[0042] 4: Construction of an adaptive multi-feature fusion model. To better fuse the feature information extracted in the three ways, a multi-feature fusion network structure containing only input and output is designed. Specifically:
[0043] 4.1: Input processing. The features extracted by the three networks are preliminarily concatenated as the input.
[0044] 4.2: Adaptive weighting. To enhance the mapping ability from input to output and address the problem of inaccurate weights during feature map allocation, the channel attention mechanism Squeeze-and-Excitation (SE) is introduced in this model. The implementation of the SE layer is mainly divided into three modules, namely Squeeze, Excitation, and Scale. Squeeze uses the global average pooling (Global Average Pooling, GAP) operation to compress the global spatial information of each channel, that is, to compress the two-dimensional features (W×H) of each channel. The compressed feature becomes 1×1×C. The formula for the global average pooling operation is:
[0045]
[0046] zc is the weight parameter after the compression operation; F sq (.) is the feature compression operation; u cis the c-th two-dimensional matrix in U, where U is a set of multiple local feature maps; H is the height of the feature matrix; W is the width of the feature matrix. Excitation generates a weight with a value range of (0, 1) for each feature channel through parameter w, where parameter w is learned to explicitly model the correlation between feature channels. In the specific implementation, two fully connected layers (FC-ReLU-FC-Sigmoid) are used to calculate the weight value, and the formula for the weight is:
[0047] s = F ex (z, w) = σ(g(z, w)) = σ(w2δ(w1z))
[0048] δ(w1z) represents the first fully connected operation. The dimension of w1 is C / r × C, where r is a scaling parameter that reduces the number of channels to reduce the computational complexity. In the present invention, r is taken as 4. The dimension of z is 1×1×C, so the result of w1z is 1×1×C / r, and then after passing through a ReLU layer, the output dimension remains unchanged. The result of δ(w1z) is multiplied by w2 for the second fully connected operation. The dimension of w2 is C × C / r, so the output dimension is 1×1×C; finally, through the sigmoid function, the final weight s is obtained. Scale weights the previously obtained normalized weight by multiplying it with the features of each channel.
[0049] 4.3: Output processing. The weighted information of the SE layer is input into two fully connected layers with 64 and 5 nodes respectively for underwater acoustic target recognition.
[0050] The comparison results of the method of the present invention with other methods are shown in the following table. It can be seen that the classification Acc, Recall, Precision, and F1-score of a single LSTM on the underwater acoustic dataset are all higher than those of other single sub-networks, which are 0.9022, 0.9017, 0.8926, and 0.8967 respectively. Since underwater acoustic data is a time-series signal and LSTM pays more attention to time-series features, among these three single sub-networks, LSTM has the best performance. When the features extracted by different networks are grouped and fused, the recognition accuracy is higher than that of all single networks. Among them, when the features extracted by the three networks are fused simultaneously, the recognition Acc, Recall, Precision, and F1-score all reach the highest, which are 0.9348, 0.9296, 0.9336, and 0.9315 respectively. Compared with the performance of a single LSTM, they are improved by 3.26%, 2.79%, 4.1%, and 3.48% respectively. Compared with the sub-optimal fusion feature set (2DCNN + LSTM), they are improved by 1.31%, 0.82%, 2.03%, and 1.47% respectively. It can be inferred from this that a single network structure extracts relatively one-sided feature information from underwater acoustic signals and can only extract the time-domain or frequency-domain information of underwater acoustic signals, without considering the complementary information between the two, resulting in room for improvement in recognition accuracy. By simply fusing the feature information extracted by multiple network structures, this problem can be effectively solved, so it is superior to a single network in performance and significantly improves the recognition accuracy.
[0051]
[0052] Referring to the above table, the classification Acc, Recall, Precision, and F1-score of the adaptive multi-feature fusion model proposed by the present invention on the underwater acoustic dataset reach the highest, which are 0.9492, 0.9448, 0.9443, and 0.9442 respectively. Compared with the performance before adding attention, they are improved by 1.44%, 1.52%, 1.07%, and 1.27% respectively. It can be deduced that simply fusing the features extracted by the three networks can indeed consider the complementary information in the time-domain and frequency-domain of underwater acoustic signals, thus improving the recognition accuracy. However, this simple feature fusion method does not consider that the features from different sources have different effects on the final recognition. The multi-feature adaptive fusion model proposed by the present invention adaptively weights and fuses the features extracted by 2D-CNN, 1D-CNN, and LSTM through channel attention, which can assign more weights to important features, so as to better play the role of important features. Therefore, the recognition accuracy can be significantly improved.
[0053] The above are only the preferred embodiments of the present invention and do not impose any limitations on the present invention. Any simple modifications, changes, and equivalent variations made to the above embodiments based on the technical essence of the invention still fall within the scope of protection of the technical solution of the present invention.
Claims
1. An underwater acoustic target recognition method based on an adaptive multi-feature fusion model, characterized in that: The method includes the following steps: (1) Data preparation: cutting the original audio data to obtain a data set; (2) Data preprocessing: extracting MFCC features for each audio and generating a two-dimensional time-frequency spectrogram; (3) Multi-dimensional feature extraction: including deep temporal feature extraction, deep spatial feature extraction, and deep frequency-domain feature extraction; (4) Construction of an adaptive multi-feature fusion model: 4.1 Input processing: preliminarily splicing the features extracted by the three networks as the input; 4.2 Adaptive weighting: inputting the spliced feature set into a channel attention layer for adaptive weighting. The channel attention layer includes 3 modules, namely Squeeze, Excitation, and Scale. Among them, Squeeze uses global average pooling operation to obtain the global spatial information of each channel. Excitation normalizes each feature channel to generate weights. Scale weights the previously obtained normalized weights by multiplying them with the features of each channel; 4.3 Output processing: inputting the weighted information into a fully connected layer for underwater acoustic target recognition; In step (3), an LSTM network is trained based on the MFCC feature data of the underwater acoustic signal, and the output of the dropout layer is extracted as the deep temporal feature set of the underwater acoustic signal. A 1D-CNN network is trained based on the MFCC features of the underwater acoustic signal, and the output of the Fully-connected layer1 is extracted as the deep spatial feature set of the underwater acoustic signal. A 2D-CNN network is trained based on the two-dimensional time-frequency spectrogram generated from the original speech signal, and the output of the Global max pool1 layer is extracted as the deep frequency-domain feature set of the underwater acoustic signal.
2. The underwater acoustic target recognition method based on an adaptive multi-feature fusion model according to claim 1, characterized in that: The constructed LSTM has 4 layers, including an input layer, an LSTM layer, a dropout layer, and a fully connected layer. Among them, the input layer is a temporal vector with a length of 1 and a dimension of 40. The number of hidden units in the LSTM layer is set to 128. A dropout layer is introduced, and the dropout rate is set to 0.
2. The fully connected layer contains 5 nodes, respectively representing the probabilities that the predicted sample is different underwater acoustic targets. Finally, the output of the dropout layer is extracted as the deep temporal feature set of the underwater acoustic signal.
3. The underwater acoustic target recognition method based on an adaptive multi-feature fusion model according to claim 2, wherein: The 1D-CNN network has a total of 9 layers, including 1 input layer, 2 convolutional layers, 2 pooling layers, 2 dropout layers, and 2 fully connected layers. Among them, the input size of the input layer is set to 40×1. The 2 convolutional layers extract the spatial features of the underwater acoustic signal. 1 max pooling layer and 1 global max pooling layer are used for feature information compression. The 2 dropout layers prevent the model from overfitting. The connection of the 2 fully connected layers outputs the probabilities that the predicted sample belongs to different underwater acoustic targets. Finally, the output of the Fully-connected layer1 is extracted as the deep spatial feature set of the underwater acoustic signal.
4. The underwater acoustic target recognition method based on an adaptive multi-feature fusion model according to claim 3, characterized in that: The 2D-CNN network has a total of 10 layers, including an input layer, three convolutional layers, three pooling layers, two dropout layers, and a fully connected layer. Among them, the input size of the input layer is set to 224×224×3; the 3 convolutional layers are used to extract image features, the 2 max pooling layers and one global max pooling layer are used for feature information compression, the 2 dropout layers prevent the model from overfitting, and a fully connected layer is connected to output the probabilities that the predicted samples belong to different underwater acoustic targets. Finally, the output of the Global max pool1 layer is extracted as the depth frequency domain feature set of the underwater acoustic signal.
Citation Information
Patent Citations
Many-to-many speaker conversion method based on improved STARGAN and x vector
CN110600046A
Underwater sound signal detection method and system based on deep learning
CN114636995A