An Underwater Target Recognition Method Based on EfficientNet

By using EfficientNet-based method in underwater target recognition, the data is expanded using Mel spectrum feature extraction and deep convolution generation adversarial network, the problems of low recognition accuracy and poor generalization of model in underwater target recognition are solved, and the effect of high accuracy and strong generalization of underwater target recognition is achieved.

CN115204214BActive Publication Date: 2025-06-27NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210693950.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-19
Publication Date
2025-06-27
Estimated Expiration
2042-06-19

AI Technical Summary

Technical Problem

The prior art has problems with low recognition accuracy and poor generalization of the model in underwater target recognition, especially in the case of strong reverberation and noise interference in underwater environments.

Method used

The underwater target recognition method based on EfficientNet is adopted to extract the Mel spectral image data set of the underwater target through Mel spectral features, and the training set is expanded using the deep convolution generation adversarial network GAN-S. Then the expanded data is input into the deep convolution neural network ConvNet-S for supervised training to obtain the recognition results.

Benefits of technology

It significantly improves the accuracy of underwater target recognition, enhances the generalization ability of the model, and can effectively identify underwater targets in the absence of samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115204214B_ABST
    Figure CN115204214B_ABST
Patent Text Reader

Abstract

The present invention discloses an underwater target recognition method based on EfficientNet. First, the Mel spectrum feature extraction method is used for the active sonar echo signals of underwater targets to obtain the Mel spectrum images of each group of echo signals; the Mel spectrum image dataset is divided into a training set, a validation set and a test set, and preprocessing is carried out; the deep convolutional generative adversarial network GAN-S is used to expand the training set; the expanded training samples are input into the deep convolutional neural network ConvNet-S for supervised training to obtain the parameters of each layer of the convolutional neural network; the Mel spectrum images in the test set are input into the trained deep convolutional neural network ConvNet-S to obtain the recognition results of each Mel spectrum image, and the recognition accuracy of the test set of various targets is statistically calculated to obtain the final recognition result. The present invention can effectively extract the category features of underwater targets and significantly improve the recognition accuracy compared with other existing algorithms.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of pattern recognition, and particularly relates to an underwater target recognition method. Background Art

[0002] Currently, traditional underwater target classification and recognition methods usually manually extract several features from sonar echoes and use them to train a classifier to complete target recognition. Recognition methods based on deep learning can directly extract features from the original signal, compress the feature vector, and fit the target mapping, learning multi-level category features, thereby avoiding feature loss in the manual extraction process and effectively improving the generalization ability. Most of the research on underwater target classification and recognition based on deep learning uses passive sonar data or synthetic aperture sonar (SAS) data, while the research on recognition based on active sonar data is relatively lacking. At the same time, due to the complex underwater environment and the rapid development of mechanical noise reduction technology, the concealment of underwater targets is getting higher and higher. In practice, there are many difficulties in underwater target recognition based on passive sonar. Therefore, the classification research on active sonar data is of great significance for underwater target recognition.

[0003] In the real underwater environment, sonar sensors are affected by random noise, water body scattering, and the sonar reflection characteristics of target materials, resulting in strong reverberation and noise interference mixed in the active sonar echo data of the target. At the same time, it is very difficult to obtain the echo data of underwater targets, resulting in serious shortage of observation data. Summary of the Invention

[0004] In order to overcome the deficiencies of the prior art, the present invention provides an underwater target recognition method based on EfficientNet. First, use the Mel spectrum feature extraction method for the active sonar echo signal of the underwater target to obtain the Mel spectrum image of each group of echo signals; divide the Mel spectrum image dataset into a training set, a validation set, and a test set, and perform preprocessing; use the deep convolutional generative adversarial network GAN-S to expand the training set; input the expanded training samples into the deep convolutional neural network ConvNet-S for supervised training to obtain the parameters of each layer of the convolutional neural network; input the Mel spectrum images in the test set into the trained deep convolutional neural network ConvNet-S to obtain the recognition results of each Mel spectrum image, and count the recognition accuracy of the test set of each type of target to obtain the final recognition result. The present invention can effectively extract the category features of underwater targets and significantly improve the recognition accuracy compared with other existing algorithms.

[0005] The technical solution adopted by the present invention to solve its technical problems includes the following steps:

[0006] Step 1: Use the Mel spectrum feature extraction method for the active sonar echo signal of the underwater target to obtain the Mel spectrum image of each group of echo signals;

[0007] Step 2: Construct a Mel-spectrum image dataset of underwater targets from the Mel-spectrum images obtained in Step 1, divide the Mel-spectrum image dataset into a training set, a validation set, and a test set, and perform preprocessing;

[0008] Step 3: Use the deep convolutional generative adversarial network GAN-S to augment the training set;

[0009] Step 4: Input the augmented training samples into the deep convolutional neural network ConvNet-S for supervised training to obtain the parameters of each layer of the convolutional neural network;

[0010] Step 5: Input the Mel-spectrum images in the test set into the trained deep convolutional neural network ConvNet-S to obtain the recognition results of each Mel-spectrum image, and count the recognition accuracy of the test set of various targets to obtain the final recognition result.

[0011] Furthermore, the Mel-spectrum feature extraction method is specifically as follows: Frame and pre-emphasize the active sonar echo signal of the underwater target, then perform FFT transformation on each frame of the signal to obtain the signal spectrum, and filter the signal spectrum according to the formula Use a Mel filter bank for filtering, calculate the energy of each filter and take the logarithm to obtain the Mel-spectrum image, where f represents the signal spectrum.

[0012] Furthermore, the method for constructing the Mel-spectrum image dataset of underwater targets is specifically as follows: Randomly select 15% from the Mel-spectrum images of each type of target as the test set, and the remaining part as the training and validation sets. The preprocessing includes data augmentation, size scaling, cropping, grayscaling, and normalization.

[0013] Furthermore, the deep convolutional generative adversarial network GAN-S includes a generative model G used to capture the details of the data characteristic distribution and a discriminative model D used to estimate the sample data from the real training data x;

[0014] Set the task of the generative model G to maximize the probability that the discriminative model D makes a wrong judgment, and at the same time set the task of the discriminative model D to distinguish the data from the sample and the data from the generative model G; Define the distribution of the data generated by the generative model G as P g , the prior variable of the input noise is P z (z), use G(z; θ g ) to represent the mapping of the data space, where G(·) is a multi-layer perceptron containing parameters θ g ; Then define D(x; θ d ) as also a multi-layer perceptron to output a single label scalar; D(x) represents that x comes from the real data distribution P data(x) rather than Pg The probability; by training the discriminative model D to maximize the probability of correctly discriminating the labels, and training the generative model G to minimize log(1 - D(G(z))), the training process of D and G is a two-player minimax game problem regarding the value function V(G; D):

[0015]

[0016] Where:

[0017] When training D and G such that D cannot distinguish the data generated by G from the real data, there is P g (x) = P data (x), and at this time D(x) = 0.5, that is, the training process reaches the optimum;

[0018] The generative model G first generates a 100-dimensional random noise, maps it to a matrix of size 147456 through a fully connected layer, and converts it into a feature map of 12×12×1024; using the grouped symmetric padding method, the size of the feature map is doubled through the transposed convolution operation of 5 transposed convolution layers TransConv, and then the feature map is reduced, and finally a feature map of 384×384×3 is output; the discriminative model D realizes feature extraction through convolution operations on the feature map of size 384×384×3 output by the generator, and finally outputs after being mapped by the fully connected layer; batch normalization layers BN and activation functions are introduced in both the generator and the discriminator to normalize the feature values; a Dropout layer is added to the discriminator to suppress overfitting;

[0019] Furthermore, the specific grouped symmetric padding method is as follows: the feature map with 1024 channels is divided into 256 groups with an equal number in the channel order, symmetric padding is performed on the 4-channel feature map of each group, and the size of the feature map obtained after padding is (Channel, Width, Hight) = (1024, 13, 13). Then, through the transposed convolution operation with a convolution kernel size of 2×2 and a stride of 2, the length and width of the matrix are expanded, and the dimension of the matrix is reduced.

[0020] Furthermore, the basic convolutional block of the deep convolutional neural network ConvNet-S is built based on the EfficientNet network model. After stacking and designing the basic convolutional block and the attention mechanism module, the classification model of the present invention is obtained, and its structure is shown in Table 1:

[0021] Table 1 Structure Table of Deep Convolutional Neural Network ConvNet-S

[0022]

[0023] Further, the specific steps for statistically calculating the recognition accuracy of the test set for various types of targets are as follows: Statistically calculate TP: predicted as positive, actually positive; TN: predicted as negative, actually negative; FP: predicted as positive, actually negative; FN: predicted as negative, actually positive. After obtaining the above TP, TN, FP, and FN, through the formula Calculate the recognition accuracy of the test set.

[0024] The beneficial effects of the present invention are as follows:

[0025] The method of the present invention has high recognition accuracy and strong generalization ability. It can be applied to the underwater environment with reverberation and noise interference. In the case of a lack of sample numbers, the data set can be effectively expanded through a deep convolutional generative adversarial network, and then multi-dimensional target category features can be extracted through a classification model to accurately identify underwater targets. It has been verified that the present invention has achieved better results on the measured data compared with the existing methods. It can effectively solve the problems of low recognition accuracy and poor model generalization ability in traditional methods, has a wide range of application prospects, and can be directly put into use. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 is a flow chart of the present invention.

[0027] Figure 2 (a), (b), (c), and (d) are examples of the time domain and Mel spectrograms of the echoes of four types of targets in sequence.

[0028] Figure 3 The structural diagram of the basic convolutional block proposed by the present invention.

[0029] Figure 4 The structural diagram of the deep convolutional generative adversarial network of the present invention.

[0030] Figure 5 The schematic diagram of the grouping padding process of the present invention.

[0031] Figure 6 The structural diagram of the deep convolutional network model (ConvNet-S) of the present invention.

[0032] Figure 7 The experimental results obtained by each algorithm of the embodiments of the present invention on the established target echo data set of the pool experiment. (a) is the curve of the validation set loss with the epoch iteration, and (b) is the curve of the validation set accuracy with the epoch iteration. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0033] The present invention will be further described below in conjunction with the drawings and embodiments.

[0034] In view of the lack of echo data, the present invention utilizes deep convolutional generative adversarial networks to augment the dataset, effectively suppressing the overfitting phenomenon. In practice, the active sonar echoes of underwater targets are mixed with strong reverberation and noise interference. The present invention can effectively extract the category features of underwater targets and significantly improve the recognition accuracy compared with other existing algorithms.

[0035] As Figure 1 shown, an underwater target recognition method based on EfficientNet includes the following steps:

[0036] Step 1: Use the Mel spectrum feature extraction method for the active sonar echo signals of underwater targets to obtain the Mel spectrum images of each group of echo signals;

[0037] Step 2: Construct a Mel spectrum image dataset of underwater targets from the Mel spectrum images obtained in Step 1, divide the Mel spectrum image dataset into a training set, a validation set, and a test set, and perform preprocessing;

[0038] Step 3: Use the deep convolutional generative adversarial network GAN-S to augment the training set;

[0039] Step 4: Input the augmented training samples into the deep convolutional neural network ConvNet-S for supervised training to obtain the parameters of each layer of the convolutional neural network;

[0040] Step 5: Input the Mel spectrum images in the test set into the trained deep convolutional neural network ConvNet-S to obtain the recognition results of each Mel spectrum image, and statistically calculate the recognition accuracy of the test set for each type of target to obtain the final recognition result.

[0041] Further, the Mel spectrum feature extraction method is specifically as follows: Frame and pre-emphasize the active sonar echo signals of underwater targets, then perform FFT transformation on each frame of the signal to obtain the signal spectrum, and filter the signal spectrum according to the formula using a Mel filter bank, calculate the energy of each filter and take the logarithm to obtain the Mel spectrum image, where f represents the signal spectrum.

[0042] Further, the method for constructing the Mel spectrum image dataset of underwater targets is specifically as follows: Randomly select 15% of the Mel spectrum images of each type of target as the test set in turn, and the remaining part as the training and validation set. The preprocessing includes data augmentation, size scaling, cropping, grayscale conversion, and normalization.

[0043] Further, the deep convolutional generative adversarial network GAN-S includes a generative model G used to capture the detailed distribution characteristics of the data and a discriminant model D used to estimate the sample data from the real training data x.

[0044] The task of the generative model G is to maximize the probability that the discriminative model D makes an incorrect judgment. At the same time, the task of the discriminative model D is to distinguish data from samples and data from the generative model G. Define the distribution of the data generated by the generative model G as P g , and the prior variable of the input noise is P z (z). Use G(z; θ g ) to represent the mapping of the data space, where G(·) is a multi-layer perceptron with parameters θ g . Then define D(x; θ d ) as also a multi-layer perceptron, used to output a single label scalar; D(x) represents the probability that x comes from the true data distribution P data(x) rather than P g . By training the discriminative model D to maximize the probability of the correct discriminative label, and training the generative model G to minimize log(1 - D(G(z))), the training process of D and G is a minimax two-player game problem regarding the value function V(G; D):

[0045]

[0046] where:

[0047] When training D and G so that D cannot distinguish the data generated by G from the real data, there is P g (x) = P data (x). At this time, D(x) = 0.5, that is, the training process reaches the best;

[0048] The generative model G first generates 100-dimensional random noise, maps it to a matrix of size 147456 through a fully connected layer, and converts it into a feature map of 12×12×1024; using the grouped symmetric padding method, the size of the feature map is doubled through the reverse convolution operation of 5 transposed convolution layers TransConv, and then the feature map is reduced, and finally a feature map of 384×384×3 is output; the discriminative model D realizes feature extraction through convolution operations on the feature map of size 384×384×3 output by the generator, and finally outputs after being mapped by the fully connected layer; Batch Normalization layers BN and activation functions are introduced in both the generator and the discriminator to normalize the feature values; a Dropout layer is added to the discriminator to suppress the overfitting phenomenon;

[0049] Further, the specific grouping symmetric padding method is as follows: the 1024-channel feature map is divided into 256 groups with equal numbers in the channel order, and symmetric padding is performed on the 4-channel feature maps of each group. The size of the feature map obtained after padding is (Channel, Width, Hight) = (1024, 13, 13). Then, a transposed convolution operation with a convolution kernel size of 2×2 and a stride of 2 is used to expand the length and width of the matrix, and the dimension of the matrix is reduced.

[0050] Further, the basic convolution block of the deep convolutional neural network ConvNet-S is built based on the EfficientNet network model. After stacking and designing the basic convolution block and the attention mechanism module, the classification model of the present invention is obtained, and its structure is shown in Table 1;

[0051] Further, the specific steps for statistically calculating the recognition accuracy of the test set for various types of targets are as follows: statistically calculate TP: predicted as positive, actually positive; TN: predicted as negative, actually negative; FP: predicted as positive, actually negative; FN: predicted as negative, actually positive. After obtaining the above TP, TN, FP, and FN, use the formula to calculate the recognition accuracy of the test set. Specific embodiments:

[0053] The experimental environment of this embodiment is a desktop workstation with an AMD Ryzen 9 3990X CPU, a 3070 GPU, and 64G of memory, and the operating system is Windows10. The software tools are Spyder (python 3.7), Pytorch1.10.0, and CUDA 11.3. During the feature extraction process, the parameters of each layer of the neural network are configured on Spyder, and after configuration, the files such as network training and network forward propagation written in Python are compiled into Python executable files to implement network training. During the experiment, CUDA is used for GPU parallel acceleration operations. The specific implementation method is as follows:

[0054] 1. The time-domain echo data of underwater targets is the measured echo data set in an anechoic tank. There are a total of four different underwater target models in this data set. The four types of targets are successively lowered into the water. Starting from the 0-degree rotation attitude with the target's beam aspect facing the receiving array directly, keeping the positions of the transmitting transducer and the receiving array unchanged, rotating the target counterclockwise around the target's longitudinal axis, with a step of 1 degree, and using the receiving array to receive the target signals at each rotation angle. 360 angle echoes of each type of target can be obtained, and a total of 1440 angle echo signals for the four types of targets.

[0055] 2. Use the Mel spectrum feature extraction method to obtain a total of 360*4 time-frequency images for the four types from the above-mentioned measured target active sonar echo data. The specific operation details are as follows:

[0056] (1) First, pre - emphasize the echo to compensate for the loss of high - frequency components and enhance the high - frequency components. By setting the sample length to 2 * 10 4 Each angular signal passes through a first - order FIR high - pass digital filter as shown in the following formula, which divides the signal into shorter frames. In each frame, it can be regarded as a steady - state signal and can be processed by the methods for processing steady - state signals. To enable a smoother transition of parameters between adjacent frames, there is partial overlap between adjacent frames.

[0057] H(z) = 1 - αz -1

[0058] (2) The purpose of applying the window function is to reduce leakage in the frequency domain. Multiply each frame of the audio signal by the Hanning window ω(n).

[0059]

[0060] (3) Perform an FFT transformation with a number of points (NFFT) of 512 on each windowed frame signal X i (m) to obtain its spectrum, and then calculate the energy spectrum E(i,k) through the following formula, where i represents the i - th frame signal after frame division.

[0061] E(i,k) = |X(i,k)| 2 = |FFT[X i (m)]| 2

[0062] (4) Filter the power spectrum using a Mel filter bank, calculate the energy in each filter, and take the logarithm to obtain the Mel spectrum S(i,m). As shown in the following formula:

[0063]

[0064] Adopt triangular band - pass filters with a frequency response of H m (k), and the number of filters M is taken as 26. As Figure 2 shown, it is an example of the time domain and Mel spectrogram of four types of target echoes.

[0065] 3. To be consistent with the small - sample situation in the real classification task, randomly select 136 Mel spectrogram images for each type of target, thus constructing a Mel spectrogram dataset based on the experimental data. Randomly divide the obtained 544 images into three subsets: there are 380 training - set images (about 70%), 82 validation - set images (about 15%), and the remaining images are used as the test set.

[0066] 4. Preprocess the training and validation sample sets and the test sample set. The processing methods include: (1) Data augmentation: Sharpen the sub-time-frequency maps, and then adjust the brightness and saturation; (2) Size scaling: Use the OpenCV vision library to perform linear interpolation on each sub-time-frequency map to achieve the scaling of the sub-time-frequency maps, so that all sub-time-frequency maps have the same size and the length is equal to the width; (3) Cropping: Crop the scaled sub-time-frequency maps so that their sizes match the size of the input images of the convolutional neural network; (4) Grayscale conversion and normalization: Use the transforms method of the torchvision library to perform grayscale conversion and normalization operations on the sample sets.

[0067] 5. The generative adversarial network proposed in the present invention is an unsupervised model. The model framework includes a generative model (Generator Model, G) used to capture the detailed distribution characteristics of the data and a discriminative model (Discriminative Model, D) used to estimate that the sample data comes from the real training data x rather than from the generative model. The task of the generative model G is set to maximize the probability of the "judgment error" of the discriminative model D, and at the same time, the task of the discriminative model D is set to accurately distinguish the data from the sample and the data from the generative model G. Define the distribution of the data generated by the generator as P g , and the prior variable of the input noise is P z (z). Use G(z; θ g ) to represent the mapping in the data space, where G is a multi-layer perceptron containing the parameter θ g . Then define D(x; θ d ) as also a multi-layer perceptron, which is used to output a single label scalar. D(x) represents the probability that x comes from the real data distribution P data (x) rather than P g . By training D to maximize the probability of the correct discriminative label and training G to minimize log(1 - D(G(z))), the training processes of D and G are a minimax two-player game problem regarding the value function V(G; D):

[0068]

[0069] where:

[0070] When training D and G so that D cannot distinguish the data generated by G from the real data, there is P g (x) = P data (x). At this time, D(x) = 0.5, that is, the training process reaches the best.

[0071] The generator G designed in the present invention first generates a 100-dimensional random noise, maps it into a matrix of size 147456 through a fully connected layer, and converts it into a matrix of 12×12×1024; then doubles the size of the feature map through the transposed convolution operations of 5 transposed convolution layers (TransConv), and finally outputs a generated image of 384×384×3. The discriminator D extracts features by performing convolution operations (Conv) on the 384×384×3-sized feature map output by the generator, and finally outputs after being mapped by a fully connected layer. Batch normalization layers (BN) and activation functions are introduced in both the generator and the discriminator to normalize the feature values, accelerate the convergence of the network, and improve the learning ability of the network. In addition, a Dropout layer is added to the discriminator to suppress the overfitting phenomenon, thereby improving the generalization ability of the model. The structure of the generative adversarial network (GAN-S) built in the present invention is as Figure 4 .

[0072] Compared with conventional optical images, the texture features of Mel spectrogram images are denser and the contour features are not obvious. The requirement for the receptive field is relatively low during the transposed convolution process of the generator. Therefore, the present invention proposes to use a 2*2 (C2) even-sized convolution kernel to replace the conventional 3*3 (C3) convolution kernel to reduce the number of parameters. According to the characteristics of convolution, an even-sized convolution kernel will cause an asymmetric receptive field (RFs), resulting in pixel offsets in the final feature layer. This positional offset accumulates during multiple convolutional superpositions, seriously eroding the spatial information. The present invention proposes to use a grouped symmetric padding method (Group-Padding), divide the 1024-channel feature map into 256 groups with an equal number in the channel order, and perform symmetric padding on the 4-channel feature maps of each group, as Figure 5 shown. The size of the feature map obtained after padding is (Channel, Width, Hight) = (1024, 13, 13). Then, through a transposed convolution operation with a convolution kernel size of 2×2 and a stride of 2, the length and width of the matrix are expanded, and the dimension of the matrix is reduced. This process is called (C2-GP).

[0073] 6. The basic convolutional block of the present invention is built based on EfficientNet, which has achieved excellent results in the field of optical image recognition. The process of building this convolutional block is as follows: A residual structure is adopted to add an identity mapping as a branch beside the regular main path to prevent the network from degrading; the standard 3*3 convolution is replaced by a combination of convolutional kernels of various sizes; the hardswish activation function is used to replace the conventional Relu activation function to reduce the number of model parameters and computational consumption; BN is used for regularization after each convolution operation; the self-attention mechanism is adopted to adaptively weight the features in the channel dimension to improve the expressive ability of the features; finally, the dropout strategy is used to alleviate model overfitting and complete the output after superimposing the main path and the branch. The structure of the basic convolutional block built by the present invention is as Figure 3 shown.

[0074] Figure 6 is the structure diagram of the classification model of the present invention. Block 1 consists of 1 convolutional layer with a convolutional kernel size of 3*3 pixels and 28 convolutional kernels; Block 2 consists of 1 basic convolutional block with a convolutional kernel size of 7*7 pixels, a stride of 1 pixel, and 28 convolutional kernels; Block 3 consists of 1 basic convolutional block with a convolutional kernel size of 5*5 pixels, a stride of 2 pixels, and 112 convolutional kernels; Block 4 consists of 3 basic convolutional blocks with a convolutional kernel size of 3*3 pixels, a stride of 2 pixels, and 448 convolutional kernels in each convolutional block; Block 5 consists of 3 basic convolutional blocks including the self-attention mechanism, with a convolutional kernel size of 3*3 pixels, a stride of 2 pixels, and 112 convolutional kernels in each convolutional block; Block 6 consists of 4 basic convolutional blocks including the self-attention mechanism, with a convolutional kernel size of 3*3 pixels, a stride of 1 pixel, and 128 convolutional kernels in each convolutional block; Block 7 consists of 6 basic convolutional blocks including the self-attention mechanism, with a convolutional kernel size of 3*3 pixels, a stride of 1 pixel, and 256 convolutional kernels in each convolutional block; Block 8 consists of a convolutional layer, a pooling layer, and a fully connected layer. The convolutional kernel size in the convolutional layer is 1*1, and the average pooling method is adopted. The output dimension of the fully connected layer is 4. The SGD optimization algorithm is used in the training experiment, the initial learning rate is set to 0.01, and the training is carried out for 50 epochs using the cosine learning rate decay strategy, and the training batch size is 8.

[0075] 7. Set the batch size of the test set input to 16, input the Mel spectrogram images of various targets in the test set into the trained deep convolutional neural network (ConvNet-S), obtain the recognition results of each sub-graph, draw the confusion matrix, and simultaneously count the recognition accuracy of the test set of various targets to obtain the final recognition result.

[0076] The above are all the operation rules of the algorithm of the present invention. In order to reflect the effectiveness of the algorithm of the present invention, four groups of experiments as shown in Table 2 were carried out under the same data set and training environment. Experiment 1 was to use only the ConvNet-S network model proposed by the present invention; Experiment 2 was to use GAN-S to augment the training set and then use ConvNet-S for recognition experiments; Experiment 3 was to use only the existing IAFNet network model; Experiment 4 was to use only the existing Efficientnet-V2S network model. Combining the experimental results in Table 2, it can be seen that the recognition accuracy of the test set reached 92.5% after the combination of ConvNet-S and GAN-S, which is much higher than other existing ones. At the same time, compared with using ConvNet-S alone, there was also an effective improvement of 2.5%. Observing the validation set loss and accuracy curves of Experiment 2, it can be seen that the phenomenon of the decrease in the validation set accuracy and the increase in the validation set loss no longer appears after 30 epochs of iteration, and the training process is more stable. This shows that the overfitting of the network is effectively suppressed after augmenting the data by GAN-S. The experimental results show that regular augmentation of the time-frequency map data set by GAN-S is an effective and practical method. It can be predicted that continuing to augment the data set can still reduce the network overfitting to a certain extent, but it also requires the cost of increased training time. Comparing Experiment 1 with Experiments 3 and 4, it can be seen that the proposed ConvNet-S classification model in this paper achieved a significantly higher recognition accuracy of the test set on the experimental data set. Due to the decrease in the depth of the convolutional model, the training time of ConvNet-S is significantly lower than that of Efficientnet-V2S. The structure of IAFNet is more lightweight. Although the training speed is fast, its test set accuracy is only 73.8%, which is much lower than the network in this paper.

[0077] Table 2 Recognition accuracy and training time consumption of each algorithm for the test set

[0078]

[0079] Figure 7 The experimental results obtained by each algorithm in the embodiments of the present invention on the established pool experiment target echo data set, (a) is the curve of the validation set loss with the epoch iteration, and (b) is the curve of the validation set accuracy with the epoch iteration.

Claims

1. An underwater target recognition method based on EfficientNet, characterized in that, It includes the following steps: Step 1: Use the Mel-spectrum feature extraction method for the active sonar echo signals of underwater targets to obtain the Mel-spectrum images of each group of echo signals; Step 2: Construct a Mel-spectrum image dataset of underwater targets from the Mel-spectrum images obtained in Step 1, divide the Mel-spectrum image dataset into a training set, a validation set, and a test set, and perform preprocessing; Step 3: Use the deep convolutional generative adversarial network GAN-S to augment the training set; The deep convolutional generative adversarial network GAN-S includes a generative model G used to capture the details of the data characteristic distribution and a discriminative model D used to estimate the sample data from the real training data x; The task of the generative model G is to maximize the probability that the discriminative model D makes an incorrect judgment. At the same time, the task of the discriminative model D is to distinguish data from samples and data from the generative model G. Define the distribution of the data generated by the generative model G as P g , and the prior variable of the input noise is P z (z). Use G(z; θ g ) to represent the mapping of the data space, where G(g) is a multi-layer perceptron with parameters θ g . Then define D(x; θ d ) as also a multi-layer perceptron, used to output a single label scalar; D(x) represents the probability that x comes from the true data distribution P data(x) rather than P g . By training the discriminative model D to maximize the probability of the correct discriminative label, and training the generative model G to minimize log(1 - D(G(z))), the training process of D and G is a minimax two-player game problem regarding the value function V(G; D): Wherein: When training D and G so that D cannot distinguish between the data generated by G and the real data, there is P g (x) = P data (x), at this time D(x) = 0.5, that is, the training process reaches the optimum; The generative model G first generates a 100-dimensional random noise, maps it to a matrix of size 147456 through a fully connected layer, and converts it into a feature map of 12×12×1024, where 1024 is the number of channels; using the grouped symmetric padding method, the feature map size is doubled through the transposed convolution operations of 5 transposed convolution layers TransConv, and then the feature map is reduced, and finally a feature map of 384×384×3 is output; the discriminative model D realizes feature extraction through convolution operations on the 384×384×3-sized feature map output by the generator, and finally outputs after being mapped by a fully connected layer; batch normalization layers BN and activation functions are introduced in both the generator and the discriminator to normalize the feature values; a Dropout layer is added to the discriminator to suppress the overfitting phenomenon; Step 4: Input the augmented training samples into the deep convolutional neural network ConvNet-S for supervised training to obtain the parameters of each layer of the convolutional neural network; Step 5: Input the Mel-spectrum images in the test set into the trained deep convolutional neural network ConvNet-S to obtain the recognition results of each Mel-spectrum image, and count the recognition accuracy of the test set of various targets to obtain the final recognition result.

2. The underwater target recognition method based on EfficientNet according to claim 1, wherein, The specific Mel spectrum feature extraction method is as follows: frame and pre-emphasize the active sonar echo signal of the underwater target, then perform FFT transformation on each frame of the signal to obtain the signal spectrum, and filter the signal spectrum according to the formula Use a Mel filter bank for filtering, calculate the energy of each filter and take the logarithm to obtain the Mel spectrum image, where f represents the signal spectrum.

3. The underwater target recognition method based on EfficientNet according to claim 1, characterized in that, The method for constructing the Mel-spectrum image dataset of underwater targets is specifically as follows: Randomly select 15% from the Mel-spectrum images of each type of target as the test set, and the remaining part as the training and validation set. The preprocessing includes data augmentation, size scaling, cropping, grayscale conversion, and normalization.

4. The underwater target recognition method based on EfficientNet according to claim 1, wherein, The grouped symmetric padding method is specifically as follows: Divide the 1024-channel feature map into 256 groups with an equal number in the channel order, perform symmetric padding on the 4-channel feature map of each group, and the size of the feature map obtained after padding is (Channel, Width, Hight)=(1024, 13, 13). Then, perform transposed convolution operations with a convolution kernel size of 2×2 and a stride of 2 to expand the length and width of the matrix and reduce the dimension of the matrix.

5. The underwater target recognition method based on EfficientNet according to claim 1, characterized in that, The basic convolutional block of the deep convolutional neural network ConvNet-S is built based on the EfficientNet network model. After stacking and designing the basic convolutional block and the attention mechanism module, a classification model is obtained, including: The first convolutional block: Conv3*3, with a convolutional stride of 2, an output channel number of 28, and a stacking number of 1; Second Convolution Block: Basic-Conv, k7*7, convolution stride is 1, number of output channels is 28, stacking times is 1; Third Convolution Block: Basic-Conv, k5*5, convolution stride is 2, number of output channels is 112, stacking times is 1; Fourth Convolution Block: Basic-Conv, k3*3, convolution stride is 2, number of output channels is 448, stacking times is 3; Fifth Convolution Block: Basic-Conv, k3*3, SE0.25, convolution stride is 2, number of output channels is 112, stacking times is 3; Sixth Convolution Block: Basic-Conv, k3*3, SE0.25, convolution stride is 1, number of output channels is 128, stacking times is 4; Seventh Convolution Block: Basic-Conv, k3*3, SE0.25, convolution stride is 2, number of output channels is 256, stacking times is 6; Eighth Convolution Block: Conv1*1 & Pooling & FC, number of output channels is 4, stacking times is 1.

6. The underwater target recognition method based on EfficientNet according to claim 1, characterized in that, The specific steps for statistically calculating the recognition accuracy of the test set for various types of targets are as follows: Statistically calculate TP: predicted as positive, actually positive; TN: predicted as negative, actually negative; FP: predicted as positive, actually negative; FN: predicted as negative, actually positive; After obtaining the above TP, TN, FP, and FN, the recognition accuracy of the test set is calculated through the formula ​

Citation Information

Patent Citations

  • Underwater acoustic data set expansion method based on wavelet images

    CN112560603A

  • Deep learning laser underwater target identification instrument for improving target clustering characteristics

    CN112926382A