An SE_ResNet_17 Method Suitable for Underwater Acoustic Target Recognition in a Dynamic Environment

Through the SE_ResNet_17 model, the residual network is used to eliminate gradient vanishing and channel weighting, which solves the accuracy and robustness of water acoustic target recognition in dynamic environments, and achieves more efficient feature extraction and recognition effects.

CN114298090BActive Publication Date: 2025-07-18NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111512006.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-07
Publication Date
2025-07-18
Estimated Expiration
2041-12-07

AI Technical Summary

Technical Problem

In a dynamic environment, the accuracy of water acoustic target recognition is difficult to stabilize. The existing deep neural network has gradient disappearance problems and the recognition accuracy is reduced. The convolution kernel characteristics lead to information redundancy, making it difficult to adapt to different data sets and signal-to-noise ratios.

Method used

Using the SE_ResNet_17 model, the gradient disappearance is eliminated through the residual network, and the channel weighting method is used to eliminate redundancy, improve the adaptive ability of feature extraction, and enhance the robustness of the network.

Benefits of technology

Effectively removes the redundancy of water acoustic time domain features, improves the accuracy of water acoustic target recognition and the system's adaptability to different data, and enhances the robustness of recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114298090B_ABST
    Figure CN114298090B_ABST
Patent Text Reader

Abstract

The present invention proposes an SE_ResNet_17 method applicable to underwater acoustic target recognition in a dynamic environment. The SE_ResNet_17 model first uses a residual network to extract the features of underwater acoustic signals. The residual network eliminates the phenomenon of gradient disappearance during the optimization process. By analyzing the output of the residual network, the channels where the feature data effective for the recognition task are located are further found, and the channels are weighted to improve the correct rate of underwater acoustic recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of underwater acoustic recognition, and particularly relates to a SE_ResNet_17 method suitable for underwater acoustic target recognition in a dynamic environment. Background Technique

[0002] Underwater target recognition is one of the most important functional requirements in passive sonar systems. However, due to the fact that the underwater noise environment varies with different sea areas, and even in the same sea area, different sound field environments are formed with the changes in time, water temperature, and underwater depth, it has always been very difficult to obtain stable results in the accuracy of underwater acoustic target recognition. Unstable recognition results are difficult to apply to important underwater recognition tasks. Currently, the main method is to rely on the unique auditory perception system of humans to train sonar operators to manually complete some important underwater recognition tasks. Deep learning mimics the neural network mechanism of humans and establishes a relatively perfect adaptive network structure, which can identify features with obvious invariance and differences in complex and changing environments. In recent years, it has been widely used in the field of underwater acoustic recognition.

[0003] In underwater target recognition methods, the abstract features represented by the original data in different convolutional kernels are different. Therefore, increasing the number of convolutional layers can extract more features. As a result, many deep neural networks have emerged for underwater target recognition. However, during the process of continuously deepening the network layers, there will be a phenomenon that the recognition rates of the training set and the test set both decrease as the number of layers deepens. This phenomenon is mainly due to the disappearance of gradients during the backpropagation process as the network layers deepen. To address the problem of gradient disappearance in deep networks, the ResNet network uses an optimized stacked residual function to replace the parameters in the optimized mapping relationship, effectively solving the problem of gradient disappearance.

[0004] The ResNet model has achieved good recognition results when applied to underwater target recognition. However, for different data sets or data with different signal-to-noise ratios in the same data set, the recognition accuracy of this network significantly decreases. The patent "An Underwater Acoustic Target Recognition Method Based on an Adversarial Residual Network" proposes an underwater acoustic target recognition method based on a residual network. In order to further improve the recognition accuracy of the network, the characteristics of convolutional kernels in the convolutional network are studied, and it is found that different original waveform signals are extracted on different convolutional kernel channels. However, the waveforms of many bands in the time-domain underwater acoustic signals are similar, resulting in similar features appearing on many channels, causing information redundancy and lacking the ability to adaptively distinguish different input data. In view of the sparsity and non-stationarity of underwater acoustic signals, a SE_ResNet_17 model based on the ResNet model is built. The model contains 17 mapping operations and is called the SE_ResNet_17 model. This model uses a channel weighting method to eliminate channel redundancy, increase the correctness of feature extraction, and uses an adaptive weighting method to improve the adaptability of the system to different data, thereby improving the robustness of the system. Summary of the Invention

[0005] In order to improve the recognition accuracy of underwater acoustic signals, the present invention proposes a SE_ResNet_17 method applicable to underwater acoustic target recognition in a dynamic environment.

[0006] The technical solution of the present invention is: a SE_ResNet_17 method applicable to underwater acoustic target recognition in a dynamic environment, comprising the following steps:

[0007] Step 1: Collect two groups of underwater acoustic target data sets of different types in different acoustic field environments, represented by data set 1 and data set 2 respectively; after collection, perform the same preprocessing on the two data sets and conduct sample normalization;

[0008] Step 2: Send the samples into the SE_ResNet_17 model to train a stable recognition model.

[0009] A further technical solution of the present invention is: the preprocessing of the data set in step 1 includes the following sub-steps:

[0010] Step 1.1: Frame the time-domain signals of the data set, with n feature points as one frame and no overlap between frames;

[0011] Step 1.2: Perform windowing processing on the framed time-domain signals in step 1.1 to obtain multiple groups of samples of the same length, analyze all sample feature points, and if the maximum feature point in the sample is less than 0.1, eliminate the small-value frame samples;

[0012] Step 1.3: Normalize all samples after eliminating the small-value frame samples.

[0013] A further technical solution of the present invention is: in step 2, it includes the following sub-steps:

[0014] Step 2.1: Select an appropriate convolution size for the time-domain waveform of the input signal, and perform convolution and max-pooling methods on the effective target features to obtain feature data;

[0015] Step 2.2: The obtained data is used as the original data of the residual network, and after passing through two residual modules, each residual module contains two convolution operations to obtain feature data of multiple channels;

[0016] Step 2.3: Process the data of each channel output by the residual module to obtain the weights of the feature data in the channel in the recognition task, and then use the weights to perform a weighting operation on the output data of the residual network.

[0017] A further technical solution of the present invention is: in step 2.1, the convolution kernel size is selected as 1×64, the stride is 1, and the number of channels changes from 1 to 16.

[0018] A further technical solution of the present invention is that in the step 2.1, the kernel size of the max pooling is 2, the stride is 2, and the size of the obtained data is halved while the number of channels remains unchanged.

[0019] A further technical solution of the present invention is that in the step 2.2, the convolution kernel size of the first convolution operation of each residual module is 1×64, the stride is 1, and the number of channels changes from 16 to 16. The mathematical expression is:

[0020]

[0021] Where x represents the original data input to the residual network in step 2.2, y represents the output of x after the first convolution, h1 represents the convolution kernel function, n represents the nth value of the output data, and i represents the ith value of the input.

[0022] A further technical solution of the present invention is that in the step 2.2, the convolution kernel size of the second convolution of each residual module is 1×64, the stride is 1, and the number of channels is 16. The mathematical expression is as follows:.

[0023]

[0024] Where y is the output of the first convolution and serves as the input of the second convolution, l represents the output of the second convolution, h2 represents the convolution kernel function, m represents the mth value of the output, and j represents the jth value of the input. Before each convolution, the data is normalized and non-linearly mapped, and the mapping function is Relu for both times to obtain the residual data l[0,m].

[0025] A further technical solution of the present invention is that in the step 2.3, each channel data output by the residual module is processed to obtain a new expression:

[0026] l[0,m] = w2(δ(w1x))) (7)

[0027] H(x) = x + l[0,m] (8)

[0028] l[0,m] contains 16-channel data. The residual data and the original data are added as the output to form the SE_ResNet_17 model:

[0029] H′(x) = x + l′[0,m] (9).

[0030] The SE_ResNet_17 model first uses the residual network to extract the underwater acoustic signal features. The residual network eliminates the gradient disappearance phenomenon in the optimization process. By analyzing the output of the residual network, the channels where the feature data effective for the recognition task are located are further found, and the channels are weighted to improve the correct rate of underwater acoustic recognition.

[0031] Advantages of the Invention

[0032] The technical advantages of the present invention are as follows: The present invention studies the characteristics of underwater acoustic signals. Aiming at the repeatability of the Residual Network (ResNet) in extracting features from sparse underwater acoustic time domains, a SE_ResNet_17 model is proposed to identify underwater acoustic targets. This model can effectively remove redundant underwater acoustic time-domain features and extract effective target recognition features. Description of the Drawings

[0033] Figure 1 It is a flowchart of the method for underwater target recognition based on the SE_ResNet_17 network provided by an embodiment of the present invention;

[0034] Figure 2 It is a structural diagram of the ResNet model provided by an embodiment of the present invention;

[0035] Figure 3 It is an architecture diagram of the SE network provided by an embodiment of the present invention;

[0036] Figure 4 It is a structural diagram of underwater target recognition based on the SE_ResNet_17 structure provided by an embodiment of the present invention;

[0037] Figure 5 It is the recognition rate of different networks when adding noises with different signal-to-noise ratios provided by an embodiment of the present invention;

[0038] Figure 6 It is a comparison diagram of the recognition rates of different datasets provided by an embodiment of the present invention; Detailed Embodiments

[0039] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present invention.

[0040] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations.

[0041] Accordingly, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0042] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other.

[0043] It should be noted that similar reference numerals and letters indicate similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0044] See Figures 1 - 6 , the embodiments of the present invention provide an underwater target recognition method based on the SE_ResNet_17 model, including:

[0045] Obtain the original random underwater acoustic signals of different types of underwater acoustic targets in two different acoustic field environments, then process the original underwater acoustic signals, and use the SE_ResNet_17 model to extract the stable features of underwater targets. The steps are as shown in the appendix Figure 1 shown, and its features include the following steps:

[0046] Step 1: Collect two datasets of different types of underwater acoustic targets in two different acoustic field environments, denoted as dataset 1 and dataset 2 respectively. The two datasets have the same preprocessing. Frame the time-domain signals, with n feature points in one frame and no overlap between frames. After windowing, analyze all feature points. If the maximum feature point in the sample is less than 0.1, remove the sample. Removing the small-value frame samples can ensure that the recognition result is not affected by special sample points, and then normalize all samples.

[0047] Step 2: Send the training set into the SE_ResNet_17 model to train a stable recognition model.

[0048] The ResNet network uses residual learning to solve network degradation. When the input is x, for a stacked layer structure, the learned feature is denoted as H(x). When the number of stacked layers is too large, the learned feature is difficult to be optimized through backpropagation of gradients. Therefore, find a function that is easier to optimize, F(x) = H(x) - x, and F(x) is called the residual function. The original feature learning can be expressed as F(x) + x. The architecture of the ResNet network is as shown in the appendix Figure 2 shown, where module 1 to module N is a stacked structure, and the data expression is shown in Equation (1):

[0049] H(x) = x + wN δ(w N-1 (δ(...δ(w1x)))) (1)

[0050] where w1…w N represent the weights of each module in the residual network, and δ represents the activation function. When backpropagating, the gradient of H(x) with respect to x is calculated, and the mathematical expression is shown in Equation (2).

[0051]

[0052] As shown in Equation (2), the first term is 1, and the second term is the gradient value of the weight function with respect to x. Since it contains 1, even if the second term is very small, the phenomenon of gradient disappearance will not occur during backpropagation.

[0053] The ResNet network uses residual learning to improve the performance of the network from the spatial dimension level, aggregating spatial information and feature dimension information. However, for underwater acoustic target recognition, it is desired to use the network to extract only the most significant feature attributes for the underwater acoustic recognition task. The SE (Squeeze-and-Excitation) network solves the problems of diverse spatial feature dimensions and redundant feature channels of underwater acoustic signals by explicitly modeling the interdependencies between feature channels, making the underwater acoustic features extracted by the network more suitable for the recognition task. The structure of the SE network is as shown in the appendix Figure 3 where X represents the original input data, U represents the residual module in the ResNet network, H represents the width of the data, W represents the height of the data, and c represents the number of channels of the data.

[0054] The SE network analyzes the global information on each channel and then uses the method of feature reconstruction to analyze the interdependencies between channels. Instead of introducing a new spatial dimension for feature channel fusion, the SE network adopts a brand-new feature recalibration strategy. Specifically, it automatically obtains the importance of each feature channel through learning, and then enhances the useful features and suppresses the features that are not very useful for the current task according to this importance. The SE network recalibrates the response of the channels to the features through the interdependencies between channels to enhance the representation ability of the network.

[0055] The SE network is divided into two parts, the squeeze part and the excitation part. Since SE is embedded in the ResNet network, the data in ResNet is the output of the convolutional neural network, and each group of data is generated using a fixed convolutional kernel, so the data with a fixed receptive field size is obtained, and global features cannot be extracted. Using global average pooling to generate channel statistics is called the squeeze part, and the mathematical model is shown in Equation (3)

[0056]

[0057] where u c represents the data of the c-th channel of the input data of the SE network, and z c represents the output data of the c-th channel of the compression part. H represents the length of the data, and W represents the width of the data.

[0058] The obtained compressed data represents the global features of each channel. In order to independently learn the non-linear characteristics between channels, a gated system with an activation function is adopted, which is called the excitation part. The mathematical model is shown in Equation (4).

[0059] s = σ(w2δ(w1z)) (4)

[0060] where z is the output of the compression part, w1 and w2 are the weights of the network mapping. In order to obtain the features of the network channels, a fully connected mapping is adopted, and the number of feature points before mapping is r times that after mapping. Therefore c is the number of channels of z. δ is the Rlue activation function, and σ is the sigmoid activation function. The fully connected network can use dimensionality reduction to extract effective features, and then perform feature recombination on the effective features, so as to improve the effectiveness of the features for specific tasks. The SE network obtains channel weights conditional on the input, that is, a self-attention mechanism on the channels.

[0061] Based on the ResNet network of the SE network, on the basis of eliminating the vanishing gradient, by analyzing the global features of the channels, the channels are weighted to further optimize the mapping relationship between the input data and the output data, and improve the recognition effect of underwater acoustic targets.

[0062] According to the characteristics of underwater acoustic target recognition, a suitable network structure and parameters are selected. The ResNet structure contains two residual structures. In each structure, two convolutional layers are stacked, and the residual network loops twice to form the SE_ResNe_17 network. The entire network structure is as shown in the appendix Figure 4As shown in the figure, first, a convolution operation is performed on the input signal. Considering the time-domain characteristics of existing underwater acoustic signals, the convolution kernel is selected as 1×64, the stride is 1, and the number of channels becomes 16. Then, the Relu activation function is used for non-linear mapping. Next, the maximum pooling method is applied to the output data to obtain the feature data of different channels. The kernel size of the maximum pooling is 2 and the stride is 2, so the data size is halved while the number of channels remains unchanged. The data obtained above is used as the original data of the SE_ResNet_17 network. Two convolution operations are performed on the original data. The convolution kernel sizes are both 1×64, the stride is 1, and the number of channels is 16. Before each convolution, the data is normalized and non-linearly mapped, and the mapping function is Relu in both cases to obtain the residual data. Analyze the channel information of the residual data and weight the channels. First, perform a global analysis on the channel data, and take the average value of each channel data as the channel feature. To further extract the effective features of the channels, fully connect and map the 16-channel data to 4 channels, and then reconstruct it to 16 channels. The reconstructed data can be used as the representation of the channel feature, that is, the channel weight, which is multiplied by the residual data to obtain the residual data after channel weighting. Add the residual data and the original data as the output. Repeat the network twice to obtain the output data. Map the output data to different classification dimensions, and use the backpropagation algorithm to optimize the network to obtain the classification result.

[0063] For the signal processed in step 1, 1 / 4 of the data is used as the test set and 3 / 4 of the data is used as the training set. In the same training, the training set and the test set are fixed and there are no repeated samples between them. For the training process of the network, it is implemented using stochastic gradient descent. The underwater acoustic samples obtained in step 1 and their corresponding class labels are used to train the network. The softmax function of the final output features of the network and the cross-entropy function of the corresponding labels are used as the optimization basis of the network.

[0064] There are mainly two reasons affecting the robustness. One is that the noise contained in the samples reduces the data quality, and the other is that the similarity between the types in the samples is relatively high, and the model is prone to confusion. To verify the strong robustness of the proposed method, the present invention adopts two groups of experiments. In the first group of experiments, noise signals with different signal-to-noise ratios are added to Data 1 to verify the robustness of the network when the data quality is reduced by noise. The second group of experiments compares the recognition results of two groups of data with extremely high similarity between the classes contained in the samples and differences between the classes contained in the samples to verify the robustness of the network when the similarity between the classes is extremely high.

[0065] An underwater target recognition method based on the SE_ResNet_17 structure. The network uses the channel weighting method to eliminate channel redundancy, increase the correctness of feature extraction, and uses the adaptive weighting method to improve the adaptability of the system to different data, thereby improving the robustness of the system. Its features include the following steps:

[0066] Step 1: Sample, frame, and window the two groups of original underwater acoustic signals respectively. The original signals are underwater acoustic signals in wav format. Among them, Data 1 contains three types of underwater acoustic data, and Data 2 contains four types of underwater acoustic data. Sample the time-domain signals at a specific frequency. For the sampled signals, every n points are taken as one frame. In order to preserve the complete characteristics of the underwater acoustic signals, apply a Hamming window to the framed signals, and then perform small-sample removal and normalization processing on all samples. Obtain two groups of experimental data, namely processed Data 1 and Data 2.

[0067] Step 2: Build an underwater target recognition model based on the SE_ResNet_17 structure.

[0068] First, perform a convolution operation on the input signal. For the time-domain waveform of the underwater acoustic signal, when the scale size is 64, effective target features can be extracted more completely. Therefore, the convolution kernel size is selected as 1×64, the stride is 1, and the number of channels changes from 1 to 16. Then, use the Relu activation function for non-linear mapping. Next, use the max-pooling method for the output data to obtain feature data of different channels. The kernel size of the max-pooling is 2, the stride is 2, and the size of the obtained data is halved while the number of channels remains unchanged.

[0069] The data obtained above is used as the original data of the residual network. When using a convolutional neural network to extract the characteristics of underwater acoustic signals, when the convolution kernel size is 1×64, 16 convolutional kernels can extract significant underwater acoustic waveform features, but cannot extract fine waveform features. Therefore, when constructing the residual function, perform two convolution operations on the original data. Due to the feature extraction of underwater acoustic waveforms, if too many convolutional kernels are used at one time, it will cause repeated extraction of waveforms, and if too few convolutional kernels are used at one time, only a small amount of features will be extracted. After experiments, 16 is selected as the number of convolutional kernels for one time, which can extract features with a small amount of feature redundancy. The convolution kernel size of the first convolution operation is 1×64, the stride is 1, and the number of channels changes from 1 to 16. The mathematical expression is as shown in Equation (5), where x represents the input, y represents the output, h1 represents the convolution kernel function, n represents the nth value of the output, and i represents the i-th value of the input.

[0070]

[0071] The convolution kernel size of the second convolution is 1×64, the stride is 1, and the number of channels remains 16. The mathematical expression is as shown in Equation (6), where y is the output of the first convolution and serves as the input of the second convolution. l represents the output of the second convolution, h2 represents the convolution kernel function, m represents the mth value of the output, and j represents the j-th value of the input. Normalize and perform non-linear mapping on the data before each convolution, and the mapping function is Relu for both, to obtain the residual data l[0,m].

[0072]

[0073] l[0,m] is the data after two - layer convolution. As can be seen from Equation (1), the formula expression can be written as Equation (7), where x is the input and w is the weight. By adding l[0,m] to the original data, a new mapping function is obtained. As shown in Equation (8), the new mapping can eliminate the problem of gradient vanishing. From the analysis of Equation (2), H(x) can effectively eliminate gradient vanishing and improve the recognition effect.

[0074] l[0,m] = w2(δ(w1x))) (7)

[0075] H(x) = x + l[0,m] (8)

[0076] l[0,m] contains 16 - channel data. Global analysis is performed on the channel data. Since each channel is the feature extracted by a group of convolution kernel functions, it represents a waveform feature. It is necessary to analyze the importance of this feature in recognition and weight the channels. First, global analysis is performed on the channel data, and the average value of each channel data is taken as the channel feature. To further extract the effective features of the channels, the 16 - channel data are fully - connected and mapped to 4 channels, and then reconstructed to 16 channels. The data after reconstruction can be used as the representation of the channel feature, that is, the channel weight. Multiply it with the residual data to get the residual data l′[0,m] after channel weighting. Add the residual data to the original data as the output, as shown in Equation (9). Repeat the residual network twice for the obtained output data. Map the output data to different classification dimensions, and use the backpropagation algorithm to optimize the network to obtain the classification result. The network structure contains 17 mapping operations, forming the SE_ResNet_17 model.

[0077] H′(x) = x + l′[0,m] (9)

[0078] Step 3: Use Dataset 1 as the experimental data to verify the robustness of the proposed method.

[0079] 1. Process Dataset 1: 1 / 4 of the data of each class in Dataset 1 is used as the test set, and 3 / 4 of the data is used as the training set. There are 1884 samples in the training set and 629 samples in the test set.

[0080] 2. Add noise signals with different signal - to - noise ratios to both the training set and the test set. The signal - to - noise ratio ranges from - 20 to 20 dB, with a group of data every 5 dB. A total of 10 groups of datasets with different signal - to - noise ratios are obtained including the original signal.

[0081] 3. Use the 10 sets of data obtained above to train and test the SE_ResNet_17 network. The training method uses the batch processing method. 64 samples are randomly selected in each batch. The selected samples will not be used as candidate samples for the next batch. The learning rate is 0.0001, the optimization method is stochastic gradient descent, and 0.9 is selected as the sparse function during optimization. Iterate 100 times. The recognition rate of the test set is calculated by averaging the recognition results 5 times with random initial parameters.

[0082] 4. Using the 10 sets of data obtained above, several commonly used underwater recognition neural networks were trained and tested. The GAN network model adopted batch training with a batch size of 64 and a learning rate of 0.001. The optimization method was the stochastic gradient descent algorithm, and the optimization was not sparse. The DBN network model was trained using a batch processing method with a batch size of 64. The gradient descent algorithm was selected for optimization during the training process, and the learning rate was 0.01. The Densenet network was trained using a batch processing method with a batch size of 32. The gradient descent method was selected as the optimization method during the training process, and the learning rate was 0.001. The experimental method of the ResNet network was the same as that of the SE_ResNet_17 network. The experimental results obtained were compared with the recognition results of the underwater target recognition model based on the SE_ResNet_17 structure. The comparison chart is shown in the attached figure. Figure 5 As shown in the figure, the dotted line represents the SE_ResNet_17 network, the solid triangle represents the ResNet network, the dotted triangle represents the GAN network, the solid circle represents the Densenet network, and the dotted circle represents the DBN network. The experimental results show that the recognition effect of the SE_ResNet_17 network is better than that of other networks within the signal-to-noise ratio range shown in the figure.

[0083] Step 4: Use dataset 1 and dataset 2 as experimental data to verify the robustness of the proposed method.

[0084] 1. Processing Dataset 1 and Dataset 2, 1 / 4 of the data of each class of samples in Dataset 1 is used as the test set, and 3 / 4 of the data is used as the training set, resulting in 1884 samples in the training set and 629 samples in the test set. 1 / 4 of the data of each class of samples in Dataset 2 is used as the test set, and 3 / 4 of the data is used as the training set, resulting in 4499 samples in the training set and 1500 samples in the test set.

[0085] 2. Noise signals with different signal-to-noise ratios are added to both the training and test sets of the two data sets. The signal-to-noise ratio range is -20 to 20 dB, with one set of data every 5 dB. Data set 1 and data set 2 each include the original signal to obtain 10 sets of data sets with different signal-to-noise ratios.

[0086] 3. Use the ten sets of data from dataset 1 and the ten sets of data from dataset 2 to train and test the SE_ResNet_17 network and the ResNet network respectively. Batch processing is used for both networks. 64 samples are randomly selected in each batch. The selected samples will not be used as candidate samples for the next batch. The learning rate is 0.0001, the optimization method is stochastic gradient descent, 0.9 is selected as the sparse function during optimization, and iterates 100 times. The recognition rate of the test set is calculated by performing 5 experiments with random initial parameters and taking the average value of the recognition results. The experimental results are shown in the attached figure. Figure 6 As shown in the figure, the dotted circle represents the recognition result of data 1 in the SE_ResNet_17 network, the solid circle represents the recognition result of data 2 in the SE_ResNet_17 network, the dotted triangle represents the recognition result of data 1 in the ResNet network, and the solid triangle represents the recognition result of data 2 in the ResNet network. As can be seen from the figure, for data 1 with a large difference between categories, the recognition rate of the SE_ResNet_17 network is higher than that of the ResNet network in the entire signal-to-noise ratio range. For data 2 with a small difference between categories, the recognition rate of the SE_ResNet_17 network is still higher than that of the ResNet network in the entire signal-to-noise ratio range, showing strong robustness.

Claims

1. An SE_ResNet_17 method applicable to underwater acoustic target recognition in a dynamic environment, characterized in that It includes the following steps: Step 1: Collect two groups of underwater acoustic target datasets of different types in different sound field environments, denoted as dataset 1 and dataset 2 respectively; after collection, perform the same preprocessing on the two datasets for sample normalization; It includes the following sub-steps: Step 1.1: Frame the time-domain signals of the dataset, with n feature points in one frame and no overlap between frames; Step 1.2: Perform windowing processing on the framed time-domain signals in Step 1.1 to obtain multiple groups of samples of the same length. Analyze all sample feature points. If the maximum feature point in the sample is less than 0.1, eliminate the small-value frame samples; Step 1.3: Normalize all samples after eliminating the small-value frame samples; Step 2: Feed the samples into the SE_ResNet_17 model to train a stable recognition model; It includes the following sub-steps: Step 2.1: Select an appropriate convolution size for the time-domain waveform of the input signal, and perform convolution and max-pooling methods on the effective target features to obtain feature data; the convolution kernel size is selected as 1×64, the stride is 1, and the number of channels changes from 1 to 16; Step 2.2: The obtained data is used as the original data of the residual network. After passing through two residual modules, each residual module contains two convolution operations to obtain feature data of multiple channels; use the Relu activation function for non-linear mapping, and then perform the max-pooling method on the output data to obtain feature data of different channels. The kernel size of the max-pooling is 2 and the stride is 2, and the size of the obtained data is halved while the number of channels remains unchanged; Step 2.3: Process each channel data output by the residual module to obtain the weights of the feature data in the channel in the recognition task, and then use the weights to perform a weighted operation on the output data of the residual network.

2. The SE_ResNet_17 method for underwater acoustic target recognition applicable to dynamic environments according to claim 1, characterized in that, In Step 2.2, the convolution kernel size of the first convolution operation of each residual module is 1×64, the stride is 1, and the number of channels changes from 16 to 16. The mathematical expression is: where x represents the original data input to the residual network in Step 2.2, y represents the output of x after the first convolution, h1 represents the convolution kernel function, n represents the nth value of the output data, and i represents the ith value of the input.

3. The SE_ResNet_17 method for underwater acoustic target recognition applicable to a dynamic environment according to claim 1, wherein In Step 2.2, the convolution kernel size of the second convolution of each residual module is 1×64, the stride is 1, and the number of channels is 16. The mathematical expression is as follows: where y is the output of the first convolution and serves as the input of the second convolution, l represents the output of the second convolution, h2 represents the convolution kernel function, m represents the mth value of the output, and j represents the jth value of the input; before each convolution, the data is normalized and non-linearly mapped, and the mapping function is Relu for both to obtain the residual data l[0,m].

4. The SE_ResNet_17 method for underwater acoustic target recognition applicable to a dynamic environment according to claim 1, characterized in that In Step 2.3, each channel data output by the residual module is processed to obtain a new expression: l[0,m] = w2(δ(w1x))) (7) H(x) = x + l[0,m] (8) l[0,m] contains 16-channel data. Add the residual data and the original data as the output to form the SE_ResNet_17 model: H′(x) = x + l′[0,m] (9) The SE_ResNet_17 model first uses the residual network to extract the features of underwater acoustic signals. The residual network eliminates the phenomenon of gradient disappearance during the optimization process. By analyzing the output of the residual network, the channels where the feature data effective for the recognition task are located are further found, and the channels are weighted to improve the correct rate of underwater acoustic recognition.

Citation Information

Patent Citations

  • Underwater acoustic target recognition method based on signal processing and deep-shallow network multi-model fusion

    CN112364779A

  • Dish identification method based on improved YOLO v3

    CN112560918A