A method for identifying ship radiated noise based on data enhancement

By constructing a data enhancement model of variational autoencoder, the depth characteristics of radiated noise are extracted using time-delay convolutional encoder and transposed time-delay convolutional decoder, and the data reconstruction mechanism generates enhanced data, the problem of low recognition accuracy caused by insufficient data and environmental complexity in the prior art is solved, and higher recognition accuracy and wider application range are achieved.

CN114818789BActive Publication Date: 2025-05-09NANJING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210362029.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-07
Publication Date
2025-05-09
Estimated Expiration
2042-04-07

AI Technical Summary

Technical Problem

Existing underwater radiation noise identification technology is limited by insufficient data and environmental complexity, resulting in low recognition accuracy and limited application range.

Method used

Using a data augmentation method, a data augmentation model of variational autoencoder is constructed, and the depth characteristics of radiated noise are extracted using time-delay convolution encoder and transposed time-delay convolution decoder, and augmented data is generated through the data reconstruction mechanism to expand the training data set.

Benefits of technology

It effectively improves the recognition accuracy and robustness of the radiation noise recognition model and expands the scope of technical applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114818789B_ABST
    Figure CN114818789B_ABST
Patent Text Reader

Abstract

The invention discloses a ship radiation noise identification method based on data enhancement, comprising: collecting radiation noise samples containing ship information; performing windowing processing on the samples to obtain a plurality of small sections of radiation noise, regularizing the windowed radiation noise samples and attaching corresponding labels to different samples; extracting MFCC features of all samples, inputting them into a time-delay convolution encoder of a variational autoencoder data enhancement model, and obtaining distribution parameters of corresponding feature spaces; sampling from the feature space distribution to obtain a sampling feature vector and inputting the sampling feature vector into a transposed time-delay convolution decoder of the variational autoencoder data enhancement model to obtain corresponding reconstructed data; training the variational autoencoder data enhancement model, generating a large amount of generated data after the training is completed to expand the original training samples; training a classifier using the expanded training samples; and predicting the test data to obtain the corresponding predicted ship category.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a ship radiation noise identification method, in particular to a ship radiation noise identification method based on data enhancement. Background Art

[0002] my country has a vast sea area and rich marine resources. With the rapid development of modern shipping industry, fishery and other industries, the research on underwater acoustic signals has received more and more attention. Underwater radiated noise is a common underwater acoustic signal, which is mostly generated by ships sailing in the water. Because the radiated noise of a ship can reflect the ship's running speed, cargo capacity, working condition of parts and other hull information, the radiated noise signal is often used as an important source of information for analyzing ships. At present, most of the ship radiated noise analysis work is completed by trained professionals. Due to the cost of professional training and the efficiency of manual identification, the current application scope of this technology is still very limited. In order to reduce the cost of radiated noise identification technology and expand the scope of application of this technology, more and more scientific researchers have invested in the research of automatic identification of radiated noise.

[0003] However, compared with signals propagating in the air, underwater acoustic signals are more susceptible to the influence of underwater channels and the ever-changing aquatic environment. Due to the random movement of the sea surface, the unevenness of the seabed and its changes over time, and the uneven distribution of water bodies, the underwater acoustic channel is not only non-uniform in space, but also randomly variable in the time domain. In addition, the propagation speed of underwater acoustic signals is slow, the code element period is long, and the complex underwater channel leads to the time-varying and non-stationary characteristics of underwater acoustic signals. At the same time, the marine environment is more diverse than on land. There are various organisms, seawater movement, and noise generated by ships underwater. The absorption frequency of sound wave signals in different waters is also inconsistent, and the speed of signal propagation in waters of different depths is also different. This causes the actual signal propagated underwater to be interfered by multipath effects, Doppler effects, and various noise signals. Due to the diversity of problems and the complexity of processing, underwater acoustic signal recognition has always been a challenging topic.

[0004] Most of the existing research on underwater radiation noise recognition is based on classic signal processing algorithms. This type of algorithm completes signal noise reduction and signal feature extraction by constructing complex signal processing models, and then completes classification by comparing the similarity of features between different signals. This type of method has good interpretability and theoretical basis, but is easily affected by environmental factors. When the experimental scene changes, it is necessary to rely on relevant field knowledge to adjust the model parameters. In recent years, with the development of deep learning theory and the replacement of data computing equipment, deep neural network technology has developed rapidly, and has achieved great results in image recognition, speech recognition, natural language processing and other fields. At present, more and more scholars have proposed the use of deep learning models to construct radiation noise signal recognition models. Compared with the previous time-frequency analysis method, the deep learning model can extract more expressive nonlinear features, and the recognition performance has also been greatly improved. However, most deep learning algorithms require large-scale data sets to support model training. Due to the high acquisition cost and complex acquisition methods of radiation noise signals, the scale of experimental data is very limited. Therefore, it is crucial to carry out data enhancement research and expand the noise signal data set to improve the recognition performance of the model. Summary of the invention

[0005] Purpose of the invention: The technical problem to be solved by the present invention is to provide a ship radiated noise identification method based on data enhancement in view of the shortcomings of the prior art.

[0006] In order to solve the above technical problems, the present invention discloses a ship radiated noise identification method based on data enhancement, comprising the following steps:

[0007] Step 1: construct a variational autoencoder data enhancement model, which consists of a time-delay convolutional encoder and a transposed time-delay convolutional decoder; collect radiation noise samples containing ship driving information, take the type of ship emitting the radiation noise as its category label, obtain the original radiation noise dataset B1, and divide the dataset B1 into a training set B2 and a test set B3 according to the ratio;

[0008] Step 2, by performing a windowing operation on the training set B2, a number of segment signals are segmented from the original radiation noise signal, the ship category labels of the segment signals are the same as the corresponding original radiation noise signal labels, and the Mel frequency cepstral coefficient MFCC features of the segment signals are used for variational autoencoder data enhancement model training to obtain training set B4;

[0009] Step 3, input the data Mel frequency cepstral coefficient I in the training set B4 into the time-delay convolution encoder, obtain the probability distribution of the input data Mel frequency cepstral coefficient I in the deep feature space, the probability distribution is a normal distribution, and the time-delay convolution encoder outputs the mean M and standard deviation S of the probability distribution;

[0010] Step 4, using the mean M and standard deviation S output by the time-delay convolution encoder, randomly sample to obtain a new feature vector V, input the feature vector V into the transposed time-delay convolution decoder, decode the feature vector V layer by layer, and obtain the reconstructed Mel frequency cepstral coefficient MFCC data O;

[0011] Step 5, calculate the reconstruction error between the input data Mel frequency cepstral coefficient I and the reconstructed Mel frequency cepstral coefficient MFCC data O, as well as the deviation between the probability distribution of the deep feature vector and the standard normal distribution, use these error and deviation terms to calculate the parameter update value of the variational autoencoder data enhancement model, and use the parameter update value to update the corresponding parameters in the variational autoencoder data enhancement model;

[0012] Step 6: Input the training set B4 data into the trained variational autoencoder data enhancement model to generate more reconstructed data. The category label of the reconstructed data is consistent with the input data. Then, the reconstructed data is used to expand the training set B4 to obtain the expanded training set B5.

[0013] Step 7: Use the expanded training set B5 to train a ResNet-18 classifier, and save the trained classifier result file as result file F;

[0014] Step 8: Use the result file F to identify each radiation noise signal in the test set B3, obtain the ship category to which the test radiation noise belongs, and complete the ship radiation noise identification based on data enhancement.

[0015] In step 1 of the present invention, a passive sonar device is used to collect the radiation noise emitted by ships traveling in the port to obtain the original radiation noise data, that is, the radiation noise sample, and the type of the ship emitting the signal is used as the category of the radiation noise signal. All the collected radiation noise samples and the category data of the radiation noise signal are combined into an original radiation noise data set B1, and the data set B1 is divided into a training set B2 and a test set B3 according to a ratio.

[0016] The windowing operation described in step 2 of the present invention includes: using a windowing method, specifying a fixed window size as W, dividing the original data in the training set B2 into several segments of windowed signals with a length of W, the category label of each segment of the windowed signal is the category label of the original data to which it belongs, extracting the corresponding Mel-frequency cepstral coefficient MFCC feature from the segment signal obtained by windowing, which is a two-dimensional time-frequency feature, and using the Mel-frequency cepstral coefficient MFCC features of the window segment signal to form a training set B4 for training the variational autoencoder data enhancement model.

[0017] In step 3 of the present invention, the Mel-frequency cepstral coefficient I is selected from the training set B4 as the input of the variational autoencoder data enhancement model. The variational autoencoder data enhancement model consists of two parts: a time-delay convolutional encoder and a transposed time-delay convolutional decoder. The input data is first learned through the time-delay convolutional encoder structure to obtain the feature space probability distribution corresponding to the input data, and then the feature space sampling vector is decoded to obtain the reconstructed data, thereby completing the task of data enhancement;

[0018] The time-delay convolution encoder is composed of three time-delay convolution units in cascade, each of which contains four parts: convolution, transposition operation, batch normalization and residual connection. The time-delay convolution encoder extracts the deep features of the input data Mel frequency cepstral coefficient I, expands the deep features into one-dimensional features, and inputs them into two independent fully connected neural networks to obtain the mean M and standard deviation S of the random distribution of the deep features.

[0019] In step 4 of the present invention, random sampling is performed from a multidimensional normal distribution with a mean value of M and a standard deviation of S obtained by the time-delay convolutional encoder to obtain a new feature vector V, and the following re-parameter sampling strategy is adopted:

[0020] V'~N(0,1)

[0021] V=V'*S+M

[0022] Where N(0,1) represents a standard normal distribution with a mean of 0 and a standard deviation of 1. The reparameter sampling strategy first samples a random vector V' from the standard normal distribution N(0,1), and then generates a feature vector V that conforms to a multidimensional normal distribution with a mean of M and a standard deviation of S through the characteristics of the normal distribution.

[0023] The transposed time-delay convolution decoder is another component of the variational autoencoder data enhancement model. The transposed time-delay convolution decoder is composed of three transposed time-delay convolution units in cascade, and each transposed time-delay convolution unit includes four parts: transposition operation, diffusion convolution, batch normalization and residual connection. Through the transposed time-delay convolution decoder, the feature vector V is decoded into a feature map, and finally a reconstructed output reconstructed Mel-frequency cepstral coefficient MFCC data O of the same size as the input Mel-frequency cepstral coefficient MFCC feature is obtained.

[0024] In step 5 of the present invention, the reconstruction error between the input data Mel frequency cepstral coefficient I and the reconstructed Mel frequency cepstral coefficient MFCC data O, as well as the deviation between the probability distribution of the deep feature vector and the standard normal distribution are calculated to obtain the loss function L for evaluating the performance of the data enhancement model, and the method is as follows:

[0025] L=mse_loss(O,I)-KL(N(M,S^2)|N(0,1))

[0026] Among them, the reconstruction error term mse_loss(O,I) in the loss function calculates the mean square error between the output reconstructed Mel frequency cepstral coefficient MFCC data O and the input Mel frequency cepstral coefficient I. The closer the output O is to the input Mel frequency cepstral coefficient I, the smaller the reconstruction error is, and the larger the inverse rule is; the regularization term KL(N(M,S^2)|N(0,1)) in the loss function calculates the KL divergence between the probability distribution of the deep feature vector N(M,S^2) and the standard normal distribution N(0,1);

[0027] The error back propagation algorithm (reference: RUMELHART DE, HINTON GE, WILLIAMS RJ. Learning representations by back-propagating errors [J]. nature, 1986, 323 (6088): 533–536.) is used to calculate the parameter update values ​​of the model using these error values, and then the parameters of the variational autoencoder data enhancement model are updated using the stochastic mini-batch gradient descent algorithm (reference: Li M, Zhang T, Chen Y, et al. Efficient mini-batch training for stochastic optimization [C] / / Proceedings of the 20th ACM SIGKDD international conference on Knowledge discovery and data mining. 2014: 661-670.).

[0028] In step 6 of the present invention, each data in the training set B4 is input into the variational autoencoder data enhancement model, the time-delay convolution encoder of the model outputs the probability distribution of the data in the feature space, and a feature vector is obtained by sampling in the feature space based on the probability distribution. The feature vector is input into the transposed time-delay convolution decoder to generate reconstructed data of the same category as the original data, and the reconstructed data is used to expand the training set B4 to obtain the expanded training set B5.

[0029] In step 7 of the present invention, the expanded training set B5 is used, and the data in the training set B5 is combined with its corresponding sample category label to be input into the ResNet-18 classifier for training, and the cross entropy loss function is used to calculate the deviation between the predicted category label output during the training process and the actual sample category label. The updated value of the ResNet-18 model parameter is calculated by the back propagation algorithm, and the stochastic gradient descent algorithm is used to update the parameters of the ResNet-18 model. Repeat multiple rounds of training until the change trend of the cross entropy loss function value converges; after obtaining the trained ResNet-18 classifier, save the training result as the result file F.

[0030] In step 8 of the present invention, the ResNet-18 classifier obtained in step 7 is used to identify the test signal in the test set B3. First, the original test signal is windowed to obtain a segmented radiated noise signal that is the same as the training data. Then, the Mel-frequency cepstral coefficient MFCC features of the segmented signal are extracted and input into the ResNet-18 classifier for prediction to obtain the ship category to which the segmented radiated noise signal belongs.

[0031] The transposed delay convolution process in step 4 of the present invention includes:

[0032] Transpose operation: This operation exchanges the channel dimension and high dimension of the input feature vector;

[0033] Diffused convolution: The diffuse convolution operation diffuses each element of the transposed feature vector into a matrix of the same size as the transposed convolution kernel; then the convolution result between the diffuse matrix and the transposed convolution kernel is calculated, and the above operation is repeated for each element in the input feature to obtain a series of intermediate results of the convolution operation, and then these intermediate results are matrix-added to obtain the result of the diffuse convolution;

[0034] Batch Normalization: To alleviate the vector distribution drift caused by multi-layer nonlinear operations, a batch normalization layer is connected after each diffusion convolution to dynamically adjust the feature distribution of batch training data;

[0035] Residual connection: A residual connection branch is added between the input of the transposed delayed convolution and the output of the batch normalization layer to alleviate the gradient vanishing and gradient exploding problems of deep models.

[0036] Beneficial effects:

[0037] The present invention overcomes the problem of insufficient training data for the current radiation noise recognition model based on deep learning, and constructs a variational autoencoder data enhancement model based on a time-delay convolution structure. The model extracts the frame-dimensional depth features of radiation noise through a time-delay convolution module, and uses the data reconstruction mechanism of the variational autoencoder data enhancement model to generate enhanced data to complete the expansion of the training data set. This method effectively increases the scale of the original training data set and improves the recognition accuracy of the radiation noise recognition model.

[0038] The significant advantage of the present invention is that it expands the scale of the radiation noise data set, effectively alleviates the problem of insufficient data for the deep learning model, and improves the robustness and recognition performance of the recognition model. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments, and the above and / or other advantages of the present invention will become more clear.

[0040] Figure 1 The figure is a flow chart of the operation of the system of the present invention.

[0041] Figure 2 This is a diagram showing the effect of the algorithm in the present invention on the radiation noise data set A.

[0042] Figure 3 This is a diagram showing the effect of the algorithm in the present invention on the radiation noise data set B.

[0043] Figure 4 This is a comparison chart of the heat map corresponding to the generated MFCC data and the original MFCC data obtained by the time-delay convolutional variational autoencoder data enhancement model in the present invention. DETAILED DESCRIPTION

[0044] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0045] like Figure 1 As shown, it is the operation flow chart of the system of the present invention, which includes 8 steps.

[0046] Step 1: Collect radiation noise samples containing ship driving information, take the type of ship emitting the radiation noise as its category label, obtain the original radiation noise dataset B1, and divide the dataset B1 into a training set B2 and a test set B3 in a ratio of 2:1;

[0047] Step 2, by performing a windowing operation on the training set B2, a number of segment signals are segmented from the original radiation noise signal, the ship category labels of the segment signals are the same as the corresponding original signal labels, and the Mel-frequency cepstral coefficient MFCC features of the segment signals are combined into the training set B4;

[0048] Step 3, input the Mel frequency cepstral coefficient I of the training set B4 data into the time-delay convolution encoder to obtain the probability distribution of the input data in the deep feature space. The probability distribution is a normal distribution, and the encoder outputs the mean M and standard deviation S of the distribution;

[0049] Step 4: Use the mean M and standard deviation S output by the time-delay convolution encoder to randomly sample a new feature vector V, input V into the transposed time-delay convolution decoder, decode the feature vector layer by layer, and obtain the reconstructed Mel frequency cepstral coefficient MFCC data O;

[0050] Step 5, calculate the reconstruction error between the input data Mel frequency cepstral coefficient I and the reconstructed Mel frequency cepstral coefficient MFCC data O, as well as the deviation between the probability distribution of the deep feature vector and the standard normal distribution, use these error and deviation terms to calculate the parameter update value of the variational autoencoder data enhancement model composed of the time delay convolution encoder and the transposed time delay convolution decoder, and use the parameter update value to update the corresponding parameters in the variational autoencoder data enhancement model;

[0051] Step 6: Input the training set B4 data into the trained variational autoencoder data enhancement model to generate more reconstructed data. The category label of the reconstructed data is consistent with the input data. Then, the reconstructed data is used to expand the training set B4 to obtain the expanded training set B5.

[0052] Step 7: Use the expanded training set B5 to train a ResNet-18 classifier, and save the trained classifier result file as result file F;

[0053] Step 8: Use file F to identify each radiation noise signal in the test set B3 to obtain the ship category to which the test radiation noise belongs.

[0054] In step 1, passive sonar equipment is used to collect the radiated noise emitted by ships in the port to obtain the original radiated noise data. Ship radiated noise is the sound signal generated by the interaction of multiple noise sources on the ship and the water medium in which it is located. The main sources of radiated noise include: propellers, rotating and reciprocating machinery, various pumps, etc. Identifying the type of ship to which the radiated noise belongs is a common underwater acoustic signal processing task. After completing the data collection, the type of ship that emits the signal is used as the category of the radiated noise signal, and the collected data is composed of the original radiated noise data set B1. The original radiated noise data in the data set B1 is divided according to the ratio of 2:1 to obtain the training set B2 and the test set B3.

[0055] In step 2, the windowing method is used. Since the length of the original radiated noise signal collected is generally long, it is too complex to directly use it as the input of the classification model. Therefore, a fixed window size W (Mel frequency cepstrum coefficients, MFCC) is used here. The fixed window size W = 10000 is specified. The original radiated noise data in the training set B2 is divided into several segments of windowed signals with a length of W. The category label of each segment of the windowed signal is the category label of the original data to which it belongs. Then the MFCC features of the windowed segmented noise signal are extracted. The MFCC features are two-dimensional time-frequency features that describe the acoustic characteristics of the signal and are often used in acoustic signal recognition tasks. The labeled MFCC features are used to form the training set B4. (Reference: Davis S, Mermelstein P. Comparison of parametric representations for monosyllabic word recognition in continuously spoken sentences [J]. IEEE transactions on acoustics, speech, and signal processing, 1980, 28 (4): 357–366.)

[0056] In step 3, data is selected from the training set B4 as the input Mel-frequency cepstral coefficient I of the variational autoencoder data enhancement model. The delayed convolution encoder structure extracts the deep features of the data through the delayed convolution unit. The delayed convolution unit includes four parts: convolution, transposition operation, batch normalization and residual connection.

[0057] 1. Convolution. In commonly used convolutional neural network models, a 3*3 convolution kernel is usually used because this size of convolution kernel can better capture the detailed texture in the image. However, for ship radiation noise, the signal sampling frequency is concentrated between 8kHz and 16kHz. At this time, the receptive field corresponding to the convolution feature obtained by the 3*3 convolution kernel is too small to describe the signal characteristics well. Delay convolution is an extended structure based on time domain convolution. The kernel height of delay convolution is a fixed value, which is the same as the height of the input feature. Delay convolution adopts the design of large convolution kernel to increase the receptive field size corresponding to the convolution feature.

[0058] 2. Transpose operation. In the output of the delayed convolution, the channel dimension and high dimension of the output feature map are exchanged, so that the delayed convolution feature still has the time-frequency distribution characteristics of the input data.

[0059] 3. Batch normalization. Since the convolution kernel used in delayed convolution is large in size, in order to alleviate the difficulty of parameter optimization caused by the large convolution kernel, a batch normalization layer is connected after the convolution kernel transposition operation to dynamically adjust the feature distribution of the batch training data.

[0060] 4. Residual connection: A residual connection branch is added between the input of the delayed convolution unit and the output of the batch normalization layer to alleviate the problem of gradient vanishing and gradient exploding in deep models.

[0061] Then the deep time-delay convolutional features are expanded into one-dimensional features and input into two independent fully connected neural networks to obtain the mean M and standard deviation S of the random distribution of deep features. (Reference: Kingma DP, Welling M. Auto-encoding variational bayes[J].arXiv preprint arXiv:1312.6114,2013)

[0062] In step 4, random sampling is then performed from a multidimensional normal distribution with a mean of M and a standard deviation of S to obtain a new feature vector V. Because the random sampling operation is not integrable, in order to update some of the encoder network parameters in the subsequent parameter update phase, the following re-parameter sampling strategy is adopted:

[0063] V'~N(0,1)

[0064] V=V'*S+M

[0065] Where N(0,1) represents a standard normal distribution with a mean of 0 and a standard deviation of 1. The resampling strategy first samples a random vector V' from the standard normal distribution N(0,1), and then generates a feature vector V that conforms to a multidimensional normal distribution with a mean of M and a standard deviation of S through the characteristics of the normal distribution;

[0066] The transposed delayed convolution decoder consists of a series of transposed delayed convolution layers with opposite input and output characteristics to the delayed convolution layers in the delayed convolution encoder, so that the feature vector generated by the encoder can be gradually decoded to obtain the reconstructed output. Through the transposed delayed convolution encoder structure, the feature vector V is gradually decoded into a feature map of larger size, and finally the reconstructed output O of the same size as the input MFCC feature is obtained. Transposed delayed convolution is the reverse process of delayed convolution, which is mainly divided into the following steps:

[0067] 1. Transpose operation. This operation exchanges the channel dimension and high dimension of the input feature vector;

[0068] 2. Diffusion convolution. The diffusion convolution operation diffuses each element of the transposed feature vector into a matrix of the same size as the transposed convolution kernel. Then the convolution result between the diffusion matrix and the transposed convolution kernel is calculated, and the above operation is repeated for each element in the input feature to obtain a series of intermediate results of the convolution operation, and then these intermediate results are matrix-added to obtain the result of the diffusion convolution;

[0069] 3. Batch normalization. In order to alleviate the problem of vector distribution drift caused by multi-layer nonlinear operations, a batch normalization layer is connected after each diffusion convolution to dynamically adjust the feature distribution of batch training data;

[0070] 4. Residual connection: A residual connection branch is added between the input of the transposed delayed convolution and the output of the batch normalization layer to alleviate the problem of gradient vanishing and gradient exploding in deep models.

[0071] In step 5, the reconstruction error between the input data Mel frequency cepstral coefficient I and the reconstructed Mel frequency cepstral coefficient MFCC data O, as well as the deviation between the probability distribution of the deep feature vector and the standard normal distribution are calculated to obtain the loss function for evaluating the performance of the data enhancement model:

[0072] L=mse_loss(O,I)-KL(N(M,S^2)|N(0,1))

[0073] The reconstruction error term mse_loss(O,I) in the loss function calculates the mean square error between the output reconstructed Mel-frequency cepstral coefficient MFCC data O and the input Mel-frequency cepstral coefficient I. The closer the output reconstructed Mel-frequency cepstral coefficient MFCC data O and the input Mel-frequency cepstral coefficient I are, the smaller the reconstruction error is, and the larger the inverse rule is; the regularization term KL(N(M,S^2)|N(0,1)) in the loss function calculates the Kullback-Leibler divergence between the probability distribution of the deep feature vector N(M,S^2) and the standard normal distribution N(0,1), referred to as KL divergence, which reflects the degree of difference between the two distributions;

[0074] The error back propagation algorithm is used to use these error values ​​to calculate the parameter update values ​​of the model, and then the parameters of the variational autoencoder data enhancement model are updated through the random mini-batch gradient descent algorithm.

[0075] In step 6, the MFCC features of each radiation noise signal in the training set B4 are input into the variational autoencoder data enhancement model. The time-delay convolution encoder structure of the data enhancement model first outputs the probability distribution of the data in the feature space, and then samples the feature space based on the probability distribution to obtain a feature vector. The feature vector is input into the transposed time-delay convolution decoder to generate reconstructed data of the same category as the original data. The reconstructed data is used to expand the training set B3 to obtain the expanded training set B5.

[0076] In step 7, the expanded training set B5 is used to input the data therein into the ResNet-18 (18-layer residual convolutional neural network, Residual Network-18, ResNet-18) classifier for training. The cross entropy loss function is used to calculate the deviation between the predicted label output by ResNet-18 and the true sample label during the training process, and the updated value of the ResNet-18 model parameters is calculated by the back propagation algorithm. The parameters of the ResNet-18 model are updated using the stochastic gradient descent algorithm, and multiple rounds of training are repeated until the change trend of the cross entropy loss function value converges. After obtaining the trained ResNet-18 classifier, the training results are saved as the result file F.

[0077] In step 8, the ResNet-18 (18-layer residual convolutional neural network, ResidualNetwork-18, ResNet-18) classifier obtained in step 7 is used to identify the signal in the test set B3. First, the original test signal is windowed to obtain a segmented radiated noise signal that is the same as the training data. Then, the MFCC features of the segmented signal are extracted and input into the ResNet-18 classifier for prediction to obtain the ship category to which the segmented radiated noise signal belongs. (Reference: He K, Zhang X, Ren S, et al. Deep residual learning for image recognition [C] / / Proceedings of the IEEE conference on computer vision and pattern recognition. 2016: 770–778.)

[0078] Example

[0079] In order to perform pre-processing before the system is operated, the present invention needs to train the system algorithm model, and the training set is the radiation noise signals of different ship types.

[0080] To obtain the radiated noise training set, the present invention uses a passive sonar device to collect the radiated noise emitted by ships entering and leaving the ports in different ports, and then manually labels the ship categories. Finally, the present invention obtains two sets of radiated noise data sets with ship category label information, and each set of data sets finally includes more than 50 original radiated noise signals.

[0081] The present invention extracts 33% of the original radiation noise signal from the above two radiation noise data sets as the test signal. Figure 1 The trained ResNet-18 classifier is used to detect the ship category, and the actual noise signal ship category and the predicted noise ship category are compared to calculate the category prediction accuracy on the test set.

[0082] Using the above-mentioned radiation noise training set and radiation noise test set, follow the steps below to train and evaluate the system model:

[0083] 1. Model training based on radiation noise signal:

[0084] 1.1 Using the windowing method, a fixed window size of W = 10000 is specified, and the original radiation noise data in the training set is divided into several segments of windowed signals with a length of W. The category label of each segment of the windowed signal is the category label of the original data to which it belongs. Then, the MFCC features of each windowed segmented signal are extracted to form the training data set of the model;

[0085] 1.2 Regularize the scale of the samples, that is, the size of the samples should be the same, so as to eliminate the influence of the scale of the training samples on the model training, and convert the labels of different categories into digital values, such as 0, 1, 2, ...;

[0086] 1.3 Use the processed sample set as the input of the variational autoencoder data enhancement model, extract the deep features of the input data through the time-delay convolution encoder structure, and then expand the deep time-delay convolution features into one-dimensional features, input two independent fully connected neural networks, and obtain the mean M and standard deviation S of the random distribution of the deep features;

[0087] 1.4 Then randomly sample from a multidimensional normal distribution with a mean of M and a standard deviation of S to obtain a new feature vector. Through the transposed delayed convolution encoder structure, the above feature vector is gradually decoded into a feature map of a larger size, and finally a reconstructed output of the same size as the input data is obtained.

[0088] 1.5 Calculate the reconstruction error between the input data and the reconstructed output data, as well as the deviation between the probability distribution of the deep feature vector and the standard normal distribution, use these error terms to calculate the parameter update value of the model, and update the parameters of the variational autoencoder data enhancement model;

[0089] 1.6 Input the training set data into the trained variational autoencoder data enhancement model to generate more reconstructed data. The category labels of the reconstructed data are consistent with the input data, and then the reconstructed data is used to expand the training set;

[0090] 1.7 Use the expanded training set to train a ResNet-18 classifier and save the model parameters obtained from the training.

[0091] 2. Testing

[0092] 2.1 For each original signal in the test set, first perform the same windowing process as the training data to obtain a series of windowed segmented signals;

[0093] 2.2 Extract the MFCC features of each windowed segmented signal, and then use it as the input of the ResNet-18 classifier to predict the ship category label;

[0094] 2.3 By comparing the predicted labels and true labels of the windowed segmented signals, the accuracy and AUC value (Area Under ROC Curve, AUC) of the ResNet-18 classifier on the test set are calculated.

[0095] 2.3 For the label category of the original signal, the voting method can be used to count the number of labels of each windowed segmented signal contained in the original signal, and the ship category label with the most occurrences is used as the predicted label of the original signal.

[0096] Based on the above training and testing steps, a ship radiated noise data enhancement and identification system was finally obtained. The accuracy of radiated noise identification based on this data enhancement method reached more than 75%. In addition, the use of data enhancement strategy solves the shortcomings of traditional deep learning methods with poor robustness under small data sets. Therefore, the application of the present invention for ship radiated noise identification has the advantages of good robustness and high prediction accuracy.

[0097] Figure 2The recognition algorithm part of the present invention is listed for the recognition of segmented signal types on the radiation noise data set 1 collected by the present invention. The results show that the present invention has excellent performance in statistical accuracy. The meanings of some indicators in the table are as follows: "algorithm" indicates different recognition models, FCN represents the full convolutional neural network classifier, and ResNet-18 represents the residual convolutional neural network structure used in the present invention; "data enhancement" indicates different data enhancement methods, "none" represents the original training set that has not been expanded by the data enhancement method, "mfcc" represents the training set expanded by the data enhancement method that directly applies perturbations to the MFCC features, "wavelet" represents the training set expanded by the data enhancement method that applies perturbations to the original signal in the wavelet domain using the discrete wavelet decomposition method, and "VAE-tdc" represents the training set expanded by the variational autoencoder data enhancement model proposed by the present invention. Auccuary and AUC represent two measurement indicators, accuracy and ROC curve area. The larger these two indicators are, the better the recognition effect of the code model.

[0098] Figure 3 The recognition algorithm of the present invention is listed in the radiation noise data set 2 collected by the present invention for the recognition of segmented signal types. The meaning of each indicator of the experimental results and Figure 2 Consistent, Figure 2 It is illustrated that the variational autoencoder data enhancement model proposed in the present invention has a significant improvement effect on different measured radiation noise signals, further proving its effectiveness.

[0099] Figure 4 A comparison chart of the heat maps corresponding to the reconstructed MFCC data obtained by the variational autoencoder data enhancement model and the original MFCC data is given. It can be clearly seen that the enhanced data generated by the variational autoencoder data enhancement model has richer changes compared with the original data, while the enhanced data retains the distribution characteristics of the original data, which is very beneficial for improving the robustness and recognition accuracy of the recognition model.

[0100] The present invention provides a method and idea for ship radiated noise identification based on data enhancement. There are many methods and approaches to implement the technical solution. The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the protection scope of the present invention. All components not specified in this embodiment can be implemented by existing technologies.

Claims

1. A ship radiated noise identification method based on data enhancement, characterized in that: The steps include: Step 1: construct a variational autoencoder data enhancement model, which consists of a time-delay convolutional encoder and a transposed time-delay convolutional decoder; collect radiation noise samples containing ship driving information, take the type of ship emitting the radiation noise as its category label, obtain the original radiation noise dataset B1, and divide the dataset B1 into a training set B2 and a test set B3 according to the ratio; Step 2, by performing a windowing operation on the training set B2, a number of segment signals are segmented from the original radiation noise signal, the ship category labels of the segment signals are the same as the corresponding original radiation noise signal labels, and the Mel frequency cepstral coefficient MFCC features of the segment signals are used for variational autoencoder data enhancement model training to obtain training set B4; Step 3, input the data Mel frequency cepstral coefficient I in the training set B4 into the time-delay convolution encoder, obtain the probability distribution of the input data Mel frequency cepstral coefficient I in the deep feature space, the probability distribution is a normal distribution, and the time-delay convolution encoder outputs the mean M and standard deviation S of the probability distribution; Step 4, using the mean M and standard deviation S output by the time-delay convolution encoder, randomly sample to obtain a new feature vector V, input the feature vector V into the transposed time-delay convolution decoder, decode the feature vector V layer by layer, and obtain the reconstructed Mel frequency cepstral coefficient MFCC data O; Step 5, calculate the reconstruction error between the input data Mel frequency cepstral coefficient I and the reconstructed Mel frequency cepstral coefficient MFCC data O, as well as the deviation between the probability distribution of the deep feature vector and the standard normal distribution, use these error and deviation terms to calculate the parameter update value of the variational autoencoder data enhancement model, and use the parameter update value to update the corresponding parameters in the variational autoencoder data enhancement model; Step 6: Input the training set B4 data into the trained variational autoencoder data enhancement model to generate more reconstructed data. The category label of the reconstructed data is consistent with the input data. Then, the reconstructed data is used to expand the training set B4 to obtain the expanded training set B5. Step 7, use the expanded training set B5 to train a ResNet-18 classifier, and save the trained classifier result file as result file F; Step 8: Use the result file F to identify each radiation noise signal in the test set B3, obtain the ship category to which the test radiation noise belongs, and complete the ship radiation noise identification based on data enhancement.

2. According to the method for identifying ship radiated noise based on data enhancement in claim 1, it is characterized in that: In step 1, a passive sonar device is used to collect the radiation noise emitted by ships passing through the port to obtain the original radiation noise data, i.e., the radiation noise sample, and the type of the ship emitting the signal is used as the category of the radiation noise signal. All the collected radiation noise samples and the category data of the radiation noise signal are combined into an original radiation noise data set B1, and the data set B1 is divided into a training set B2 and a test set B3 according to the ratio.

3. The ship radiated noise identification method based on data enhancement according to claim 2 is characterized in that: The windowing operation in step 2 includes: using the windowing method, setting a fixed window size of W, dividing the original data in the training set B2 into several segments of windowed signals with a length of W, the category label of each segment of the windowed signal is the category label of the original data to which it belongs, extracting the corresponding Mel frequency cepstral coefficient MFCC feature from the windowed segment signal, which is a two-dimensional time-frequency feature, and using the Mel frequency cepstral coefficient MFCC features of the window segment signal to form a training set B4 for training the variational autoencoder data enhancement model.

4. The ship radiated noise identification method based on data enhancement according to claim 3 is characterized in that: In step 3, the Mel-frequency cepstral coefficient I is selected from the training set B4 as the input of the variational autoencoder data enhancement model. The variational autoencoder data enhancement model consists of two parts: a time-delay convolutional encoder and a transposed time-delay convolutional decoder. The input data is first learned through the time-delay convolutional encoder structure to obtain the feature space probability distribution corresponding to the input data, and then the feature space sampling vector is decoded to obtain the reconstructed data to complete the data enhancement task; The time-delayed convolutional encoder is composed of three cascaded time-delayed convolutional units, each of which contains four parts: convolution, transposition operation, batch normalization and residual connection. The time-delayed convolutional encoder extracts the deep features of the Mel-frequency cepstral coefficient I of the input data, expands the deep features into one-dimensional features, and inputs them into two independent fully connected neural networks to obtain the mean M and standard deviation S of the random distribution of the deep features.

5. The ship radiated noise identification method based on data enhancement according to claim 4 is characterized in that: In step 4, random sampling is performed from the multidimensional normal distribution with mean M and standard deviation S obtained from the time-delay convolution encoder to obtain a new feature vector V, using the following re-parameter sampling strategy: V'~N(0,1) V=V'*S+M Where N(0,1) represents a standard normal distribution with a mean of 0 and a standard deviation of 1. The reparameter sampling strategy first samples a random vector V' from the standard normal distribution N(0,1), and then generates a feature vector V that conforms to a multidimensional normal distribution with a mean of M and a standard deviation of S through the characteristics of the normal distribution. The transposed time-delay convolution decoder is another component of the variational autoencoder data enhancement model. The transposed time-delay convolution decoder is composed of three transposed time-delay convolution units in cascade, and each transposed time-delay convolution unit includes four parts: transposition operation, diffusion convolution, batch normalization and residual connection. Through the transposed time-delay convolution decoder, the feature vector V is decoded into a feature map, and finally a reconstructed output reconstructed Mel-frequency cepstral coefficient MFCC data O of the same size as the input Mel-frequency cepstral coefficient MFCC feature is obtained.

6. The ship radiated noise identification method based on data enhancement according to claim 5 is characterized in that: In step 5, the reconstruction error between the input data Mel frequency cepstral coefficient I and the reconstructed Mel frequency cepstral coefficient MFCC data O, as well as the deviation between the probability distribution of the deep feature vector and the standard normal distribution are calculated to obtain the loss function L for evaluating the performance of the data enhancement model, as follows: L=mse_loss(O,I)-KL(N(M,S^2)|N(0,1)) Among them, the reconstruction error term mse_loss(O,I) in the loss function calculates the mean square error between the output reconstructed Mel-frequency cepstral coefficient MFCC data O and the input Mel-frequency cepstral coefficient I. The closer the output reconstructed Mel-frequency cepstral coefficient MFCC data O is to the input Mel-frequency cepstral coefficient I, the smaller the reconstruction error is, and the larger the inverse rule is; the regularization term KL(N(M,S^2)|N(0,1)) in the loss function calculates the KL divergence between the probability distribution of the deep feature vector N(M,S^2) and the standard normal distribution N(0,1); The error back propagation algorithm is used to use these error values ​​to calculate the parameter update values ​​of the model, and then the parameters of the variational autoencoder data enhancement model are updated through the random mini-batch gradient descent algorithm.

7. The ship radiated noise identification method based on data enhancement according to claim 6 is characterized in that: In step 6, each data in the training set B4 is input into the variational autoencoder data enhancement model. The time-delay convolution encoder of the model outputs the probability distribution of the data in the feature space, and a feature vector is obtained by sampling in the feature space based on the probability distribution. The feature vector is input into the transposed time-delay convolution decoder to generate reconstructed data of the same category as the original data. The reconstructed data is used to expand the training set B4 to obtain the expanded training set B5.

8. The ship radiated noise identification method based on data enhancement according to claim 7 is characterized in that: In step 7, the expanded training set B5 is used to input the data in the training set B5 into the ResNet-18 classifier in combination with its corresponding sample category label for training. The cross entropy loss function is used to calculate the deviation between the predicted category label output during the training process and the actual sample category label. The updated value of the ResNet-18 model parameter is calculated by the back propagation algorithm. The stochastic gradient descent algorithm is used to update the parameters of the ResNet-18 model. Multiple rounds of training are repeated until the change trend of the cross entropy loss function value converges. After obtaining the trained ResNet-18 classifier, the training result is saved as the result file F.

9. The ship radiated noise identification method based on data enhancement according to claim 8 is characterized in that: In step 8, the ResNet-18 classifier obtained in step 7 is used to identify the test signal in the test set B3. First, the original test signal is windowed to obtain a segmented radiated noise signal that is the same as the training data. Then, the Mel-frequency cepstral coefficient MFCC features of the segmented signal are extracted and input into the ResNet-18 classifier for prediction to obtain the ship category to which the segmented radiated noise signal belongs.

10. A ship radiated noise identification method based on data enhancement according to claim 9, characterized in that: Step 4 includes: Transpose operation: This operation exchanges the channel dimension and high dimension of the input feature vector; Diffused convolution: The diffuse convolution operation diffuses each element of the transposed feature vector into a matrix of the same size as the transposed convolution kernel; then the convolution result between the diffuse matrix and the transposed convolution kernel is calculated, and the above operation is repeated for each element in the input feature to obtain a series of intermediate results of the convolution operation, and then these intermediate results are matrix-added to obtain the result of the diffuse convolution; Batch Normalization: To alleviate the vector distribution drift caused by multi-layer nonlinear operations, a batch normalization layer is connected after each diffusion convolution to dynamically adjust the feature distribution of batch training data; Residual connection: A residual connection branch is added between the input of the transposed delayed convolution and the output of the batch normalization layer to alleviate the gradient vanishing and gradient exploding problems of deep models.

Citation Information

Patent Citations

  • A training sample data expansion method and device based on a variational auto-encoder

    CN109886388A

  • Seismic data expansion method based on variational auto-encoder

    CN111258992A