An underwater acoustic data generation method based on generative adversarial networks
By constructing an adaptive stable deep convolution generation adversarial network model and a progressive learning strategy, the problems of scarcity and instability of training in underwater acoustic target classification are solved, and high-quality samples are generated quickly and efficiently, which significantly improves classification accuracy.
Patent Information
- Application Number
- CN202411846687.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-16
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-12-16
AI Technical Summary
The existing underwater acoustic target classification methods perform poorly in the context of complex noise, and the scarcity of data leads to overfitting of deep learning models and unstable training, and fail to effectively utilize the step-by-step learning strategy from low to high frequency, and the quality of the generated results is not ideal.
Adaptive stable deep convolution generation adversarial network model is built, combined with adaptive controllers and progressive learning strategies, adjust the training ratio of generators and discriminators through adaptive controllers, accelerate network convergence, and adopt progressive learning to gradually learn from low frequency to high frequency to generate high-quality Mel frequency cepspectral coefficients.
The network convergence speed is significantly accelerated, training time is reduced by about 40%, the generated sample quality is improved, and the classification accuracy of underwater acoustic targets is significantly improved through data augmentation, with classification accuracy reaching 82.74% and 86.41% respectively.
Smart Images

Figure CN119293570B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of underwater acoustic target classification, and particularly to an underwater acoustic data generation method based on a generative adversarial network. Background Art
[0002] In the field of underwater acoustic target classification (UATC), commonly used methods include those based on traditional signal processing means, such as wavelet transform, low-frequency analysis and recording (LOFAR), and envelope modulation noise detection (DEMON). These methods can obtain good results under specific conditions, but perform poorly in complex noise backgrounds and changing acoustic environments.
[0003] In recent years, deep learning methods have also been introduced. For example, convolutional neural networks (CNNs) are used for underwater target classification. However, deep learning methods rely on a large amount of labeled data, especially high-quality labeled data. The problem of data scarcity leads to the model being prone to overfitting. Generative adversarial networks (GANs) are used to generate underwater acoustic data, but the training process is very unstable and the convergence speed is slow, which limits their wide application in underwater acoustic target classification. The complex underwater environment generates a large amount of sample noise, which further increases the training difficulty of GANs. Moreover, existing methods usually directly train on complex data and fail to effectively utilize the strategy of gradually learning from low frequency to high frequency, resulting in the quality of the generated results not being ideal enough.
[0004] Therefore, there is an urgent need for an underwater acoustic data generation method based on a generative adversarial network to overcome the above technical defects. Summary of the Invention
[0005] The purpose of the present invention is to provide an underwater acoustic data generation method based on a generative adversarial network, by constructing an adaptive stable deep convolutional generative adversarial network model for generating synthetic samples of underwater acoustic targets to overcome the problem of existing data scarcity and improve the network convergence speed; and by generating high-quality mel-frequency cepstral coefficients (MFCCs) to improve the classification accuracy of underwater acoustic targets.
[0006] To achieve the above purpose, the present invention provides an underwater acoustic data generation method based on a generative adversarial network, including the following steps:
[0007] Step S1: Based on the generative adversarial network model, in combination with an adaptive controller, construct an adaptive stable deep convolutional generative adversarial network model;
[0008] Step S2: Based on the adaptive controller, adjust the training steps of the generator and the discriminator to accelerate the network convergence process;
[0009] Step S3: During the training process, adopt a progressive learning strategy to gradually learn from low frequency to high frequency, stabilize the generation process, and improve the quality of sample generation.
[0010] Preferably, in step S1, based on the generative adversarial network model and combined with an adaptive controller, construct an adaptive stable deep convolutional generative adversarial network model;
[0011] Among them, the adaptive stable deep convolutional generative adversarial network model includes: an adaptive controller and a progressive learning strategy.
[0012] Preferably, in step S2, based on the adaptive controller, adjust the training steps of the generator and the discriminator to accelerate the convergence process of the network. The workflow is as follows:
[0013] Step S21: Initialize the proportional parameters of the generator G and the discriminator D;
[0014] Step S22: During the training process, generate a random probability P;
[0015] Step S23: If P meets the controller setting, when P tends to the generator, train the generator network; otherwise, train the discriminator network;
[0016] Step S24: If P does not meet the controller setting, then adjust the proportional parameters of G and D according to the controller;
[0017] Step S25: Compare the adjusted ratio with P to determine the probability of training G or D in the next iteration;
[0018] Step S26: Loop the above steps until the end.
[0019] Preferably, in step S3, during the training process, adopt a progressive learning strategy to gradually learn from low frequency to high frequency, stabilize the generation process, and improve the quality of sample generation. The specific implementation steps are as follows:
[0020] Step S31: Initialize the fuzzy processing of Mel-frequency cepstral coefficients MFCCs;
[0021] Step S32: During the training process, perform fuzzy processing on the dataset and the generated MFCCs;
[0022] Step S33: If in the early stage of training, allow the generator to learn without giving priority to details and train the generator network;
[0023] Step S34: If in the later stage of training, gradually reduce the degree of fuzziness so that the generator focuses on generating clearer MFCCs;
[0024] Step S35: Loop the above steps until the end.
[0025] Therefore, the present invention adopts the above-mentioned underwater acoustic data generation method based on a generative adversarial network, and the beneficial effects are as follows:
[0026] 1) The present invention improves the convergence speed of the network: Compared with the traditional generative adversarial network (GAN) model, the AS-DCGAN model with an adaptive controller can significantly accelerate the convergence speed of the network and reduce the training time by about 40%;
[0027] 2) The samples generated by the present invention have higher quality: The progressive learning strategy enables the model to gradually learn high-frequency features from low-frequency features, thereby generating high-quality Mel-frequency cepstral coefficients (MFCCs). These synthetic samples can significantly improve the classification accuracy of underwater acoustic targets;
[0028] 3) The data augmentation effect of the present invention is obvious: On two public datasets (DeepShip and ShipsEar), through data augmentation with synthetic samples, the classification accuracies are respectively improved to 82.74% and 86.41%, which are significantly better than the classification accuracies of the original datasets.
[0029] The technical solution of the present invention will be further described in detail below with reference to the drawings and embodiments. Description of the Drawings
[0030] Figure 1 is the model framework of the Adaptive Stable Deep Convolutional Generative Adversarial Network (AS-DCGAN) for the underwater acoustic data generation method based on a generative adversarial network of the present invention. Detailed Embodiments
[0031] The technical solution of the present invention will be further described below with reference to the drawings and embodiments.
[0032] The underwater acoustic data generation method based on a generative adversarial network of the present invention includes the following steps:
[0033] Step S1: Based on the generative adversarial network model, combined with an adaptive controller, construct an adaptive stable deep convolutional generative adversarial network model.
[0034] As Figure 1 shown, the present invention constructs an Adaptive Stable Deep Convolutional Generative Adversarial Network (AS-DCGAN) model based on the generative adversarial network model to generate synthetic samples of underwater acoustic targets to overcome the problem of scarce existing data.
[0035] Among them, the adaptive stable deep convolutional generative adversarial network model includes two modules: an adaptive controller (AC) and a progressive learning (PL) strategy. The training of the generator and discriminator in the adversarial network is adjusted by the adaptive controller (AC), and the training ratio is changed to achieve network convergence while also accelerating the convergence speed. The progressive learning (PL) strategy is introduced at the generator end of the adversarial network to gradually stabilize the process of generating samples. By introducing low-frequency feature blurring and gradually weakening the blurring effect to ensure high-frequency features, the quality of the generated samples can be effectively improved, and finally, underwater acoustic data can be generated quickly, efficiently, and with high quality.
[0036] Step S2: Based on the adaptive controller (AC), adjust the training steps of the generator and discriminator to accelerate the convergence process of the network. The workflow is as follows:
[0037] Step S21: Initialize the proportional parameters of the generator (G) and discriminator (D).
[0038] Step S22: During the training process, generate a random probability P.
[0039] Step S23: If P meets the controller settings, when P tends to the generator, train the generator network; otherwise, train the discriminator network.
[0040] Step S24: If P does not meet the controller settings, adjust the proportional parameters of G and D according to the controller.
[0041] Step S25: Compare the adjusted ratio with P to determine the probability of training G or D in the next iteration.
[0042] Step S26: Loop the above steps until the end.
[0043] The above process effectively reduces the training time by about 40% through the dynamic adjustment of the training ratio by the adaptive controller.
[0044] Step S3: During the training process, adopt the progressive learning (PL) strategy to learn gradually from low frequency to high frequency, stabilize the generation process, and improve the quality of sample generation. The specific implementation steps are as follows:
[0045] Step S31: Initialize the blurring process of the Mel-frequency cepstral coefficients (MFCCs).
[0046] Step S32: During the training process, perform blurring processing on the dataset and the generated MFCCs.
[0047] Step S33: If in the early stage of training, allow the generator to learn without giving priority to details and train the generator network.
[0048] Step S34: If it is in the later stage of training, gradually reduce the degree of blurriness so that the generator focuses on generating clearer MFCCs.
[0049] Step S35: Loop the above steps until the end.
[0050] Embodiment
[0051] This embodiment conducts a performance test on a method for generating underwater acoustic data based on a generative adversarial network proposed by the present invention.
[0052] Based on the constructed Adaptive Stable Deep Convolutional Generative Adversarial Network (AS-DCGAN) model, high-quality spectrograms can be generated for UATC. In the experimental comparison by controlling the ratio of original data and generated data, regardless of the ratio of original data, the recognition accuracy of the model increases with the increase in the ratio of generated images in the training set. When the ratio of generated images reaches 80%, the recognition accuracy of the model reaches 84.53% in ShipsEar and 80.49% in DeepShip, and finally reaches 86.41% in ShipsEar and 82.74% in DeepShip as the ratio of generated images increases. This result exceeds the recognition accuracy of the model when only using the original images as input data, which are 80.48% in ShipsEar and 78.38% in DeepShip respectively. It shows that by using high-quality generated images to augment data, the proposed model can effectively improve the recognition accuracy of underwater acoustic targets. It should be noted that when only using the generated images for training, the recognition accuracy of the model in ShipsEar and DeepShip is still 60.74% and 57.55% respectively. This result indicates that the generated images are of high quality and can be used to train the classification network, and the recognition accuracy of the model will gradually increase as the ratio of generated images gradually increases.
[0053] Therefore, the present invention adopts the above method for generating underwater acoustic data based on a generative adversarial network. By constructing an Adaptive Stable Deep Convolutional Generative Adversarial Network model, it is used to generate synthetic samples of underwater acoustic targets to overcome the problem of scarce existing data, and significantly accelerate the convergence speed of the network, reducing the training time by about 40%; the progressive learning strategy enables the model to gradually learn high-frequency features from low-frequency features, thereby generating high-quality Mel Frequency Cepstral Coefficients (MFCCs), significantly improving the classification accuracy of underwater acoustic targets; on two public datasets (DeepShip and ShipsEar), through data augmentation with synthetic samples, the classification accuracy is improved to 82.74% and 86.41% respectively, significantly superior to the classification accuracy of the original dataset.
[0054] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions of the present invention or make equivalent replacements, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. An underwater acoustic data generation method based on a generative adversarial network, characterized in that, It includes the following steps: Step S1: Based on the generative adversarial network model and combined with an adaptive controller, construct an adaptive stable deep convolutional generative adversarial network model; Among them, the adaptive stable deep convolutional generative adversarial network model includes: an adaptive controller and a progressive learning strategy; Step S2: Based on the adaptive controller, adjust the training steps of the generator and the discriminator to accelerate the network convergence process. Its working process is as follows: Step S21: Initialize the proportional parameters of the generator G and the discriminator D; Step S22: Generate a random probability P during the training process; Step S23: If P meets the controller setting, when P tends to the generator, train the generator network; otherwise, train the discriminator network; Step S24: If P does not meet the controller setting, adjust the proportional parameters of G and D according to the controller; Step S25: Compare the adjusted proportion with P to determine the probability of training G or D in the next iteration; Step S26: Loop the above steps until the end; Step S3: During the training process, adopt a progressive learning strategy to gradually learn from low frequency to high frequency, stabilize the generation process, and improve the quality of sample generation. The specific implementation steps are as follows: Step S31: Initialize the fuzzy processing of Mel-frequency cepstral coefficients MFCCs; Step S32: During the training process, perform fuzzy processing on the dataset and the generated MFCCs; Step S33: If in the early stage of training, allow the generator to learn without giving priority to details and train the generator network; Step S34: If in the later stage of training, gradually reduce the degree of fuzziness so that the generator focuses on generating clearer MFCCs; Step S35: Loop the above steps until the end.
Citation Information
Patent Citations
A method for improving a generative adversarial network by using adaptive control learning
CN109902824A
Controllable video generation method and system based on multi-modal fusion
CN119091362A