An Underwater Acoustic Data Augmentation Method and System Based on a Hybrid Diffusion Model
The hybrid diffusion model addresses data scarcity in underwater acoustic signal classification by generating diverse training data through mel-frequency spectrum mixing, improving model generalization and robustness.
Patent Information
- Application Number
- CN202510559113.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-04-30
AI Technical Summary
In the underwater acoustic target recognition task, it is difficult for the existing technology to generate diverse synthetic data, resulting in insufficient generalization capabilities of the model and scarce data leading to overfitting and performance degradation.
Using a data augmentation method based on a hybrid diffusion model, a Mel spectrogram is extracted for feature extraction, and inter-class data is generated by combining the conditional projection layer and the diffusion model. A variational autoencoder is used to model and predict noise and data distribution to generate a diverse synthetic Mel spectrogram.
The generalization ability and robustness of the underwater acoustic target recognition model are improved, and the classification performance is significantly improved.
Smart Images

Figure CN120089154B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of underwater acoustic target recognition, and particularly relates to an underwater acoustic data augmentation method and system based on a hybrid diffusion model. Background Art
[0002] Underwater acoustic target recognition has a wide range of applications in marine environmental monitoring, port ship management, and underwater ecological monitoring. In recent years, with the successful application of deep learning in fields such as speech recognition and audio noise reduction, many researchers have begun to introduce deep learning-based methods into underwater acoustic target recognition tasks. On the one hand, due to the combined influence of underwater environmental conditions and ship maintenance status, underwater acoustic signals (such as ship radiated noise) usually have noise and complexity. This makes the acoustic data vary greatly and poses high requirements for the generalization ability of the model. On the other hand, these deep learning-based methods usually require a large amount of data for training. However, it is time-consuming, laborious, and costly to obtain various data for underwater acoustic target recognition tasks. The scarcity of data can lead to the risk of overfitting, which further damages the performance of the model. Specifically:
[0003] (1) Combining underwater channel modeling to generate synthetic data and using transfer learning can narrow the domain gap between synthetic data and real data, but the data generated by this method is relatively single and does not have the characteristic of data diversity;
[0004] (2) The generative model can enhance the diversity of data by synthesizing new, real, and credible samples, but the generative adversarial network-based method has instability;
[0005] (3) Due to the difference in the distribution between synthetic data and real underwater acoustic data, the generated data often leads to a performance decline. Summary of the Invention
[0006] The purpose of the present invention is to overcome the defects of the prior art and propose an underwater acoustic data augmentation method and system based on a hybrid diffusion model.
[0007] In view of this, the present invention proposes an underwater acoustic data augmentation method based on a hybrid diffusion model, including:
[0008] Step 1: Extract the Mel spectrogram from the original audio for feature extraction to obtain a hybrid Mel spectrogram, which constitutes a training set;
[0009] Step 2: Adopt a diffusion-based hybrid strategy to generate inter-class data and corresponding hybrid conditional labels;
[0010] Step 3: Send the hybrid conditional label into a conditional projection layer to obtain conditional embedding features, and then input them into the attention U-Net of the noise prediction module in the diffusion model to obtain predicted noise;
[0011] Step 4: Use the variational autoencoder in the diffusion model to model and predict the noise, and the data distribution of the mixed mel spectrogram, and generate the corresponding synthetic mel spectrogram through variational sampling to obtain the augmented training data.
[0012] Preferably, randomly select two different labels from the mixed conditional labels and , and generate the mixed label according to the following formula :
[0013]
[0014] where the coefficient obeys the Beta distribution;
[0015] Based on the conditional diffusion model, according to and respectively find the corresponding samples and from the training set, according to the following formula:
[0016]
[0017] to obtain the corresponding mixed data .
[0018] Preferably, the diffusion-based mixing strategy includes: when and are not equal, generate inter-class data.
[0019] Preferably, the conditional projection layer includes two cascaded fully connected layers.
[0020] On the other hand, the present invention provides an underwater acoustic data augmentation system based on a hybrid diffusion model, including:
[0021] An extraction module for extracting features of the original audio by extracting the mel spectrogram to obtain a mixed mel spectrogram and constituting a training set;
[0022] A mixed conditional label generation module for generating inter-class data and corresponding mixed conditional labels by adopting a diffusion-based mixing strategy;
[0023] A predicted noise generation module for sending the mixed conditional label into the conditional projection layer to obtain conditional embedding features, and then inputting them into the attention U-Net of the noise prediction module in the diffusion model to obtain the predicted noise;
[0024] An augmentation module for using the variational autoencoder in the diffusion model to model the predicted noise and the data distribution of the mixed mel spectrogram, and generating the corresponding synthetic mel spectrogram through variational sampling to obtain the augmented training data.
[0025] Compared with the prior art, the advantages of the present invention are as follows:
[0026] 1. The present invention first introduces the conditional diffusion model into underwater acoustic data augmentation, synthesizes the Mel spectrograms of different categories of underwater acoustic signals, and generates sufficient training data for the classification model;
[0027] 2. The present invention proposes a "Mix-Gen" (Mix-Generation) strategy, pre-mixes the labels and the corresponding Mel spectrograms to generate new mixed data-condition pairs, enabling the diffusion model to generate not only different categories of data but also data between classes;
[0028] 3. By combining conditional diffusion with the Mix-Gen strategy, the present invention synthesizes in-class and between-class Mel spectrograms. Compared with the prior art, it enhances the generalization and robustness of the model. Extensive experiments on the ShipsEar and DeepShip datasets have demonstrated the superiority of Mix-Gen. The experimental results show that the method proposed by the present invention significantly improves the classification performance, highlighting its effectiveness in underwater acoustic data augmentation. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 is a flowchart of the underwater acoustic data augmentation method based on the hybrid diffusion model of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0030] The present invention first introduces the conditional diffusion model into the underwater acoustic target recognition task, synthesizes the Mel spectrograms of different categories of underwater acoustic signals, and provides sufficient training data for the classification model. Considering that a large amount of data is required to train the diffusion model and the existing underwater acoustic datasets are insufficient, a diffusion-based hybrid (Mix-Gen) strategy is proposed. Specifically, we pre-mix the labels and the corresponding Mel spectrograms to generate new mixed data-condition pairs, thereby driving the diffusion model to generate not only different categories of data but also data between classes. This method can guide the training of the diffusion model, thereby generating diverse training data and providing guarantee for the recognition of underwater acoustic targets.
[0031] The technical solutions of the present invention will be described in detail below with reference to the drawings and embodiments.
[0032] Embodiment 1
[0033] The present invention consists of two main parts. The first part is the conditional diffusion model, which uses conditional vectors to guide the generation of data corresponding to specific labels. The second component is the hybrid generation strategy, which transforms the in-class data augmentation into a hybrid strategy, further improving the training accuracy of the classification model. Specifically, we first extract Mel spectrograms from the original underwater acoustic signals for feature extraction to form a training dataset. Then, the Mix-Gen strategy is applied to generate inter-class data and corresponding hybrid conditional labels. These hybrid conditional labels are processed through a conditional projection layer to obtain conditional embedding features, which are subsequently input into the noise prediction module Attention U-Net in the diffusion model. These features are combined with the hybrid Mel spectrograms. Finally, the variational autoencoder (VAE) generates the corresponding synthetic Mel spectrograms and uses them as training data. The detailed framework is as shown in Figure 1 where, and respectively represent the noise generated at the -th and
[0034] Conditional Diffusion Model:
[0035] The conditional diffusion model can generate different samples x according to the input label y, denoted by . It is mainly divided into two types: classifier guidance and classifier-free guidance. Given the scarcity of data in the underwater acoustic dataset, it is challenging to train an effective classifier. Therefore, this paper adopts a classifier-free method. Specifically, it follows the Bayesian model:
[0036]
[0037] where, represents the probability that the data label is y given the known data sample x, represents the probability that the obtained sample is x, represents the probability that the obtained label is y, represents the gradient of the function.
[0038] Hybrid Generation Strategy:
[0039] Specifically, given two data labels of different classes and , we define the hybrid label as follows:
[0040]
[0041] where, is sampled from the Beta distribution (θ Beta(α,α)). Similarly, we define the hybrid data It is defined as:
[0042]
[0043] Where and correspond to the tags and . respectively. During the conditional diffusion training process, we randomly generate two single-shot encoding vectors ( , ) as tags, mix them to generate as a conditional constraint, and use as the training data. According to the above formula, we can find that when the random tags satisfy is equal to , the conditional diffusion model will perform in-class data increment. On the contrary, when is not equal to , the model will learn to generate mixed data , thus achieving inter-class data augmentation.
[0044] Example 2
[0045] Example 2 of the present invention provides an underwater acoustic data augmentation system based on a hybrid diffusion model, which is implemented based on the method of Example 1. The system includes:
[0046] An extraction module, configured to extract features from the original audio by extracting the Mel spectrogram, obtain a mixed Mel spectrogram, and form a training set;
[0047] A mixed conditional label generation module, configured to generate inter-class data and corresponding mixed conditional labels by adopting a diffusion-based mixing strategy;
[0048] A predicted noise generation module, configured to send the mixed conditional label into a conditional projection layer to obtain conditional embedding features, and then input them into the attention U-Net of the noise prediction module in the diffusion model to obtain predicted noise;
[0049] An augmentation module, configured to model the predicted noise and the data distribution of the mixed Mel spectrogram using the variational autoencoder in the diffusion model, and generate corresponding synthetic Mel spectrograms through variational sampling to obtain augmented training data.
[0050] It should be noted that in the embodiments of the above system, the included modules are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional modules are only for the convenience of mutual distinction and do not limit the protection scope of the present invention.
[0051] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the embodiments, those of ordinary skill in the art should understand that any modification or equivalent replacement of the technical solutions of the present invention does not depart from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.
Claims
1. An underwater acoustic data augmentation method based on a hybrid diffusion model, comprising: Step 1: Extract the Mel spectrogram to perform feature extraction on the original audio, obtain the hybrid Mel spectrogram, and form a training set; Step 2: Adopt a diffusion-based hybrid strategy to generate inter-class data and corresponding hybrid conditional labels; Step 3: Feed the hybrid conditional labels into the conditional projection layer to obtain conditional embedded features, and then input them into the attention U-Net of the noise prediction module in the diffusion model to obtain predicted noise; Step 4: Use the variational autoencoder in the diffusion model to model the predicted noise and the data distribution of the hybrid Mel spectrogram, and generate the corresponding synthetic Mel spectrogram through variational sampling to obtain the augmented training data; Randomly select two different tags from the mixed condition tags and , and generate a mixed tag according to the following formula : ; Among them, the coefficient obeys the Beta distribution; Based on the conditional diffusion model, find the corresponding samples from the training set according to y1 and y2 respectively and , according to the following formula: ; Obtain the corresponding mixed data ; The diffusion-based hybrid strategy includes: when and are not equal, generate inter-class data.
2. The underwater acoustic data augmentation method based on the hybrid diffusion model according to claim 1, wherein The conditional projection layer includes two cascaded fully connected layers.
3. A system for the underwater acoustic data augmentation method based on the hybrid diffusion model according to claim 1, characterized in that, Comprising: An extraction module for extracting the Mel spectrogram to perform feature extraction on the original audio, obtaining the hybrid Mel spectrogram, and forming a training set; A hybrid conditional label generation module for adopting a diffusion-based hybrid strategy to generate inter-class data and corresponding hybrid conditional labels; A predicted noise generation module for feeding the hybrid conditional labels into the conditional projection layer to obtain conditional embedded features, and then inputting them into the attention U-Net of the noise prediction module in the diffusion model to obtain predicted noise; And An augmentation module for using the variational autoencoder in the diffusion model to model the predicted noise and the data distribution of the hybrid Mel spectrogram, and generating the corresponding synthetic Mel spectrogram through variational sampling to obtain the augmented training data.
Citation Information
Patent Citations
CNN underwater acoustic signal target recognition method based on data enhancement and time-frequency separation
CN112257521A
Marine mammal sound data enhancement method based on improved Inception block and SACGAN
CN118506792A