Ship audio generation method based on deep learning

By combining the technology of variational autoencoder, contrast language-audio pre-training and diffusion model, realistic ship audio is generated, which solves the problem of data scarcity and improves the audio generation quality and recognition capabilities.

CN120279937APending Publication Date: 2025-07-08CHANGZHOU UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510462072.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

Existing deep learning models perform poorly in the case of scarce ship audio data, making it difficult to effectively generate high-quality audio data, limiting their acoustic target recognition capabilities under small sample conditions.

Method used

Using a deep learning-based ship audio generation method, the variational autoencoder (VAE), contrasting language-audio pre-training (CLAP) and diffusion model (Diffusion) technology is used to generate natural and realistic audio content through text prompts, including preprocessing, encoding, decoding and iterative denoising processes, to achieve cross-modal alignment and detail capture between audio and text.

Benefits of technology

It improves the quality and fidelity of audio generation, provides strong support for ship audio data under small sample conditions, and improves the ability of acoustic target recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279937A_ABST
    Figure CN120279937A_ABST
Patent Text Reader

Abstract

The invention relates to the field of deep learning audio processing, in particular to a ship audio generation method based on deep learning. The method comprises the following steps: acquiring a ship audio set, and preprocessing all ship audios in the ship audio set to obtain a Mel spectrogram corresponding to each ship audio; an audio generation model is constructed, the audio generation model comprises a VAE module, a Diffusion module and a Clap module, the preprocessed data set is utilized to train the audio generation model, and the trained audio generation model is obtained; and inputting a text prompt into a Clap module of the trained audio generation model, and generating a ship audio by using the trained audio generation model. According to the method, the ship audio data can be effectively generated under the condition that the audio data is scarce, and a sample data guarantee is provided for continuous improvement of the acoustic target intelligent recognition capability under the small sample condition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of deep learning audio processing, and particularly to a method for generating ship audio based on deep learning. Background Art

[0002] With the rapid development of global economic integration, maritime traffic has become increasingly busy. As the most important means of transportation at sea, the number of ships has been increasing year by year. At the same time, the development of deep learning technology has brought revolutionary progress to the field of audio processing. In this context, using deep learning technology to automatically learn and extract key features from a large amount of audio data for tasks such as ship classification and supervision has certain potential.

[0003] For an identification algorithm to achieve good identification results, it is necessary to have a large amount of data for network training. However, it is often difficult to obtain underwater target sound signals. Due to the limitation of the insufficient ship audio dataset, the performance of existing deep learning models for ship tasks often fails to meet expectations, and the trained models are also difficult to be put into practical applications.

[0004] Therefore, studying how to generate audio using a small amount of target sound signals has certain theoretical value and important practical significance. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to overcome the defects of the prior art and provide a method for generating ship audio based on deep learning, which can effectively generate ship audio data in the case of scarce audio data, and provide sample data guarantee for the continuous improvement of the intelligent recognition ability of acoustic targets under small sample conditions.

[0006] To solve the above technical problem, the technical solution of the present invention is: A method for generating ship audio based on deep learning, comprising:

[0007] Obtain a ship audio set, preprocess all ship audios in the ship audio set to obtain the Mel spectrogram corresponding to each ship audio;

[0008] Construct an audio generation model. The audio generation model includes a VAE module, a Diffusion module, and a Clap module. Use the preprocessed dataset to train the audio generation model to obtain a trained audio generation model; wherein,

[0009] During the training process, the VAE module encodes the mel spectrogram to obtain the corresponding latent features of the ship audio; the Clap module maps the ship audio and the relevant text description to the same semantic space to obtain the association between the ship audio and the text description; during the forward propagation of the latent features, the Diffusion module gradually adds Gaussian noise to the latent features, and during the backward propagation of the latent features, the Diffusion module uses the text encoding processed by the Clap module as a condition to guide the gradual denoising and restore the latent features; the VAE module then converts the restored latent features back to the mel spectrogram.

[0010] Input the text prompt into the Clap module of the trained audio generation model, and use the trained audio generation model to generate ship audio.

[0011] Furthermore, preprocess the ship audio, and the specific steps are as follows:

[0012] First, cut the ship audio into segments with a preset duration and store them by category;

[0013] Then load and resample the ship audio to a unified sampling rate;

[0014] Then transform it to the frequency domain through the short-time Fourier transform and convert the amplitude value to the decibel value;

[0015] Finally, calculate the mel spectrogram using the decibel value.

[0016] Furthermore, the VAE module includes an encoder, a latent space, and a decoder; among them,

[0017] The encoder is used to map the mel spectrogram to the parameters of the latent space;

[0018] The latent space is used to generate latent features based on the parameters mapped to it by the encoder;

[0019] The decoder is used to convert the latent features processed by the Diffusion module back to the mel spectrogram.

[0020] Furthermore, during the training process, the working process of the VAE module is expressed as:

[0021] μ, log(σ 2 ) = Encoder(x)

[0022] Z = μ + σ☉∈

[0023]

[0024] In the formula, X represents the input data of the encoder; the mean μ and variance σ 2The logarithms are all parameters for the encoder to map the input data x to the latent space; ∈ represents a noise vector sampled by the VAE from the standard normal distribution; z represents the latent features generated in the latent space; z d represents the latent features after being processed by the Diffusion module.

[0025] Furthermore, during the training process, the working process of the Clap module is as follows:

[0026] First, the text is converted into a high-dimensional vector representation through the text encoder to capture the text semantic information; the mel spectrogram is converted into a high-dimensional vector representation through the audio encoder to capture the audio acoustic features;

[0027] Then, a contrastive learning is carried out to select the matching text-audio pairs as the positive sample pairs and the non-matching text-audio pairs as the negative sample pairs, and the cosine similarity is used to calculate the similarity between the positive sample pairs and the negative sample pairs respectively;

[0028] Finally, a contrastive loss function is used for training to make the similarity of the positive sample pairs greater than that of the negative sample pairs. By minimizing the contrastive loss, the Clap module learns to map semantically related text and audio to similar vector representations.

[0029] Furthermore, the text prompt is input into the Clap module of the trained audio generation model, and the trained audio generation model is used to generate ship audio; specifically:

[0030] First, the text prompt is input into the text encoder of the Clap module, and the text encoder converts the text prompt into a text encoding;

[0031] Then, the text encoding result is used as a condition to input into the Diffusion module. The Diffusion module uses the text encoding as a guidance and generates an audio latent representation matching the text prompt through the trained iterative denoising process;

[0032] After that, the audio latent representation is input into the decoder of the VAE module, and the decoder converts the audio latent representation into a mel spectrogram;

[0033] Finally, the vocoder converts this mel spectrogram into an audible audio waveform.

[0034] After adopting the above technical solution, the present invention introduces and applies the Variational Autoencoder (VAE), Contrastive Language-Audio Pretraining (CLAP), and Diffusion Model (Diffusion) technologies. The VAE technology compresses audio data into a low-dimensional latent representation, helping the model learn the core features of the data, reducing storage and computational requirements. At the same time, VAE uses self-supervised learning and does not require a large amount of labeled data, and can learn useful feature representations from unlabeled audio. Then, the Clap technology maps text and audio to the same semantic space through contrastive learning, achieving cross-modal alignment and effectively associating audio and text. This enables the model to understand the correspondence between text descriptions and audio content, providing a basis for generating audio that matches the text. Finally, combined with the Diffusion technology, the data is gradually refined by adding noise step by step and then denoising iteratively, which helps to capture the details of the audio and generate natural and realistic audio samples. With the guidance of text conditions, the present invention can generate natural and realistic audio content according to given text prompts, not only improving the quality and realism of audio generation, but also providing strong support for deep learning tasks related to ship audio, effectively solving the problem of insufficient existing datasets. Description of the Drawings

[0035] Figure 1 It is a flowchart of the ship audio generation method based on deep learning of the present invention;

[0036] Figure 2 It is a flowchart of training the audio generation model of the present invention using the preprocessed dataset;

[0037] Figure 3 It is a structural diagram of the VAE module of the present invention;

[0038] Figure 4 It is a structural diagram of the Clap module of the present invention;

[0039] Figure 5 It is a structural diagram of the Diffusion module of the present invention. Detailed Embodiments

[0040] In order to make the content of the present invention easier to be clearly understood, the present invention will be further described in detail below according to specific embodiments in conjunction with the drawings.

[0041] As Figure 1 and Figure 2 shown, a ship audio generation method based on deep learning includes:

[0042] Step S1, obtaining a ship audio set, preprocessing all ship audios in the ship audio set to obtain Mel spectrograms corresponding to each ship audio;

[0043] Step S2, construct an audio generation model. The audio generation model includes a VAE module (Variational Autoencoder), a Diffusion module, and a Clap module (Contrastive Language-Audio Pretraining model). Use the preprocessed dataset to train the audio generation model to obtain a trained audio generation model. Among them,

[0044] During the training process, the VAE module encodes the Mel spectrogram to obtain the corresponding potential features of the ship audio; the Clap module maps the ship audio and the relevant text description to the same semantic space to obtain the association between the ship audio and the text description; during the forward propagation of the potential features, the Diffusion module gradually adds Gaussian noise to the potential features. During the backward propagation of the potential features, the Diffusion module uses the text encoding processed by the Clap module as a condition to guide the gradual denoising and restore the potential features; the VAE module then converts the restored potential features back to the Mel spectrogram;

[0045] Step S3, input the text prompt into the Clap module of the trained audio generation model, and use the trained audio generation model to generate ship audio.

[0046] In this embodiment, in step S1, the specific steps of preprocessing are as follows:

[0047] First, cut the ship audio into segments with a preset duration and store them by category. The preset duration can be 10s;

[0048] Then load and resample the ship audio to a unified sampling rate;

[0049] Then transform it to the frequency domain through short-time Fourier transform and convert the amplitude value to the decibel value;

[0050] Finally, calculate the Mel spectrogram using the decibel value;

[0051] That is:

[0052] x resampled = Resample(x original , f s,original , f s,new )

[0053] X = STFT(x resampled , window, hop length)

[0054]

[0055] M = MelSpectrogram(X dB , n fft , f s,new , nmels )

[0056] In the formula, Resample() represents the resampling operation; x original represents the original audio signal (time-domain signal); f s,original represents the sampling rate of the original signal (unit: Hz); f s,new represents the target sampling rate (unit: Hz);

[0057] STFT() represents the short-time Fourier transform; x resampled represents the resampled audio signal as the input; Window represents the window function (such as Hanning window, Hamming window, etc.) used to segment the signal; Hop length represents the frame shift length (unit: number of samples), indicating the time interval between adjacent frames;

[0058] X dB represents the spectrum in decibels as the input; X represents the complex spectrum matrix of STFT; ∥X∥ represents the amplitude (magnitude) of the spectrum; ref represents the reference value (usually 1 or the maximum amplitude of the signal) used for normalization;

[0059] MelSpectrogram() represents the operation of converting to a Mel spectrogram; n fft represents the number of points of the fast Fourier transform (FFT), which determines the resolution of the spectrum; n mels represents the number of Mel filters used to map the spectrum to the Mel scale;

[0060] In this embodiment, as Figure 3 shown, the VAE module includes an encoder, a latent space, and a decoder; the encoder is used to map the Mel spectrogram to the parameters in the latent space (i.e., the mean μ and the logarithm of the variance σ 2 ) with the aim of learning the distribution of the input data in the latent space; the latent space is used to generate latent features based on the parameters mapped to it by the encoder (sampling a noise vector ∈ from the standard normal distribution and then generating the latent variable z through the mean μ and the standard deviation σ output by the encoder, so that the gradient can be backpropagated through the random sampling process, enabling the VAE module to be trained by gradient descent); the decoder is used to convert the latent features processed by the Diffusion module back to the Mel spectrogram. During the training process, the working process of the VAE module is expressed as:

[0061] μ, log(σ 2 ) = Encoder(x)

[0062] z = μ + σ ⊙ ∈

[0063]

[0064] In the formula, x represents the input data of the encoder, i.e., the original Mel spectrogram; the logarithm of the mean μ and variance σ 2 are both parameters for the encoder to map the input data x to the latent space; ∈ represents a noise vector sampled by the VAE from the standard normal distribution; z represents the latent feature generated in the latent space; z d represents the latent feature processed by the Diffusion module; represents the Mel spectrogram obtained by the decoder conversion.

[0065] In this embodiment, as Figure 4 shown, the Clap module includes a text encoder, an audio encoder, and a projection layer. The core goal of CLAP is to enable the model to understand and associate text and audio in the same semantic space by contrastive learning of the joint representation between text and audio, so as to provide rich semantic information for the audio generation task. During the training process, the working process of the Clap module is as follows:

[0066] First, the text is converted into a high-dimensional vector representation through the text encoder to capture the text semantic information; the Mel spectrogram is converted into a high-dimensional vector representation through the audio encoder to capture the audio acoustic features; among them,

[0067] before the input text enters the text encoder, operations such as word segmentation, stop word removal, and lowercasing are performed on the input text to adapt to the input of the text encoder.

[0068] Then, contrastive learning is carried out, and the matching text-audio pairs are selected as positive sample pairs, and the non-matching text-audio pairs are selected as negative sample pairs, and the cosine similarity is used to calculate the similarity between the positive sample pairs and the negative sample pairs respectively;

[0069]

[0070] In the formula, z txt represents the embedding vector of the text (text feature representation); z aud represents the embedding vector of the audio (audio feature representation); ∥·∥ represents the norm (magnitude) of the vector, indicating the size of the vector; Similarity(z txt ,z aud ) represents calculating the similarity between the text and audio embedding vectors.

[0071] Finally, a contrastive loss function is used for training to make the similarity of the positive sample pairs greater than that of the negative sample pairs. By minimizing the contrastive loss, the Clap module learns to map semantically related text and audio to similar vector representations. At a certain level of the Clap module, the feature vectors of text and audio are fused so that the model can learn cross-modal associations.

[0072]

[0073] In the formula, θ txt represents the parameters of the text encoder; θ aud represents the parameters of the audio encoder; represents the parameters of the optimized text encoder; represents the parameters of the optimized audio encoder; L contrastive represents the contrast loss function;

[0074] This formula indicates that the parameters of the text encoder and the audio encoder are optimized by minimizing the contrast loss function.

[0075] In summary, the Clap module achieves an effective mapping between text and audio through a contrastive learning framework, providing a strong semantic alignment ability for audio generation tasks. The text representation provided by the Clap module is used as conditional information to guide the Diffusion model to generate audio that matches the text description.

[0076] In this embodiment, as Figure 5 shown, the Diffusion module converts data into high-dimensional noise by gradually adding noise, and then recovers the original data from the noise through an iterative denoising process. This process can be divided into two main stages: the forward diffusion process and the reverse denoising process.

[0077] In the forward diffusion process, the model gradually converts structured data into unordered noise. The process can be expressed as:

[0078]

[0079] In the formula, x0 is the original data, x 1:T is a series of intermediate states from x0 to the noise data x T , and q(x t |x t-1 ) is the probability distribution of adding noise at each step

[0080] The reverse denoising process is the inverse of the forward diffusion process, and the goal is to recover the original data from the noise data. The condition in this step is generated by Clap to guide the model to denoise. The process can be expressed as:

[0081]

[0082] In the formula, p θ (x t-1 |x t ) is the probability distribution of recovering the previous state x t given the current noise state x t-1 , and θ represents the model parameters.

[0083] The training objective of the Diffusion module is to maximize the Evidence Lower Bound (ELBO). The process can be expressed as follows:

[0084]

[0085] By maximizing the ELBO, the Diffusion module learns how to generate data similar to the original data.

[0086] The purpose of the forward diffusion process is to gradually add noise to the data until the data is completely transformed into noise. This process can be regarded as a Markov chain, where each step follows a Gaussian distribution and the final state is close to the standard normal distribution. The main role of the forward diffusion process is to simulate the diffusion of data from an ordered state to a disordered state, providing a basis for the subsequent reverse diffusion process.

[0087] The purpose of the reverse diffusion process is to recover the original data from the noise state. This process is also a parameterized Markov chain. By gradually removing the noise, a series of intermediate states are generated, and finally the original data is recovered. The main role of the reverse diffusion process is to generate new data samples, and these samples have a similar distribution to the training data.

[0088] The forward diffusion process and the reverse diffusion process together constitute the complete life cycle of the diffusion model. The forward diffusion process gradually adds noise to make the data gradually become a Gaussian noise distribution, while the reverse diffusion process gradually removes the noise to recover the original data from the Gaussian noise. These two processes enable the diffusion model to generate complex high-quality data samples from a simple noise distribution.

[0089] The Diffusion module recovers the audio that matches the text description from the noisy signal through a step-by-step denoising iteration process. This method ensures the coherence and smoothness of the latent space, and then produces high-quality and high-fidelity audio content.

[0090] In this embodiment, the specific steps of step S3 are as follows:

[0091] First, the text prompt is input into the text encoder of the Clap module, and the text encoder converts the text prompt into a text encoding;

[0092] Then, the text encoding result is used as a condition to input into the Diffusion module. The Diffusion module uses the text encoding as a guide and generates an audio latent representation that matches the text prompt through a trained iterative denoising process;

[0093] After that, the audio latent representation is input into the decoder of the VAE module, and the decoder converts the audio latent representation into a mel spectrogram;

[0094] Finally, the vocoder converts this mel spectrogram into an audible audio waveform.

[0095] The advantages of the solutions involved in the above embodiments will be described below in combination with specific experimental evaluations.

[0096] The present invention compares the trained model with the available text-to-audio (T2A) generation model DiffSound, and evaluates its performance through several indicators such as FID, KL divergence, MOS-Q, and MOS-F. The comparison results are shown in Table 1:

[0097] Table 1

[0098] Model Params FID KL MOS-Q MOS-F Diffsound 520M 7.17 3.57 67.1±1.03 70.9±1.05 OurModel 332M 4.61 2.79 72.5±0.90 78.6±1.01

[0099] Among them, FID is an indicator that measures the perceptual similarity between the generated image (or audio) and the real image (or audio). The lower the FID value, the closer the generated audio is to the real audio in terms of perceptual quality. KL divergence is an indicator that measures the difference between two probability distributions. In the context of audio generation, it is usually used to measure the difference between the distribution of the generated audio and the distribution of the real audio. The lower the KL divergence, the closer the distribution of the generated audio is to the distribution of the real audio. MOS is a subjective evaluation indicator that measures the quality or characteristics of audio through user ratings. MOS-Q (Quality) specifically scores the quality of audio, measuring the clarity, naturalness, etc. of the audio; MOS-F (Fidelity) scores the consistency between the audio and the text description, measuring to what extent the generated audio is faithful to its corresponding text description.

[0100] According to the experimental results, it can be analyzed that the audio generation model in the present invention shows significant performance advantages in the comparison with Diffusion, both in terms of objective audio quality evaluation and subjective sound quality and fidelity evaluation. This demonstrates the strong generalization ability of the model and its potential in ship audio generation, which can effectively solve the dilemma of insufficient existing datasets and provide strong support for other ship audio-related deep learning tasks. These experimental results confirm the effectiveness of the method in this paper, indicating its potential and adaptability in practical applications.

[0101] Taking the above ideal embodiments based on the present invention as an inspiration, through the above description, relevant staff can completely make various changes and modifications without departing from the technical idea of this invention. The technical scope of this invention is not limited to the content in the specification, and its technical scope must be determined according to the scope of the claims.

Claims

1. A method for generating ship audio based on deep learning, characterized in that it includes: Obtain a ship audio set, preprocess all the ship audios in the ship audio set to obtain the Mel spectrogram corresponding to each ship audio; Construct an audio generation model. The audio generation model includes a VAE module, a Diffusion module, and a Clap module. Use the preprocessed data set to train the audio generation model to obtain a trained audio generation model; where During the training process, the VAE module encodes the Mel spectrogram to obtain the latent features of the corresponding ship audio; the Clap module maps the ship audio and the relevant text description to the same semantic space to obtain the association between the ship audio and the text description; during the forward propagation of the latent features, the Diffusion module gradually adds Gaussian noise to the latent features, and during the backward propagation of the latent features, the Diffusion module uses the text encoding processed by the Clap module as a condition to guide the gradual denoising and restore the latent features; the VAE module then converts the restored latent features back to the Mel spectrogram; Input the text prompt into the Clap module of the trained audio generation model, and use the trained audio generation model to generate ship audio.

2. The method for generating ship audio based on deep learning according to claim 1, characterized in that For the preprocessing of ship audio, the specific steps are: First, cut the ship audio into segments with a preset duration and store them by category; Then load and resample the ship audio to a unified sampling rate; Then convert it to the frequency domain through short-time Fourier transform, and convert the amplitude value to decibel value; Finally, calculate the Mel spectrogram using the decibel value.

3. The method for generating ship audio based on deep learning according to claim 1, characterized in that The VAE module includes an encoder, a latent space, and a decoder; where The encoder is used to map the Mel spectrogram to the parameters of the latent space; The latent space is used to generate latent features based on the parameters mapped to it by the encoder; The decoder is used to convert the latent features processed by the Diffusion module back to the Mel spectrogram.

4. The method for generating ship audio based on deep learning according to claim 3, characterized in that During the training process, the working process of the VAE module is expressed as: μ, log(σ 2 ) = Encoder(x) z = μ + σ ⊙ ∈ Wherein, X represents the input data of the encoder; the logarithm of the mean μ and the variance σ 2 are both parameters for the encoder to map the input data x to the latent space; ∈ represents a noise vector sampled by the VAE from the standard normal distribution; z represents the latent feature generated in the latent space; z d represents the latent feature after being processed by the Diffusion module.

5. The method for generating ship audio based on deep learning according to claim 1, characterized in that During the training process, the working process of the Clap module is: First, convert the text into a high-dimensional vector representation through a text encoder to capture the semantic information of the text; convert the Mel spectrogram into a high-dimensional vector representation through an audio encoder to capture the acoustic features of the audio; Then, perform a contrast study, select the matching text-audio pair as the positive sample pair, the non-matching text-audio pair as the negative sample pair, and calculate the similarity between the positive sample pair and the negative sample pair using the cosine similarity respectively. Finally, it is trained using a contrastive loss function to make the similarity of positive sample pairs greater than that of negative sample pairs. By minimizing the contrastive loss, the Clap module learns to map semantically related text and audio to similar vector representations.

6. The method for generating ship audio based on deep learning according to claim 1, wherein Input the text prompt into the Clap module of the trained audio generation model, and use the trained audio generation model to generate ship audio; specifically: First, the text prompt is input into the text encoder of the Clap module, and the text encoder converts the text prompt into a text encoding. Then, the text encoding result is used as a condition to input into the Diffusion module. The Diffusion module uses the text encoding as a guide and generates an audio latent representation that matches the text prompt through the trained iterative denoising process. After that, the audio latent representation is input into the decoder of the VAE module, and the decoder converts the audio latent representation into a mel spectrogram. Finally, the vocoder converts this mel spectrogram into an audible audio waveform.

Citation Information

Cited By

  • Small sample ship voiceprint data enhancement and high-precision identification method

    CN121459849A

  • Small sample ship soundprint data enhancement and high-precision identification method

    CN121459849B