Ship classification method, device and equipment based on deep learning

Through deep learning methods, the diffusion model and proxy attention mechanism are used to optimize feature extraction, which solves the problems of insufficient accuracy and generalization ability of ship audio classification and achieves more efficient ship sound recognition.

CN120632680APending Publication Date: 2025-09-12CHANGZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510717227.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing technologies in ship audio classification have problems with low classification accuracy and insufficient generalization ability, especially in the changing ocean environment, it is difficult to effectively distinguish sub-categories of ship sounds.

Method used

A deep learning-based ship classification method is adopted. A diffusion model is used to generate a spectrogram matching the text prompt. The agent attention mechanism and multi-dimensional feature extraction are introduced. The residual block and Droppath iterative training strategy are combined to optimize feature extraction and data enhancement.

Benefits of technology

The accuracy and generalization ability of ship sound classification have been significantly improved, which enables better identification of different types of ship sounds and improves the performance of the model in complex marine environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120632680A_ABST
    Figure CN120632680A_ABST
Patent Text Reader

Abstract

The invention relates to the field of ship classification, in particular to a ship classification method, device and equipment based on deep learning. The method comprises the following steps: S1, preprocessing ship audio in a ship audio data set to obtain a spectrogram X; a pre-trained diffusion model is utilized to generate a spectrogram G matched with the text prompt, the spectrogram G and the corresponding spectrogram X are respectively introduced into masks and fused, an enhanced spectrogram Xaug is obtained, and an enhanced spectrogram set is constructed; s2, constructing a ship classification model, extracting features from a time dimension and a frequency dimension, fusing the features, and introducing an agency attention mechanism to enhance key features; s3, training a ship classification model by using the enhanced spectrum image set; and S4, preprocessing the audio of the ship to be classified to obtain a spectrogram X to be classified, and inputting the spectrogram X to be classified into the trained ship classification model to perform ship classification. According to the method, the accuracy and generalization ability of ship sound classification can be remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of ship classification, and in particular to a ship classification method, device and equipment based on deep learning. Background Art

[0002] With the rapid development of the shipping industry, accurate and efficient ship classification is crucial for maintaining maritime traffic order, ensuring maritime safety, improving port operations efficiency, and performing environmental monitoring. Ship classification not only involves ship supervision and identification but also connects to multiple aspects such as maritime search and rescue, route planning, and marine resource management. With the increasing frequency of maritime activities, traditional ship identification methods that rely on images and radar are often limited in volatile ocean environments and adverse weather conditions. Audio-based ship identification technology, however, is gaining attention due to its unique advantages.

[0003] However, ship sound classification in the field of audio signal processing faces technical challenges. Although the development of deep learning technology has brought new opportunities for automatic feature extraction and audio processing, traditional audio classification methods are often only able to distinguish very different types of audio, such as dog barking and engine roaring, which do not belong to the same general category. For subcategories within the same general category, the classification accuracy is very low. At the same time, most existing technologies use a single convolutional neural network (CNN) structure, which is too simple to fully capture the complex patterns and timing information in audio signals with very small differences, such as ship audio. This limits the model's classification accuracy and generalization ability. Therefore, a more advanced network structure design is needed to improve the performance of ship sound classification, especially in the changing real-world ocean environment. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to overcome the defects of the existing technology and provide a ship classification method based on deep learning, which can significantly improve the accuracy and generalization ability of ship sound classification.

[0005] In order to solve the above technical problems, the technical solution of the present invention is: a ship classification method based on deep learning, comprising:

[0006] Step S1, processing the ship audio dataset to obtain an enhanced spectrum atlas;

[0007] Step S11, preprocessing the ship audio in the ship audio dataset to obtain a spectrogram X;

[0008] Step S12: Use the pre-trained diffusion model to generate a spectrogram G that matches the text prompt, and introduce the spectrogram G and the corresponding spectrogram X into the mask and fuse them to obtain the enhanced spectrogram X. aug; Among them, the text prompt comes from the label;

[0009] Step S13, multiple spectrum graphs X aug Constructed into an enhanced spectrum atlas;

[0010] Step S2: constructing a ship classification model, wherein the ship classification model extracts and fuses features from the time dimension and frequency dimension respectively, and also introduces a proxy attention mechanism to enhance key features;

[0011] Step S3, training the ship classification model using the enhanced spectrum atlas to obtain a trained ship classification model;

[0012] Step S4: pre-process the audio of the ship to be classified to obtain a spectrogram X to be classified, and input the spectrogram X to be classified into the trained ship classification model to perform ship classification.

[0013] Furthermore, the ship audio is pre-processed, specifically including:

[0014] First, cut the ship audio into segments of preset duration and store them by category;

[0015] The ship audio is then loaded and resampled to a uniform sampling rate;

[0016] Then convert it to the frequency domain through short-time Fourier transform, and convert the amplitude value into decibel value;

[0017] Finally, the decibel value is used to calculate the spectrum graph X.

[0018] Furthermore, the diffusion model includes a VAE module, a Diffusion module, and a Clap module. In the process of training the diffusion model:

[0019] The VAE module encodes the spectrogram X to obtain the corresponding ship audio potential features;

[0020] The Clap module maps the ship audio and related text prompts into the same semantic space to obtain the association between the ship audio and the text description;

[0021] During the forward propagation of the latent features, the Diffusion module gradually adds Gaussian noise to the latent features. During the backward propagation of the latent features, the Diffusion module uses the text encoding processed by the Clap module as a condition to guide the gradual denoising and restore the latent features.

[0022] The VAE module then converts the recovered latent features back to the spectrogram G.

[0023] Furthermore, the spectrum map G and the spectrum map X are introduced into the mask and fused to obtain the enhanced spectrum map X aug; Specifically include:

[0024] Mask the spectrogram X and the spectrogram G using a mask randomly selected from the predefined n masks, and then concatenate the spectrogram X and the spectrogram G; the formula is:

[0025] X aug =α·(X☉M i )+(1-α)·(G☉(1-M i ))

[0026] Where α is the mixing coefficient, which is used to control the weight of the spectrogram X and the spectrogram G after fusion; ⊙ is the element-by-element multiplication, which is used to apply the mask to the spectrogram; i is a randomly selected mask index, i∈[1,n]).

[0027] Furthermore, features are extracted and fused from the time dimension and frequency dimension respectively, including:

[0028] For the spectrum graph input to the ship classification model, average pooling and convolution operations are performed in the Y direction to extract features in the time dimension; average pooling and convolution operations are performed in the X direction to extract features in the frequency dimension;

[0029] The features extracted from the time dimension and frequency dimension are subjected to depth-wise separable convolution, group normalization and Sigmoid activation function respectively, and finally element-wise multiplication operation is performed to obtain the final feature representation.

[0030] Furthermore, the ship classification model also introduces multiple residual blocks connected in sequence; among them,

[0031] The first residual block takes the features obtained by fusing the features extracted from the time dimension and the frequency dimension as input;

[0032] In the process of training the ship classification model, the Droppath iterative training strategy is adopted.

[0033] The present invention also relates to a ship classification device based on deep learning, comprising:

[0034] The dataset processing module is used to process the ship audio dataset to obtain an enhanced spectrogram set. The specific process is as follows: pre-process the ship audio in the ship audio dataset to obtain a spectrogram X; use the pre-trained diffusion model to generate a spectrogram G that matches the text prompt, and introduce the spectrogram G and the corresponding spectrogram X into the mask and fuse them to obtain the enhanced spectrogram X. aug ; Multiple spectrograms X aug Constructed into an enhanced spectrum atlas; where the text prompts come from the labels;

[0035] The classification model construction module is used to build a ship classification model. The ship classification model extracts and fuses features from the time dimension and frequency dimension respectively, and also introduces a proxy attention mechanism to enhance key features.

[0036] A training module is used to train the ship classification model using the enhanced spectrum atlas to obtain a trained ship classification model;

[0037] The classification module is used to pre-process the audio of the ship to be classified to obtain the spectrum map X to be classified, and input the spectrum map X to be classified into the trained ship classification model to perform ship classification.

[0038] The present invention also relates to a device comprising:

[0039] memory for storing computer programs;

[0040] A processor is used to execute the computer program, and when the computer program is executed by the processor, the steps of the ship classification method based on deep learning are implemented.

[0041] The present invention also relates to a readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the ship classification method based on deep learning.

[0042] The present invention also relates to a computer program product, comprising a computer program, which, when executed by a processor, implements the steps of the ship classification method based on deep learning.

[0043] After adopting the above technical solution, the present invention improves data enhancement technology through the diffusion model, optimizes feature extraction, and integrates the proxy attention mechanism. Specifically, the method first uses the diffusion model to generate a spectrogram that matches the text prompt, and then masks it and fuses it with the original spectrogram to form a hybrid image, thereby enhancing data diversity and ensuring label consistency. Then, the improved feature extraction module can extract features from the time dimension and frequency dimension respectively to obtain more detailed multi-dimensional feature information. In addition, the introduction of the proxy attention mechanism enables the model to focus on the key parts of the audio features, improving the recognition ability of different types of ship sounds. Therefore, the present invention significantly improves the accuracy and generalization ability of ship audio classification, and provides an efficient and reliable solution for ship classification tasks in actual marine environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 Flowchart of the ship classification method based on deep learning of the present invention;

[0045] Figure 2 This is a flow chart of the ship classification model training of the present invention;

[0046] Figure 3 This is a framework diagram of the diffusion model and spectrum graph fusion of the present invention;

[0047] Figure 4 It is a framework diagram of the feature extraction module of the present invention;

[0048] Figure 5 This is a framework diagram of the proxy attention mechanism of the present invention;

[0049] Figure 6 This is a visualization of the loss and accuracy of the method in the present invention and the traditional CNN model during training. DETAILED DESCRIPTION

[0050] In order to make the contents of the present invention more clearly understood, the present invention is further described in detail below based on specific embodiments in conjunction with the accompanying drawings.

[0051] like Figure 1 and Figure 2 As shown, a ship classification method based on deep learning includes:

[0052] Step S1, processing the ship audio dataset to obtain an enhanced spectrum atlas;

[0053] Step S11, preprocessing the ship audio in the ship audio dataset to obtain a spectrogram X;

[0054] Step S12: Use the pre-trained diffusion model to generate a spectrogram G that matches the text prompt, and introduce the spectrogram G and the corresponding spectrogram X into the mask and fuse them to obtain the enhanced spectrogram X. aug ; Among them, the text prompts come from the labels of the corresponding ship audio in the ship audio dataset;

[0055] Step S13, multiple spectrum graphs X aug Constructed into an enhanced spectrum atlas;

[0056] Step S2: constructing a ship classification model. The ship classification model extracts and fuses features from both the time and frequency dimensions, and also introduces a proxy attention mechanism to enhance key features.

[0057] Step S3, training the ship classification model using the enhanced spectrum atlas to obtain a trained ship classification model;

[0058] Step S4: pre-process the audio of the ship to be classified to obtain a spectrogram X to be classified, and input the spectrogram X to be classified into the trained ship classification model to perform ship classification.

[0059] Specifically, this embodiment combines the diffusion model to generate a spectrum that matches the text prompt, and then constructs a fused spectrum. In this way, the diversity of the data is greatly enriched, while ensuring the consistency of the labels, thereby providing more diverse and high-quality training samples for the model training process. The feature extraction is optimized and improved so that it has the ability to extract features from both the time dimension and the frequency dimension, thereby obtaining more detailed and comprehensive multi-dimensional features. This enables the model to learn the inherent characteristics of ship audio in a deeper and more comprehensive way, thereby effectively improving its ability to distinguish different types of ship sounds. By introducing the proxy attention mechanism, the model can focus on the key parts of the audio features and enhance the recognition ability of different types of ship sounds. Therefore, this embodiment can significantly improve the accuracy and generalization ability of ship audio classification.

[0060] In this embodiment, the ship classification model can also introduce multiple residual blocks connected in sequence; wherein,

[0061] The first residual block takes the features obtained by fusing the features extracted from the time dimension and the frequency dimension as input; in the process of training the ship classification model, the Droppath iterative training strategy is adopted.

[0062] Specifically, residual connections allow the model to transfer information directly between layers. This design helps the gradient flow in the network, especially in deep networks. The model contains multiple residual blocks, each of which performs a convolution operation to extract features, followed by a batch normalization layer to adjust the distribution of features. Finally, the ReLU activation function is applied to increase the nonlinear expression ability of the model. The design of these residual blocks helps alleviate the gradient vanishing problem in deep network training and ensures that information and gradients can flow effectively in the network. In addition, the use of a residual discarding path strategy to optimize the network structure can enhance the efficiency of gradient propagation and feature transfer, thereby improving the performance of the model in deep networks.

[0063] like Figure 1 As shown in the figure, the entire ship classification model works as follows: a feature extraction module extracts and fuses features from both the time and frequency dimensions. The fused features serve as the input to the first residual block. After a convolution operation is performed on the output of the last residual block, a proxy attention mechanism is introduced to enhance key features. A global average pooling operation is then performed, and the features are fused with the features fed into the proxy attention mechanism. Finally, a fully connected operation is performed to achieve ship classification based on ship audio. The fusion of local features extracted by conventional convolutional blocks and global information extracted by the proxy attention module enhances the expressiveness of features, thereby improving classification accuracy.

[0064] Introduction to Droppath iterative training strategy:

[0065] The training strategy of Droppath iteration, namely Residual Droppath, is a training method used to enhance feature reuse in residual connections. It enhances the model's feature reuse capabilities by randomly dropping some layers. Specifically, in each training iteration, the system randomly selects a portion of layers as dropouts, and the outputs of these layers are temporarily set to zero or ignored. In this way, the model has to rely on the features of the remaining layers for learning and prediction during the forward propagation phase. This random dropout operation can encourage the model to make more full use of the features of the non-dropped layers, thereby enhancing feature reuse and improving the generalization ability of the model. When calculating the loss, the output of the model may be affected because the outputs of some layers are dropped, but this effect can be compensated by adjusting the model parameters. During the backpropagation process, the parameters of the dropped layers are not updated because their outputs are set to zero or ignored.

[0066] The model then enters the stage of training the dropout portion. During this stage, the previously dropped layers will be reactivated and participate in training, while the parameters of the non-dropped layers are frozen and no longer updated. The purpose of this is to allow the model to focus on training those previously dropped layers to compensate for the impact of being dropped in the Droppath iteration. During forward propagation, the non-dropped layers are frozen, and the output of the model mainly depends on the features of the dropped layers. In this way, the model can better learn the features of these layers and further enhance feature reuse. During backpropagation, the parameters of the dropped layers will be updated, while the parameters of the non-dropped layers remain frozen and will not be updated. This training strategy helps the model better learn to train with feature reuse, thereby further improving the performance and generalization ability of the model.

[0067] The core idea of ​​the Droppath iterative training strategy is to force the ship classification model to learn more effective feature representations by randomly dropping the residual connections of some layers during training. The specific implementation steps are as follows:

[0068] After the feature extraction module extracts the time and frequency features and fuses them, the model enters the residual block iteration stage. In each residual block, the residual block first extracts features through the convolution layer, and then combines the input feature E with the extracted feature E. step To enhance the generalization ability of the model and prevent overfitting, the model randomly discards some residual connections during training.

[0069] Specifically, if the discard stage M of the current iteration is even, a mask mask is randomly generated d , and update E to:

[0070] E=E+E step ⊙mask d

[0071] If M is odd, update E to:

[0072] E=E+E step .detach()☉mask d +E step ☉(1-mask d )

[0073] Among them, detach() is a method in PyTorch, which is used to separate the tensor from the current computational graph so that it no longer participates in gradient calculations.

[0074] In this way, the model is forced to learn more effective feature representations during training. Finally, the final output is processed by the post-processing layer to obtain the model's prediction results. The loss function L is used to calculate the loss between the predicted result and the true label Y, and backpropagation is performed to update the model parameters.

[0075] The residual dropout path randomly drops some residual connections during training, forcing the model to learn more effective feature representations, thereby enhancing its generalization and robustness. This random dropout mechanism prevents the model from over-reliance on a specific residual path, promotes feature diversification and reuse, and helps the model function effectively across diverse network architectures. Furthermore, the residual dropout path introduces an implicit regularization that optimizes the model training process, reduces the risk of overfitting, and enables the model to better maintain performance despite input variations.

[0076] In this embodiment, the ship audio is pre-processed, specifically including:

[0077] First, cut the ship audio into segments of preset duration and store them by category;

[0078] The ship audio is then loaded and resampled to a uniform sampling rate;

[0079] Then convert it to the frequency domain through short-time Fourier transform, and convert the amplitude value into decibel value;

[0080] Finally, the decibel value is used to calculate the spectrum graph X; the specific formula is:

[0081] x resampled =Resample(x original , f s,original , f s,new )

[0082] X = STFT(x resampled, window, hop length)

[0083]

[0084] M=MelSpectrogram(X dB , n fft , f s,new , n mels )

[0085] Where Resample() represents the resampling operation; x original Represents the original audio signal (time domain signal); f s,original Indicates the sampling rate of the original signal (unit: Hz); f s,new Indicates the target sampling rate (unit: Hz);

[0086] STFT() stands for short-time Fourier transform; x resampled Represents the input resampled audio signal; Window represents the window function (such as Hanning window, Hamming window, etc.), which is used to segment the signal; Hoplength represents the frame shift length (unit: number of samples), which represents the time interval between adjacent frames;

[0087] X dB represents the input spectrum expressed in decibels; X represents the complex spectrum matrix of STFT; ||X|| represents the amplitude (modulus) of the spectrum; ref represents the reference value (usually 1 or the maximum amplitude of the signal) used for normalization;

[0088] MelSpectrogram() means converting to Mel spectrum graph operation; fft Indicates the number of points of the Fast Fourier Transform (FFT), which determines the resolution of the spectrum; n mels Indicates the number of Mel filters used to map the spectrum to the Mel scale.

[0089] Specifically, the audio preprocessing steps effectively extract the key features of the audio signal. These operations not only ensure the consistency of the audio data, but also enable the model to better simulate the characteristics of the human auditory system by converting the audio signal into a Mel-spectrogram, thereby extracting richer and more discriminative features.

[0090] In this embodiment, if Figure 2 and Figure 3 As shown in Figure 3, the diffusion model includes a VAE module, a Diffusion module, and a Clap module.

[0091] The process of training the diffusion model specifically includes:

[0092] The VAE module encodes the spectrogram X to obtain the corresponding ship audio potential features;

[0093] The Clap module maps the ship audio and related text prompts into the same semantic space to obtain the association between the ship audio and the text description;

[0094] During the forward propagation of the latent features, the Diffusion module gradually adds Gaussian noise to the latent features. During the backward propagation of the latent features, the Diffusion module uses the text encoding processed by the Clap module as a condition to guide the gradual denoising and restore the latent features.

[0095] The VAE module then converts the recovered latent features back to the spectrogram G.

[0096] In this embodiment, a pre-trained diffusion model is used to generate a spectrogram G that matches the text prompt. Specifically, the process includes:

[0097] First, the text prompt is input into the text encoder of the Clap module, and the text encoder converts the text prompt into a text encoding;

[0098] The text encoding result is then used as a conditional input to the Diffusion module. The Diffusion module uses the text encoding as a guide to generate an audio latent representation that matches the text prompt through a trained iterative denoising process.

[0099] After that, the audio potential representation is input into the decoder of the VAE module, and the decoder converts the audio potential representation into a spectrogram G;

[0100] Specifically, conditional cues and diffusion models are used to generate images that are semantically consistent with the original images, avoiding the label confusion problem.

[0101] The following is a detailed introduction to the various modules of the diffusion model.

[0102] About VAE module: Figure 2 and Figure 3 As shown, the VAE module includes an encoder, a latent space, and a decoder; the encoder is used to map the spectrogram X to the parameters of the latent space (i.e., mean μ and variance σ 2The goal is to learn the distribution of input data in the latent space; the latent space is used to generate latent features based on the parameters mapped to it by the encoder (sampling a noise vector ∈ from a standard normal distribution, and then generating the latent variable z through the mean μ and standard deviation σ of the encoder output. This process allows the gradient to be backpropagated through the random sampling process, so that the VAE module can be trained through gradient descent); the decoder is used to convert the latent features processed by the Diffusion module back into the Mel-spectrogram. During training, the working process of the VAE module is expressed as:

[0103] μ,log(σ 2 )=Encoder(x)

[0104] z=μ+σ⊙∈

[0105]

[0106] Where x represents the input data of the encoder, that is, the spectrum graph X; the mean μ and variance σ 2 The logarithms of are the parameters of the encoder that maps the input data x to the latent space; ∈ represents a noise vector sampled by VAE from a standard normal distribution; z represents the potential features generated by the latent space; z d Represents the potential features after processing by the Diffusion module; Represents the mel-spectrogram converted by the decoder.

[0107] About Clap module: Figure 2 and Figure 3 As shown in Figure 1, the Clap module consists of a text encoder, an audio encoder, and a projection layer. The core goal of CLAP is to learn the joint representation between text and audio through comparison, enabling the model to understand and associate text and audio in the same semantic space, thereby providing rich semantic information for audio generation tasks. During training, the Clap module works as follows:

[0108] First, the text is converted into a high-dimensional vector representation through a text encoder to capture the text semantic information; the Mel spectrum map is converted into a high-dimensional vector representation through an audio encoder to capture the audio acoustic features;

[0109] Before the text is input into the text encoder, operations such as word segmentation, stop word removal, and lowercase are performed on the input text to adapt it to the text encoder input.

[0110] Then, contrastive learning is performed, matching text-audio pairs are selected as positive sample pairs, and mismatched text-audio pairs are selected as negative sample pairs, and the cosine similarity is used to calculate the similarity between the positive sample pairs and the negative sample pairs respectively;

[0111]

[0112] Where z txt Embedding vector representing text (text feature representation); z aud Represents the embedded vector of the audio (audio feature representation); ||·|| represents the modulus (norm) of the vector, which represents the size of the vector; Similarity (z txt ,z aud ) represents the calculation of the similarity between the text and audio embedding vectors.

[0113] Finally, a contrastive loss function is used for training to make the similarity of positive sample pairs greater than that of negative sample pairs. By minimizing the contrastive loss, the Clap module learns to map semantically related text and audio to similar vector representations. At a certain level of the Clap module, the feature vectors of text and audio are fused so that the model can learn cross-modal associations.

[0114]

[0115] Where θ txt represents the parameters of the text encoder; θ aud Represents the parameters of the audio encoder; Represents the parameters of the optimized text encoder; Represents the parameters of the optimized audio encoder; L contrastive represents the contrast loss function;

[0116] This formula represents optimizing the parameters of the text encoder and audio encoder by minimizing the contrastive loss function.

[0117] The Clap module achieves effective mapping between text and audio through a contrastive learning framework, providing powerful semantic alignment capabilities for audio generation tasks. The text representation provided by the Clap module is used as conditional information to guide the Diffusion model to generate audio that matches the text description.

[0118] About the Diffusion module: The Diffusion module converts data into high-dimensional noise by gradually adding noise, and then recovers the original data from the noise through an iterative denoising process. This process can be divided into two main stages: the forward diffusion process and the reverse denoising process.

[0119] In the forward diffusion process, the model gradually transforms structured data into disordered noise. The process can be expressed as:

[0120]

[0121] In the formula, x0 is the original data, x 1:Tis from x0 to the noise data x T A series of intermediate states, q(x t |x t-1 ) is the probability distribution of adding noise at each step

[0122] The reverse denoising process is the inverse of the forward diffusion process, and its goal is to recover the original data from the noisy data. The conditions in this step are generated by Clap, guiding the model to denoise. The process can be expressed as:

[0123]

[0124] Where p θ (x t-1 |x t ) is given the current noise state x t Under the condition of t-1 The probability distribution of , θ represents the model parameters.

[0125] The training goal of the Diffusion module is to maximize the Evidence Lower Bound (ELBO). The process can be expressed as:

[0126]

[0127] By maximizing the ELBO, the Diffusion module learns how to generate data that is similar to the original data.

[0128] The goal of the forward diffusion process is to gradually add noise to the data until the data is completely transformed into noise. This process can be viewed as a Markov chain, where each step follows a Gaussian distribution, and the final state approaches a standard normal distribution. The main function of the forward diffusion process is to simulate the diffusion of data from an ordered state to a disordered state, providing a foundation for the subsequent reverse diffusion process.

[0129] The goal of the reverse diffusion process is to recover the original data from a noisy state. This process is also a parameterized Markov chain that gradually removes noise, generates a series of intermediate states, and ultimately recovers the original data. The main function of the reverse diffusion process is to generate new data samples that have a similar distribution to the training data.

[0130] The forward and reverse diffusion processes together constitute the complete lifecycle of a diffusion model. The forward diffusion process gradually transforms the data into a Gaussian noise distribution by gradually adding noise, while the reverse diffusion process gradually removes the noise to recover the original data. These two processes enable the diffusion model to generate complex, high-quality data samples from simple noise distributions.

[0131] The Diffusion module recovers audio that matches the text description from the noisy signal through a step-by-step denoising iterative process.

[0132] In this embodiment, if Figure 2 and Figure 3 As shown, the spectrum map G and the spectrum map X are introduced into the mask and fused to obtain the enhanced spectrum map X aug ; Specifically include:

[0133] Mask the spectrogram X and the spectrogram G using a mask randomly selected from the predefined n masks, and then concatenate the spectrogram X and the spectrogram G; the formula is:

[0134] X aug =α·(X☉M i )+(1-α)·(G☉(1-M i ))

[0135] Where α is the mixing coefficient, which is used to control the weight of the spectrogram X and the spectrogram G after fusion, and is usually 0.5; ⊙ represents the element-by-element multiplication, which is used to apply the mask to the spectrogram; i is a randomly selected mask index, i∈[1,n]), and n can be 4.

[0136] Specifically, in this way, the ship classification model can be exposed to more diverse and high-quality training samples during the training process. aug Feeding this data into a ship classification model for training improves its ability to recognize ship audio, thereby enhancing the model's generalization and robustness. This data augmentation technique not only enriches data diversity but also ensures label consistency, enabling the model to better maintain performance across different scenarios.

[0137] In this embodiment, if Figure 2 and Figure 4 As shown in the figure, features are extracted and fused from the time dimension and frequency dimension respectively, including:

[0138] The spectrogram input to the ship classification model is subjected to Y-axis average pooling and convolution operations to extract significant features in the time dimension. (The pooling operation reduces the dimensionality of the audio signal on the frequency axis, retaining key features on the time axis, while the convolution operation further captures the temporal trends and patterns of the audio signal.)

[0139] Extract significant features in the frequency dimension through X-axis average pooling and convolution operations (pooling reduces the dimensionality of the audio signal on the time axis to highlight key features on the frequency axis, while convolution further captures the energy distribution and spectral pattern of the audio signal in different frequency bands);

[0140] The features extracted from the time dimension and frequency dimension are sequentially subjected to depth-wise separable convolution, group normalization, and Sigmoid activation function, and finally element-wise multiplication is performed on them to obtain the final feature representation F.

[0141] The input to the ship classification model during training is the spectrum graph X aug For example, the formula is:

[0142] F time =Conv y (Pool y (X aug ))

[0143] F freq =Conv x (Pool x (X aug ))

[0144] F=Sigmoid(GroupNorm(DepthwiseConv(F time )))⊙Sigmoid(GroupNorm(DepthwiseConv(F freq )))

[0145] Where, Pool y () indicates the average pooling operation in the Y direction; Conv y () represents the convolution operation in the Y direction; Pool x () represents the average pooling operation in the X direction; Conv x () represents the convolution operation in the X direction; F time Represents the features extracted in the time dimension; F freq Represents the features extracted in the frequency dimension; DepthwiseConv() represents the depthwise separable convolution operation; GroupNorm() represents the group normalization operation; Sigmoid() represents the Sigmoid activation operation.

[0146] This feature extraction module can perform comprehensive feature extraction on audio signals from both time and frequency dimensions to more comprehensively capture the temporal changes and spectral distribution of audio signals, thereby more effectively extracting audio features. Feature extraction in the time dimension helps capture the dynamic temporal changes of ship audio, such as the start, end, and duration of the audio, which is crucial for analyzing the time series characteristics of ship sounds. Feature extraction in the frequency dimension can highlight the energy distribution of ship audio in different frequency bands, helping to distinguish the frequency characteristics of different types of ship sounds. The final fused feature representation provides a more comprehensive and in-depth feature representation for subsequent processing of the model, thereby improving the model's ability to distinguish different types of ship sounds.

[0147] In this embodiment, Figure 5 This is a framework diagram of the agent attention mechanism. The agent attention mechanism (AgentAttention) improves the computational efficiency and expressiveness of the model by introducing agent tokens. Its core idea is to use a small number of agent tokens to aggregate and broadcast global information, thereby achieving efficient feature reuse. The specific implementation steps are as follows:

[0148] First, we receive the input feature maps Q, K, and V, where Q is the query token, K is the key token, and V is the value token. These input feature maps provide the feature information that the model needs to process. Next, we pool or otherwise process Q to generate a proxy token A. The proxy token A acts as a proxy for the query and key, aggregating and broadcasting global information. Then, we use the proxy token A as the query to aggregate information from K and V to obtain the proxy feature V. A This step is achieved by calculating the similarity between A and K and performing a weighted sum on V. Next, using the proxy token A as the key, the proxy feature V A Broadcast to each query token Q to form the final output. This step is done by calculating the similarity between Q and A and A Finally, the attention weights are calculated by the Softmax function and weighted summation is performed to obtain the final output features.

[0149] Specifically, the proxy attention mechanism introduces a set of additional proxy tokens (agenttokens) that act as proxies for query tokens (Q), aggregate information from keys (K) and values ​​(V), and then broadcast this information back to Q. Since the number of proxy tokens can be designed to be much smaller than the number of query tokens, efficiency is significantly improved while retaining the ability to model global context. This design not only improves the model's discriminative ability, but also improves its generalization performance in complex audio environments, enabling the model to demonstrate higher accuracy and robustness in practical applications of ship sounds. The design of the attention mechanism improves the accuracy of identifying different types of ship sounds by giving the model the ability to identify key features in audio signals.

[0150] The solutions involved in the above embodiments are evaluated in conjunction with specific experiments.

[0151] The original CNN classification model was compared with the model in the above embodiment (ResAtt, which uses enhanced spectrum atlas training, improved feature extraction method, residual dropout path and proxy attention mechanism). The changes in loss and accuracy during the training process are shown in the following figure. Figure 6 Finally, the performance of the two models was tested and compared using the original Deepship dataset. The comparison results are shown in Table 1:

[0152] Table 1

[0153]

[0154] Among them, both models have four rows of data, each row is the result obtained by testing using a different test set (no repetition), which is used to reflect the accuracy of the model.

[0155] The experimental results show that the method in the above embodiment significantly improves model performance across all categories, particularly in the Cargo and Tanker categories, where accuracy increases by approximately 16.20% and 16.18%, respectively, for an overall improvement of approximately 11.64%. This demonstrates that the method in the above embodiment can improve model accuracy and generalization capabilities. These experimental results provide strong support for the method in the above embodiment, demonstrating its potential and adaptability in practical applications.

[0156] A ship classification device based on deep learning, comprising:

[0157] The dataset processing module is used to process the ship audio dataset to obtain an enhanced spectrogram set. The specific process is as follows: pre-process the ship audio in the ship audio dataset to obtain a spectrogram X; use the pre-trained diffusion model to generate a spectrogram G that matches the text prompt, and introduce the spectrogram G and the corresponding spectrogram X into the mask and fuse them to obtain the enhanced spectrogram X. aug ; Multiple spectrograms X aug Constructed into an enhanced spectrogram atlas; where the text prompts come from the labels of the corresponding ship audio in the ship audio dataset;

[0158] The classification model construction module is used to build a ship classification model. The ship classification model extracts and fuses features from the time dimension and frequency dimension respectively, and also introduces a proxy attention mechanism to enhance key features.

[0159] A training module is used to train the ship classification model using the enhanced spectrum atlas to obtain a trained ship classification model;

[0160] The classification module is used to pre-process the audio of the ship to be classified to obtain the spectrum map X to be classified, and input the spectrum map X to be classified into the trained ship classification model to perform ship classification.

[0161] A device comprising:

[0162] memory for storing computer programs;

[0163] A processor is used to execute the computer program, and when the computer program is executed by the processor, the steps of the ship classification method based on deep learning as described in the above embodiment are implemented.

[0164] A readable storage medium stores a computer program thereon, which, when executed by a processor, implements the steps of the ship classification method based on deep learning as described in the above embodiment.

[0165] A computer program product includes a computer program, which, when executed by a processor, implements the steps of the ship classification method based on deep learning as described in the above embodiment.

[0166] With the above-described preferred embodiments of the present invention as a guide, and with reference to the above description, relevant personnel are fully capable of making various changes and modifications without departing from the technical scope of this invention. The technical scope of this invention is not limited to the contents of the specification and must be determined according to the scope of the claims.

Claims

1. A ship classification method based on deep learning, characterized in that: include: Step S1, processing the ship audio dataset to obtain an enhanced spectrum atlas; Step S11, preprocessing the ship audio in the ship audio dataset to obtain a spectrogram X; Step S12: Use the pre-trained diffusion model to generate a spectrogram G that matches the text prompt, and introduce the spectrogram G and the corresponding spectrogram X into the mask and fuse them to obtain the enhanced spectrogram X. aug ; Among them, the text prompt comes from the label; Step S13, multiple spectrum graphs X aug Constructed into an enhanced spectrum atlas; Step S2: constructing a ship classification model. The ship classification model extracts and fuses features from both the time and frequency dimensions, and also introduces a proxy attention mechanism to enhance key features. Step S3, training the ship classification model using the enhanced spectrum atlas to obtain a trained ship classification model; Step S4: pre-process the audio of the ship to be classified to obtain a spectrogram X to be classified, and input the spectrogram X to be classified into the trained ship classification model to perform ship classification.

2. The ship classification method based on deep learning according to claim 1, characterized in that: Pre-process the ship audio; specifically including: First, cut the ship audio into segments of preset duration and store them by category; The ship audio is then loaded and resampled to a uniform sampling rate; Then convert it to the frequency domain through short-time Fourier transform, and convert the amplitude value into decibel value; Finally, the decibel value is used to calculate the spectrum graph X.

3. The ship classification method based on deep learning according to claim 1, characterized in that: The diffusion model includes the VAE module, the Diffusion module, and the Clap module. During the training of the diffusion model: The VAE module encodes the spectrogram X to obtain the corresponding ship audio potential features; The Clap module maps the ship audio and related text prompts into the same semantic space to obtain the association between the ship audio and the text description; During the forward propagation of the latent features, the Diffusion module gradually adds Gaussian noise to the latent features. During the backward propagation of the latent features, the Diffusion module uses the text encoding processed by the Clap module as a condition to guide the gradual denoising and restore the latent features. The VAE module then converts the recovered latent features back to the spectrogram G.

4. The ship classification method based on deep learning according to claim 1, characterized in that: The spectrogram G and the spectrogram X are introduced into the mask and fused to obtain the enhanced spectrogram X. aug ; Specifically include: Mask the spectrogram X and the spectrogram G using a mask randomly selected from the predefined n masks, and then concatenate the spectrogram X and the spectrogram G; the formula is: X aug =α·(X⊙M i )+(1-α)·(G⊙(1-M i )) Where α is the mixing coefficient, which is used to control the weight of the spectrogram X and the spectrogram G after fusion; ⊙ is the element-by-element multiplication, which is used to apply the mask to the spectrogram; i is a randomly selected mask index, i∈[1,n]).

5. The ship classification method based on deep learning according to claim 1, characterized in that: Features are extracted and fused from the time dimension and frequency dimension respectively, including: For the spectrum graph input to the ship classification model, average pooling and convolution operations are performed in the Y direction to extract features in the time dimension; average pooling and convolution operations are performed in the X direction to extract features in the frequency dimension; The features extracted from the time dimension and frequency dimension are subjected to depth-wise separable convolution, linking, group normalization and Sigmoid activation function respectively, and finally element-wise multiplication operation is performed to obtain the final feature representation.

6. The ship classification method based on deep learning according to claim 1, characterized in that: The ship classification model also introduces multiple residual blocks connected in sequence; among them, The first residual block takes the features obtained by fusing the features extracted from the time dimension and the frequency dimension as input; In the process of training the ship classification model, the Droppath iterative training strategy is adopted.

7. A ship classification device based on deep learning, characterized in that: include: The dataset processing module is used to process the ship audio dataset to obtain an enhanced spectrogram set. The specific process is as follows: pre-process the ship audio in the ship audio dataset to obtain a spectrogram X; use the pre-trained diffusion model to generate a spectrogram G that matches the text prompt, and introduce the spectrogram G and the corresponding spectrogram X into the mask and fuse them to obtain the enhanced spectrogram X. aug ; Multiple spectrograms X aug Constructed into an enhanced spectrum atlas; where the text prompts come from the labels; The classification model construction module is used to build a ship classification model. The ship classification model extracts and fuses features from the time dimension and frequency dimension respectively, and also introduces a proxy attention mechanism to enhance key features. A training module is used to train the ship classification model using the enhanced spectrum atlas to obtain a trained ship classification model; The classification module is used to pre-process the audio of the ship to be classified to obtain the spectrum map X to be classified, and input the spectrum map X to be classified into the trained ship classification model to perform ship classification.

8. A device, characterized in that include: memory for storing computer programs; A processor is used to execute the computer program, and when the computer program is executed by the processor, the steps of the ship classification method based on deep learning according to any one of claims 1 to 6 are implemented.

9. A readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the ship classification method based on deep learning are implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the ship classification method based on deep learning are implemented.