Underwater sound target small sample identification method based on time-frequency spectrum fusion and embedded feature reconstruction
Through the time spectrum fusion and embedding feature reconstruction of the CAF-ViT-CSC-EFR model, the problem of high data dependence in water acoustic target recognition is solved, and efficient recognition under small sample conditions is achieved.
Patent Information
- Application Number
- CN202510425321.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-08
AI Technical Summary
The existing water acoustic target recognition methods have high data dependence and low generalization ability under small sample conditions, making it difficult to achieve accurate identification.
The CAF-ViT-CSC-EFR model based on time spectrum fusion and embedded feature reconstruction is adopted. The features are extracted through the CAF-ViT-CSC network and combined with the embedded feature reconstruction module of the early stop strategy to improve the generalization ability of the model in small sample tasks.
It improves the accuracy of water acoustic target recognition, enhances the generalization ability of the model in small sample scenarios, and improves the accuracy and stability of the recognition.
Smart Images

Figure CN120279255A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an underwater acoustic target small-sample recognition method based on time-frequency spectrum fusion and embedded feature reconstruction, belonging to the field of underwater acoustic target recognition. Background Technique
[0002] Underwater acoustic target recognition is one of the important tasks of underwater acoustic detection and also a difficult problem in the field of underwater acoustic signal processing. Especially for accurately recognizing underwater acoustic targets under small-sample conditions, there are still great challenges. Traditional machine learning methods, such as support vector machines and K-nearest neighbor algorithms, generally have problems of strong dependence on data and limited generalization performance in the underwater acoustic target small-sample recognition task.
[0003] In addition, the performance of algorithms based on the pre-training - adaptation paradigm highly depends on the generalization ability of the pre-trained model, which poses strict requirements on the feature extraction ability of the basic model. Currently commonly used pre-trained models, such as deep neural networks, convolutional neural networks, and recurrent neural networks, all depend on the scale of training samples. Therefore, how to improve the generalization ability of the pre-trained model for small-sample tasks under the existing sample scale is an urgent problem to be solved currently.
[0004] Therefore, aiming at the deficiencies of the existing underwater acoustic target small-sample recognition methods, it is urgent to propose a new method with the capabilities of feature fusion, classification optimization, and generalization enhancement to improve the recognition accuracy in small-sample scenarios. Summary of the Invention
[0005] The present invention aims to solve the problems of high dependence on the data sample scale and low generalization ability of the existing methods in the underwater acoustic target small-sample recognition task, and proposes an underwater acoustic target small-sample recognition method based on time-frequency spectrum fusion and embedded feature reconstruction.
[0006] An underwater acoustic target small-sample recognition method based on time-frequency spectrum fusion and embedded feature reconstruction includes the following steps:
[0007] Step 1: Obtain the audio files of underwater acoustic target samples in the Deepship database and the Shipsear database, preprocess the audio files of the Deepship and Shipsear databases to obtain spectrogram samples, construct a pre-training data set and a small-sample data set, divide the spectrogram samples to obtain the training set and test set of the pre-training data set, and obtain the support set and query set of the small-sample data set. The specific steps are as follows:
[0008] Step 1.1: Obtain the audio files of underwater acoustic target samples in the Deepship database and the Shipsear database, and perform downsampling processing on the audio files;
[0009] Steps 1-2: Intercept audio segments x(n) with a duration of T from each audio file in a non-overlapping manner, where the sampling points n = 1, 2,..., L;
[0010] Steps 1-3: Frame the audio segments x(n) in Step 1-2 in a non-overlapping manner to obtain the i-th frame signal x i (n);
[0011] Steps 1-4: Perform Fourier transform on x i (n) to obtain the spectrum of the i-th frame signal:
[0012]
[0013] X i (f) is the spectrum of the i-th frame signal obtained, f is the frequency, and e -j2πnf / L represents the complex rotation factor, which decomposes the time-domain signal into sine / cosine components of different frequencies;
[0014] Calculate the power spectral density according to X i (f):
[0015] P i (f) = |X i (f)| 2
[0016] Form a two-dimensional graph with the power spectral density of each frame to obtain the LOFAR spectrogram of each frame;
[0017] Steps 1-5: Map the power spectral density P i (f) in Step 1-4 to the Mel frequency scale and filter the power spectrum using a Mel filter bank:
[0018]
[0019] H m (f) represents the Mel filter, usually in the form of a triangular filter, m represents the m-th filter, and M m is the output of the Mel filter, representing the m-th point of the i-th frame signal in the Mel spectrogram;
[0020] Form a two-dimensional graph with each frame of M m to obtain the Mel spectrogram of each frame;
[0021] Steps 1-6: Perform wavelet packet transform on x i (n) to decompose it into multiple levels,
[0022] d iv (n) = WP(x i (n))
[0023] div (n) represents a wavelet packet node, v represents the node index of the decomposition level, and WP() represents the wavelet packet transform;
[0024] Calculate the energy E of each wavelet packet node d iv (n): iv (n):
[0025]
[0026] Obtain the wavelet packet spectrogram of each frame according to the energy distribution;
[0027] Step Seventeen: Construct a pre-training set: The pre-training set is composed of LOFAR spectrograms, Mel spectrograms, and wavelet packet spectrograms generated from the audio files in the Deepship database. The pre-training set is divided into a training set and a test set according to the ratio of 80%:20%;
[0028] Step Eighteen: Construct a small-sample data set: Among the LOFAR spectrograms, Mel spectrograms, and wavelet packet spectrograms generated from the audio files in the ShipsEar database, randomly select a audio files from each class of underwater acoustic target samples, and select b spectrogram samples for each audio file to form the support set of the small-sample data set; Select another c audio files from each class of underwater acoustic target samples, and select d spectrogram samples for each audio file to form the query set of the small-sample data set.
[0029] Step Two: Construct the CAF-ViT-CSC-EFR model: The CAF-ViT-CSC-EFR model consists of a CAF-ViT-CSC network and an embedded feature reconstruction module based on an early stopping strategy. The CAF-ViT-CSC network consists of a feature extractor module based on an attention mechanism and a cosine similarity classifier module;
[0030] The feature extractor module based on the attention mechanism of the CAF-ViT-CSC network includes: an encoding unit, a feature extraction unit, and a feature fusion unit;
[0031] Encoding unit: For the input spectrogram, form a sequence by slicing each frame in order, add class encoding and position encoding to obtain x0,
[0032] x0 = [x class ||x slice +x pos
[0033] is the sequence formed by slicing each frame in order, N is the total number of time frames, C is the dimension of the slice, represents the real number field; is the trainable class encoding; || represents the matrix concatenation operation; x posIs a trainable position encoding;
[0034] Feature extraction unit: Extracts self-attention features from x0,
[0035] y1 = x0 + MSA(LN(x0))
[0036] x1 = y1 + MLP(LN(y1))
[0037] MSA() represents the attention module, MLP() represents the feed-forward neural network, LN() represents layer normalization, y1 represents the feature matrix output by the attention module, and x1 represents the feature matrix output by the feed-forward neural network;
[0038] Feature fusion unit: Inputs the self-attention features of the LOFAR spectrogram, Mel spectrogram, and wavelet packet spectrogram of the same audio file. The feature fusion unit combines the self-attention features of the three spectrograms in pairs to form six parallel cross-attention branches, performs cross-attention calculations on the two self-attention features of each cross-attention branch to obtain a cross-attention matrix, and takes the x of each matrix class As the embedded feature output of the model, that is, outputs z j , j = 1, 2,..., 6;
[0039] The cosine similarity classifier module is used to calculate the cosine similarity between the embedded feature z j And the weight feature w of each category i :
[0040] s i = (z j ) T w i / ||z j ||||w i ||
[0041] Furthermore, obtain the similarity scores [s1, s2,..., s c , and then normalize the similarity scores through the Softmax function to obtain the class prediction probability.
[0042] Step 3: Use the pre-training dataset in Step 1 to pre-train the CAF-ViT-CSC network until the specified number of training times is reached to obtain the pre-trained CAF-ViT-CSC network.
[0043] Step 4: Input the support set and query set of a small sample task of a category in the small sample dataset in Step 1 into the pre-trained CAF-ViT-CSC network to obtain the embedded feature output set S Z And the embedded feature output set Q Z , in SZ ∪Q Z Perform unsupervised training on the embedding feature reconstruction module based on the early stopping strategy, and the feature-level reconstruction loss function is:
[0044]
[0045] is the reconstruction network, S Z ∪Q Z is the number of samples in the set, d cos represents the cosine distance, z is the embedding feature vector, is the reconstructed embedding feature vector;
[0046] Calculate the LID value of the penultimate layer of the embedding feature reconstruction module based on the early stopping strategy after each round of training:
[0047]
[0048] is the output of the penultimate layer of the embedding feature reconstruction module based on the early stopping strategy, Calculate and the Euclidean distance between the i-th nearest neighbor in, m is the number of nearest neighbors;
[0049] Stop training when the LID value reaches the minimum or the maximum number of iterations k, and obtain the trained embedding feature reconstruction module based on the early stopping strategy.
[0050] Step Five: Input the support set of the small sample dataset in Step One into the pre-trained CAF-ViT-CSC network to obtain the embedding feature set of the support set, input the embedding feature set of the support set into the trained embedding feature reconstruction module based on the early stopping strategy to obtain the reconstructed support set of the embedding feature, use the reconstructed support set of the embedding feature to train the cosine similarity classifier, and use the Adam algorithm for training to obtain the trained CAF-ViT-CSC-EFR model.
[0051] Step Six: Use the query set of the small sample dataset to test the trained CAF-ViT-CSC-EFR model, and test the underwater acoustic target classification and recognition ability of the CAF-ViT-CSC-EFR model for small sample tasks;
[0052] If the recognition ability is qualified, obtain the qualified CAF-ViT-CSC-EFR model;
[0053] If the recognition ability is unqualified, return to Step Three for retraining.
[0054] The beneficial effects of the present invention are:
[0055] The present invention constructs a CAF-ViT-CSC-EFR model, which includes a CAF-ViT-CSC network and an embedded feature reconstruction module based on an early stopping strategy. The CAF-ViT-CSC network includes a feature extractor module based on an attention mechanism and a cosine similarity classification module. The CAF-ViT-CSC network can accurately extract sample features, and the embedded feature reconstruction module based on the early stopping strategy can improve the generalization ability of the model for few-shot tasks, solve the problems of high dependence on the data sample scale and low generalization ability in the few-shot recognition task of underwater acoustic targets in existing methods, and improve the accuracy of underwater acoustic target recognition in few-shot scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 It is a training flow chart of the pre-training stage and the adaptation stage of the present invention;
[0057] Figure 2 It is a training flow chart of the embedded feature reconstruction module based on the early stopping strategy;
[0058] Figure 3 It is a schematic diagram of six cross-attention branches of the feature fusion module;
[0059] Figure 4 It is the relationship between the reconstruction feature recognition accuracy and the LID value under different support set sizes;
[0060] Figure 5 It is a result comparison chart of the feature optimization strategy. DETAILED DESCRIPTION OF THE INVENTION DETAILED DESCRIPTION OF THE INVENTION 1:
[0062] This embodiment is a few-shot recognition method for underwater acoustic targets based on time-frequency spectrum fusion and embedded feature reconstruction, including the following steps:
[0063] Step 1. Obtain the audio files of underwater acoustic target samples in the Deepship database and the Shipsear database, preprocess the audio files of the Deepship and Shipsear databases to obtain spectrogram samples, construct a pre-training data set and a few-shot data set, divide the spectrogram samples to obtain the training set and test set of the pre-training data set, and obtain the support set and query set of the few-shot data set. The specific steps are as follows:
[0064] Step 1-1. Obtain the audio files of underwater acoustic target samples in the Deepship database and the Shipsear database, and perform downsampling processing on the audio files to reduce the sampling rate to 16 kHz;
[0065] Step 1-2. For each audio file, intercept an audio segment x(n) with a duration of 3 s in a non-overlapping manner, where the sampling points n = 1, 2,..., L;
[0066] Step 13: Frame the audio segment x(n) in Step 12 in a non - overlapping manner to obtain the i - th frame signal x i (n);
[0067] Step 14: Perform Fourier transform on x i (n) to obtain the spectrum of the i - th frame signal:
[0068]
[0069] X i (f) is the spectrum of the i - th frame signal, f is the frequency, e -j2πnf / L represents the complex rotation factor, and decomposes the time - domain signal into sine / cosine components of different frequencies;
[0070] Calculate the power spectral density according to X i (f):
[0071] P i (f)=|X i (f)| 2
[0072] Form a two - dimensional graph with the power spectral density of each frame to obtain the LOFAR spectrogram of each frame;
[0073] Step 15: Map the power spectral density P i (f) in Step 14 to the Mel frequency scale, and filter the power spectrum using a Mel filter bank:
[0074]
[0075] H m (f) represents the Mel filter, usually in the form of a triangular filter, m represents the m - th filter, M m is the output of the Mel filter, representing the m - th point of the i - th frame signal in the Mel spectrogram;
[0076] Form a two - dimensional graph with each frame's M m to obtain the Mel spectrogram of each frame;
[0077] Step 16: Perform wavelet packet transform on x i (n) to decompose it into multiple levels,
[0078] d iv (n)=WP(x i (n))
[0079] d iv (n) represents the wavelet packet node, v represents the node index of the decomposition level, and WP() represents the wavelet packet transform;
[0080] Calculate the energy E of each wavelet packet node d iv (n): iv (n):
[0081]
[0082] Obtain the wavelet packet spectrogram of each frame according to the energy distribution;
[0083] Step 17. Construct a pre-training set: The pre-training set is composed of LOFAR spectrograms, Mel spectrograms, and wavelet packet spectrograms generated from the audio files in the Deepship database. The pre-training set is divided into a training set and a test set according to the ratio of 80%:20%;
[0084] Step 18. Construct a small sample data set: Among the LOFAR spectrograms, Mel spectrograms, and wavelet packet spectrograms generated from the audio files in the ShipsEar database, randomly select 3 audio files from each class of underwater acoustic target samples, and select 15 spectrogram samples for each audio file to form the support set of the small sample data set; Select another 1 audio file from each class of underwater acoustic target samples, and select 16 spectrogram samples for each audio file to form the query set of the small sample data set.
[0085] Step 2. Construct the CAF-ViT-CSC-EFR model: The CAF-ViT-CSC-EFR model consists of the CAF-ViT-CSC network and the embedded feature reconstruction module based on the early stopping strategy. The CAF-ViT-CSC network consists of the feature extractor module based on the attention mechanism and the cosine similarity classifier module;
[0086] The feature extractor module based on the attention mechanism of the CAF-ViT-CSC network includes: an encoding unit, a feature extraction unit, and a feature fusion unit;
[0087] Encoding unit: For the input spectrogram, form a sequence by slicing each frame in order, add class encoding and position encoding to obtain x0,
[0088] x0 = [x class ||x slice + x pos
[0089] is the sequence formed by slicing each frame in order, N is the total number of time frames, C is the dimension of the slice, represents the real number field; is the trainable class encoding; || represents the matrix concatenation operation; x pos is the trainable position encoding;
[0090] Feature extraction unit: Extract self-attention features from x0,
[0091] y1 = x0 + MSA(LN(x0))
[0092] x1 = y1 + MLP(LN(y1))
[0093] MSA() represents the attention module, MLP() represents the feed-forward neural network, LN() represents layer normalization, y1 represents the feature matrix output by the attention module, and x1 represents the feature matrix output by the feed-forward neural network;
[0094] Feature fusion unit: Input the self-attention features of the LOFAR spectrogram, Mel spectrogram, and wavelet packet spectrogram of the same audio file. The feature fusion unit combines the self-attention features of the three spectrograms in pairs to form six parallel cross-attention branches, as Figure 3 shown. Each cross-attention branch takes a pair of self-attention features as input, uses the x of the first self-attention feature as an agent, and performs cross-attention calculation with the x of the second self-attention feature to obtain the cross-attention matrix. The x of each matrix is used as the embedded feature output of the model, that is, the output z class as an agent, and performs cross-attention calculation with the x of the second self-attention feature to obtain the cross-attention matrix. The x of each matrix is used as the embedded feature output of the model, that is, the output z slice for cross-attention calculation to obtain the cross-attention matrix. The x of each matrix is used as the embedded feature output of the model, that is, the output z class , j = 1, 2,..., 6; j , j = 1, 2,..., 6;
[0095] The cosine similarity classifier module is used to calculate the cosine similarity between the embedded feature z j and the weight feature w i of each category:
[0096] s i = (z j ) T w i / ||z j || ||w i ||
[0097] Furthermore, the similarity scores [s1, s2,..., s c of all categories are obtained, and then the similarity scores are normalized by the Softmax function to obtain the class prediction probability.
[0098] Step 3: Use the pre-training dataset in Step 1 to pre-train the CAF-ViT-CSC network, as shown in the pre-training stage Figure 1 shown in the pre-training stage. After reaching the specified number of training times, the pre-trained CAF-ViT-CSC network is obtained.
[0099] Step 4: Input the support set and query set of the few-shot task of one category in the few-shot dataset in Step 1 into the pre-trained CAF-ViT-CSC network to obtain the output set S of the embedded features of the support set samples Zand the embedding feature output set Q of the query set sample Z , on S Z ∪Q Z , perform unsupervised training on the embedding feature reconstruction module based on the early stopping strategy, as Figure 2 shown, the feature-level reconstruction loss function is:
[0100]
[0101] is the reconstruction network, S Z ∪Q Z is the number of samples in the set, d cos represents the cosine distance, z is the embedding feature vector, is the reconstructed embedding feature vector;
[0102] Calculate the LID value of the penultimate layer of the embedding feature reconstruction module based on the early stopping strategy after each round of training:
[0103]
[0104] is the output of the penultimate layer of the embedding feature reconstruction module based on the early stopping strategy, Calculate and the Euclidean distance between the i-th nearest neighbors in, m is the number of nearest neighbors;
[0105] Stop training when the LID value reaches the minimum or reaches the maximum number of iterations k, and obtain the trained embedding feature reconstruction module based on the early stopping strategy.
[0106] Step 5. Input the support set of the small sample data set in Step 1 into the pre-trained CAF-ViT-CSC network to obtain the embedding feature set of the support set, input the embedding feature set of the support set into the trained embedding feature reconstruction module based on the early stopping strategy to obtain the reconstructed support set of the embedding feature, and use the reconstructed support set of the embedding feature to train the cosine similarity classifier, as Figure 1 shown in the adaptation stage, and use the Adam algorithm for training to obtain the trained CAF-ViT-CSC-EFR model.
[0107] Step 6. Use the query set of the small sample data set to test the trained CAF-ViT-CSC-EFR model, and test the underwater acoustic target classification and recognition ability of the CAF-ViT-CSC-EFR model for small sample tasks;
[0108] If the recognition ability is qualified, obtain the qualified CAF-ViT-CSC-EFR model;
[0109] If the recognition ability fails, return to step three for retraining.
[0110] Example 1:
[0111] First, obtain the audio files of underwater acoustic target samples in the Deepship database. The total recorded data length of DeepShip is nearly 48h, with a total of 265 audio files. The target categories are divided into four categories: cargo ship, passenger ship, oil tanker, and tugboat. Obtain the audio files of underwater acoustic target samples in the ShipsEar database. The ShipsEar database contains a total of 90 records, and four types of samples, namely motorboat, dredger, mussel boat, and small passenger ship, are selected.
[0112] Perform downsampling on the audio files to reduce the sampling rate to 16kHz. For each audio file, non-overlapping audio segments with a duration of 3s are intercepted. For each audio segment, using parameter settings with a frame length of 250ms and a frame shift of 64ms, LOFAR spectra, Mel spectra, and wavelet packet spectra are generated respectively. The dimensions of the spectrograms are 43×100, 43×40, and 43×32 respectively.
[0113] The spectrogram samples generated by DeepShip are divided into a training set and a test set according to a ratio of 80%:20%. To improve the training quality of the pre-trained model, the dataset is strictly divided to ensure that all samples generated from the same audio file do not appear in the training set and the test set at the same time. For the spectrogram samples generated by ShipsEar, 3 audio files are randomly selected from each category of samples, and 15 spectrogram samples are selected from each audio file to generate the support set; 16 spectrogram samples from 1 file in each category are selected to generate the query set, constructing a small sample dataset.
[0114] Set the hyperparameters of the model. Hyperparameters are non-training parameters of the model. The maximum number of training cycles for embedded feature reconstruction is set to 30, the encoder depth is set to 4, the number of multi-head attentions is set to 2, the expansion factor of the MLP_Head layer is set to 4, and the Dropout parameter is set to 0.1.
[0115] Construct the CAF-ViT-CSC-EFR model: The model uses the Pytorch deep learning framework to build a deep learning network. The software environment is Python3.8, Cuda10.0, and the hardware configuration is GPU: NvidiaRtx3060, CPU: Intel i512500h. The model training is divided into a pre-training stage and an adaptation stage, and the Adam algorithm is selected for network optimization.
[0116] In the pre-training stage, set the initial learning rate to 0.005, which decays as the training cycle increases. Set the batch size of each training batch to 256, and the maximum number of training cycles to 150.
[0117] In the adaptation stage, set the initial learning rate to 0.0005, use all samples as a batch for training, the maximum number of training epochs is 30, introduce Dropout regularization in each layer of the network, and set the dropout rate to 0.2.
[0118] Use the trained embedding feature reconstruction module based on the early stopping strategy to reconstruct the small sample dataset, and use the reconstructed small sample dataset to train the cosine similarity classifier module. The size of each training batch (batchsize) is 4, and the maximum number of training epochs is 180. Set the learning rate to vary between 0.001 and 0.0001 through the cosine annealing algorithm, and terminate the training in a timely manner according to the early stopping strategy.
[0119] Based on the trained CAF-ViT-CSC-EFR model, test the test set data in the small sample dataset, and statistically calculate the recognition accuracy rate and 95% confidence interval of the prediction results under the test set.
[0120] The results are shown in the following table. The values in parentheses are the average accuracy rate and its 95% confidence interval. The pre-trained model of the baseline algorithm is CAF-ViT, and a fully connected classifier is used in the adaptation stage.
[0121] Table 1 Comparison of accuracy rates
[0122]
[0123]
[0124] It can be seen that in various classification tasks, compared with the baseline algorithm, applying the cosine similarity classifier or the feature reconstruction module alone can improve the recognition accuracy rate. Among them, the improvement effect of introducing the cosine similarity classifier is more significant. When the cosine similarity classifier and the feature reconstruction module are combined (i.e., the algorithm in this paper), the recognition accuracy rate reaches the highest in various tasks. Specifically, the recognition accuracy rates of 2-class, 3-class, and 4-class classification are increased by 5.6%, 9.6%, and 6.6% respectively compared with the baseline algorithm. This shows that the synergistic effect of the feature reconstruction module and the cosine similarity classifier can effectively enhance the recognition ability of the model under small sample conditions. In addition, with the optimization of the model configuration, not only the recognition accuracy rate is improved, but also the 95% confidence interval is further shrunk, indicating that the improved model has higher stability while improving the recognition accuracy.
[0125] Example 2:
[0126] In the present invention, the embedded feature reconstruction module is a key component for realizing small-sample recognition of underwater acoustic targets. In order to study the dynamic variation law of the LID value and the recognition accuracy during the training process and reveal the influence of reconstruction training on the generalization of embedded features, the following experiments were designed, and the specific steps are as follows:
[0127] Step 1: In the experiment, the feature reconstruction module was applied to the embedded features output by the six cross-attention branches, and a small-sample recognition experiment was carried out based on the reconstructed features.
[0128] Step 2: Since multiple reconstruction modules presented similar experimental results under the same small-sample recognition task, the feature reconstruction training process of one of the cross-attention branches was selected. Figure 4 For the relationship between the LID value and the recognition accuracy under the support set sizes (15-Shot, 60-Shot, and 120-Shot) of different small-sample data sets, the vertical solid line in the figure represents the training cycle when the peak of the recognition accuracy appears, and the dotted line marks the training cycle when the trough of the LID value appears.
[0129] It can be seen that under the 15-Shot, 60-Shot, and 120-Shot tasks, the LID value shows a trend of first decreasing and then increasing, and the appearance time of the peak of the recognition accuracy lags slightly behind the appearance time of the LID trough. Specifically, in the 3-Way 15-Shot task, the peak of the recognition accuracy lags behind the LID trough by 1 training cycle, while in the 3-Way 60-Shot and 3-Way 120-Shot tasks, this lag time expands to 2 training cycles. This phenomenon can be explained as follows: as the sample size of the support set increases, the model first extracts shared features through structured learning during the reconstruction stage (corresponding to the LID decrease stage), and then enters a short feature adaptation period to fuse the details of small samples (corresponding to the initial stage of the LID recovery). At this time, the feature space enhances the discriminative detail representation ability while maintaining generalization, thus achieving the best recognition performance. It is proved that the change trend of the LID value is strongly correlated with the generalization of embedded features and can be used to improve the recognition accuracy of the model.
[0130] Example 3:
[0131] In the present invention, the embedded feature reconstruction module is a key component for realizing small-sample recognition of underwater acoustic targets. In order to further analyze the performance advantages of the embedded feature reconstruction module, the following experiments were designed, and the specific steps are as follows:
[0132] Step 1: The comparison algorithm adopts the backbone fine-tuning method, which freezes the front layer of the network and only performs supervised fine-tuning on the back layer of the network to optimize the feature expression ability of the backbone network.
[0133] Step 2: The comparison experiment results are shown in Figure 5. The thicker the line in the figure, the more parameters are involved in fine-tuning.
[0134] The results show that under the condition of small samples, fine-tuning a large number of parameters of the backbone network will instead have a negative impact on the model performance, and there is a negative correlation between the number of fine-tuned parameters and the recognition accuracy. This phenomenon indicates that excessive parameter adjustment will exacerbate the risk of overfitting, resulting in the deterioration of the generalization ability of the model in small-sample tasks. In the 3-Way 15-Shot task, the recognition accuracy of the embedded feature reconstruction module has increased by 3.0% compared with that before the embedded feature reconstruction, which proves that the embedded feature reconstruction module shows significant performance advantages under small-sample conditions.
Claims
1. An underwater acoustic target small-sample recognition method based on time-frequency spectrum fusion and embedded feature reconstruction, characterized in that, It includes the following steps: Step 1: Obtain the audio files of underwater acoustic target samples in the Deepship database and the Shipsear database, preprocess the audio files of the Deepship and Shipsear databases to obtain spectrogram samples, construct a pre-training dataset and a few-shot dataset, divide the spectrogram samples to obtain the training set and test set of the pre-training dataset, and obtain the support set and query set of the few-shot dataset; Step 2: Construct the CAF-ViT-CSC-EFR model: The CAF-ViT-CSC-EFR model consists of the CAF-ViT-CSC network and an embedded feature reconstruction module based on the early stopping strategy. The CAF-ViT-CSC network consists of a feature extractor module based on the attention mechanism and a cosine similarity classifier module; Step 3: Use the pre-training dataset in Step 1 to pre-train the CAF-ViT-CSC network until the specified number of training times is reached to obtain a pre-trained CAF-ViT-CSC network; Step 4: Use the few-shot dataset in Step 1 to train the embedded feature reconstruction module based on the early stopping strategy. When the LID value of the penultimate layer of the embedded feature reconstruction module based on the early stopping strategy is the smallest or reaches the maximum iteration number k, obtain a trained embedded feature reconstruction module based on the early stopping strategy; Step 5: Use the trained embedded feature reconstruction module based on the early stopping strategy to reconstruct the few-shot dataset, and use the reconstructed few-shot dataset to train the cosine similarity classifier module to obtain a trained CAF-ViT-CSC-EFR model; Step 6: Use the query set of the few-shot dataset to test the trained CAF-ViT-CSC-EFR model and test the underwater acoustic target classification and recognition ability of the CAF-ViT-CSC-EFR model for few-shot tasks; If the recognition ability is qualified, obtain a qualified CAF-ViT-CSC-EFR model; If the recognition ability is unqualified, return to Step 3 for re-training.
2. The small-sample recognition method for underwater acoustic targets based on time-frequency spectrum fusion and embedded feature reconstruction according to claim 1, characterized in that Step 1 includes the following steps: Step 1.1: Obtain the audio files of underwater acoustic target samples in the Deepship database and the Shipsear database, and perform downsampling processing on the audio files; Step 1.2: For each audio file, intercept an audio segment x(n) with a duration of T in a non-overlapping manner, where the sampling points n = 1, 2,..., L; Step 13: Frame the audio segment x(n) in Step 12 in a non-overlapping manner to obtain the i-th frame signal x i (n); Step 14. Take x i (n) and perform Fourier transform to obtain the spectrum of the i-th frame signal: X i (f) To obtain the spectrum of the i-th frame signal, f is the frequency, e -j2πnf / L represents the complex rotation factor, and decomposes the time-domain signal into sine / cosine components of different frequencies; According to X i (f) Calculate the power spectral density: P i (f) = |X i (f)| 2 Form a two-dimensional graph with the power spectral density of each frame to obtain the LOFAR spectrogram of each frame; Step 15. Map the power spectral density P i (f) obtained in Step 14 to the Mel frequency scale, and filter the power spectrum using a Mel filter bank: H m (f) represents a Mel filter, usually in the form of a triangular filter, m represents the m-th filter, M m is the output of the Mel filter, representing the m-th point of the i-th frame signal in the Mel spectrogram; Combine the M of each frame m to form a two-dimensional graph and obtain the Mel spectrogram of each frame; Step 1-6. Perform wavelet packet transform decomposition on x i (n) into multiple levels, d iv (n) = WP(x i (n)) d iv (n) represents a wavelet packet node, v represents the node index of the decomposition level, and WP() represents the wavelet packet transform; Calculate the energy E of each wavelet packet node d(n): iv (n) iv (n): Obtain the wavelet packet spectrogram of each frame according to the energy distribution; Step 1.7: Construct a pre-training set: Combine the LOFAR spectrograms, Mel spectrograms, and wavelet packet spectrograms generated from the audio files of the Deepship database to form a pre-training set, and divide the pre-training set into a training set and a test set according to a ratio of 80%:20%. Step 1-8: Construct a small-sample dataset: Among the LOFAR spectrograms, Mel spectrograms, and wavelet packet spectrograms generated from the audio files in the ShipsEar database, randomly select a audio files from each type of underwater acoustic target sample, and select b spectrogram samples from each audio file to form the support set of the small-sample dataset; select another c audio files from each type of underwater acoustic target sample, and select d spectrogram samples from each audio file to form the query set of the small-sample dataset.
3. A small sample recognition method for underwater acoustic targets based on time-frequency spectrum fusion and embedded feature reconstruction according to claim 1, characterized in that, The attention mechanism-based feature extractor module of the CAF-ViT-CSC network in Step 2 includes: an encoding unit, a feature extraction unit, and a feature fusion unit; Encoding unit: For the input spectrogram, form a sequence by slicing each frame in order, and add class encoding and position encoding to obtain x0. x0 = [x class ||x slice +x pos is a sequence formed by frame slices in order, N is the total number of time frames, and C is the dimension of the slices. represents the real number field; is a trainable class encoding; || represents matrix concatenation operation; x pos is a trainable position encoding; Feature extraction unit: Extract self-attention features from x0. y1 = x0 + MSA(LN(x0)) x1 = y1 + MLP(LN(y1)) MSA( ) represents the attention module, MLP() represents the feed-forward neural network, LN() represents layer normalization, y1 represents the feature matrix output by the attention module, and x1 represents the feature matrix output by the feed-forward neural network; Feature fusion unit: Input the self-attention features of the LOFAR spectrogram, Mel spectrogram, and wavelet packet spectrogram of the same audio file. The feature fusion unit combines the self-attention features of the three spectrograms in pairs to form six parallel cross-attention branches. Each cross-attention branch takes a pair of self-attention features as input, uses the x of the first self-attention feature class as an agent, and performs cross-attention calculation with the x of the second self-attention feature slice to obtain a cross-attention matrix, and uses the x of each matrix class as the embedded feature output of the model, that is, output z j , j = 1, 2,..., 6.
4. A small sample recognition method for underwater acoustic targets based on time-frequency spectrum fusion and embedded feature reconstruction according to claim 1, characterized in that, The cosine similarity classifier module in Step 2 is used to calculate the cosine similarity between the embedded feature z j and the weighted features w i of each category: s i =(z j ) T w i / ||z j ||||w i || Furthermore, similarity scores [s1, s2,..., s c are obtained, and then the similarity scores are normalized by the Softmax function to obtain the class prediction probabilities.
5. A small sample recognition method for underwater acoustic targets based on time-frequency spectrum fusion and embedded feature reconstruction according to claim 1, characterized in that The specific process of Step 4 is as follows: Input the support set and query set of a small-sample task of one category in the small-sample dataset in Step 1 into the pre-trained CAF-ViT-CSC network to obtain the output set S of the embedded features of the support set samples Z and the output set Q of the embedded features of the query set samples Z , and perform unsupervised training on the embedded feature reconstruction module based on the early stopping strategy on S Z ∪Q Z . The feature-level reconstruction loss function is as follows: For the reconstructed network, S Z ∪Q Z is the number of samples in the set, d cos represents the cosine distance, z is the embedded feature vector, (z) is the reconstructed embedded feature vector; Calculate the LID value of the penultimate layer of the embedding feature reconstruction module based on the early stopping strategy after each round of training; is the output of the penultimate layer of the embedding feature reconstruction module based on the early stopping strategy, Calculate and the Euclidean distance between the i-th nearest neighbor in, where m is the number of nearest neighbors; Stop training when the LID value reaches the minimum to obtain a trained embedding feature reconstruction module based on the early stopping strategy.
6. The small-sample recognition method for underwater acoustic targets based on time-frequency spectrum fusion and embedded feature reconstruction according to claim 1, wherein The specific process of Step 5 is as follows: Input the support set of the small-sample dataset in Step 1 into the pre-trained CAF-ViT-CSC network to obtain the embedding feature set of the support set, input the embedding feature set of the support set into the trained embedding feature reconstruction module based on the early stopping strategy to obtain the support set after embedding feature reconstruction, use the support set after embedding feature reconstruction to train the cosine similarity classifier, and use the Adam algorithm for training to obtain the trained CAF-ViT-CSC-EFR model.
Citation Information
Cited By
Internal arteriovenous fistula blood flow sound signal processing method and electronic equipment
CN120748457A