A Few-Shot Modulation Recognition Method Based on Diffusion Model and Attention Mechanism

High-quality samples are generated through diffusion model DDIM and attention mechanism, combined with classifier guidance and multi-scale Fourier decomposition, the problem of insufficient data in wireless signal small sample modulation recognition is solved, and the accuracy of modulation recognition and the interpretability of the model is improved.

CN119996133BActive Publication Date: 2025-07-01PLA PEOPLES LIBERATION ARMY OF CHINA STRATEGIC SUPPORT FORCE AEROSPACE ENG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510221985.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-07-01
Estimated Expiration
2045-02-26

AI Technical Summary

Technical Problem

In wireless signal small sample modulation recognition, due to insufficient data sets, it is difficult for the prior art to effectively improve the modulation recognition rate. Especially in complex channel environments and multiple modulation methods, the deep learning model generalization capabilities are insufficient, resulting in poor recognition effect.

Method used

A small sample modulation recognition method based on diffusion model and attention mechanism is adopted to generate high-quality samples through diffusion model DDIM, and a classifier is used to guide the generation process. Combining multi-scale Fourier decomposition and attention mechanism to extract frequency and nonlinear features, the classifier model FATT is constructed for modulation type recognition.

Benefits of technology

The modulation recognition accuracy under small sample conditions is improved, the problems of underfitting and overfitting are alleviated, and the phase interference of traditional data augmentation methods and pattern crashes in the GAN generation process are avoided, providing a solution with high accuracy and good interpretability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119996133B_ABST
    Figure CN119996133B_ABST
Patent Text Reader

Abstract

The present invention discloses a few-shot modulation recognition method based on a diffusion model and an attention mechanism, belonging to the field of modulation carrier systems, which solves the problem of low modulation recognition rate under few-shot conditions, and includes: inputting the original data into DDIM for progressive denoising training and adopting an inverse denoising process for progressive denoising to generate restored data; adjusting the generation trajectory based on the logarithmic probability gradient of the classifier to obtain synthetic data and a trained DDIM; inputting the original data, the restored data and the synthetic data into FATT, mapping them to the frequency domain through Fourier basis functions, explicitly capturing periodic and aperiodic features, and extracting comprehensive features; through the attention layer, splitting the comprehensive features into frequency features and non-linear features, splicing the weighted frequency features and the non-linear features and splicing them again with the output of the residual convolution block to obtain a predicted modulation type vector; visualizing and displaying the predicted modulation type results through a confusion matrix to obtain a trained FATT; the present invention improves the modulation recognition accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of modulation carrier systems, and more particularly to the technical field of small-sample modulation recognition of wireless signals, and relates to a small-sample modulation recognition method based on a diffusion model and an attention mechanism. Background Art

[0002] In the field of modulation carrier systems, especially in the research of small-sample modulation recognition of wireless signals, in Internet access and related services, due to insufficient data sets, the neural network in deep learning suffers from underfitting, resulting in poor recognition effects, thus leading to the problem of low recognition rate of modulation recognition under small-sample conditions. How to solve the problem of insufficient data sets is a major difficulty in the field of small-sample modulation recognition of wireless signals.

[0003] In the prior art, some scholars believe that using classifiers in modulation carrier systems, such as machine learning methods like SVM, Bayesian inference, and sparse recovery, can, to a certain extent, solve the problem of low recognition rate of modulation recognition under small-sample conditions in Internet access and related services. However, machine learning methods rely heavily on a large amount of data to find the optimal solution. Such an approach is highly interpretable. However, when the data is insufficient, it is often impossible to find the optimal solution and the effect is poor when using machine learning for optimization.

[0004] Although existing AMC technologies perform well in certain specific scenarios, in practical applications, modulation carrier systems often face challenges in complex channel environments and various modulation methods. Especially under small-sample conditions, these problems are particularly prominent. On the one hand, in Internet access and related services, channel defects such as incomplete channel state information, carrier frequency offset, symbol timing offset, and phase offset significantly reduce the classification performance of modulation signals. On the other hand, under small-sample conditions, existing deep learning models, such as convolutional neural network CNN, long short-term memory network LSTM, etc., are difficult to obtain sufficient labeled training data to capture the feature differences of different modulation formats, resulting in insufficient model generalization ability and low classification accuracy. As Figure 1 shown is the basic method for automatic modulation recognition under small-sample conditions.

[0005] As Figure 2As shown, it is a general flowchart of small-sample modulation recognition in the field of modulated carrier systems. It is a flowchart of small-sample modulation recognition using data augmentation methods such as rotation and flipping. Liang et al. published L. Huang, W. Pan, Y. Zhang, L. Qian, N. Gao and Y. Wu, "Data Augmentation for Deep Learning-Based Radio Modulation Classification," in IEEE Access, vol. 8, pp. 1498-1506, 2020, doi: 10.1109 / ACCESS.2019.2960775. Applying the data amplification technology in the image field to wireless signal modulation recognition, the open radio signal dataset RadioML2016.10a is used, which contains signal samples of 11 different modulation categories. According to the characteristics of the modulation signals, three augmentation methods, namely rotation, flipping, and Gaussian noise, are used for preprocessing to enhance the dataset. Experiments are conducted using the LSTM network architecture, and the network structure includes two LSTM layers and a fully connected layer for classification. The experimental results show that all three augmentation methods can improve the classification accuracy. Data augmentation also allows for the successful classification of radio signals with fewer sampling points, thus simplifying the deep learning model and shortening the classification response time.

[0006] Subsequently, when the GAN network achieved great success in the field of image generation, many scholars converted the modulation signals into constellation diagrams and directly generated constellation diagrams for modulation recognition just like image generation. However, this is not fundamental. Figure 3This is a flowchart for few-shot modulation recognition using GAN. Tang et al. published Z. Tang, M. Tao, J. Su, Y. Gong, Y. Fan and T. Li, "Data Augmentation for Signal Modulation Classification using Generative Adverse Network," 2021 IEEE 4th International Conference on Electronic Information and Communication Technology (ICEICT), Xi'an, China, 2021, pp. 450 - 453, doi: 10.1109 / ICEICT53123.2021.9531296. A data augmentation method for signal modulation classification based on generative adversarial network is used. Instead of converting the signal, it processes the signal directly and takes into account various signal-to-noise ratio situations. The open radio signal dataset RadioML2016.10a is used, which contains signal samples of 11 different modulation classes. After augmenting the original few-shot data using the generative adversarial network, a CNN network is used for classification and recognition.

[0007] However, rotation can cause some modulation types to be very sensitive to phase changes. The rotation operation may change the phase information of the signal, thus affecting the classification result. Moreover, the rotation operation generates multiple similar samples, which may lead to an increase in redundant information in the dataset, thereby increasing the computational overhead and training time. Flipping, such as horizontal flipping and vertical flipping, may change the physical meaning of the signal, causing the model to learn incorrect features.

[0008] The training process of the GAN network is unstable and prone to problems such as mode collapse and vanishing gradients, resulting in the generated data lacking diversity. It requires a large amount of computational resources and time. Especially when the network structures of the generator and discriminator are complex, the generation process is complex and it is difficult to explain the internal mechanism of the generated data, which is a disadvantage for some application scenarios that require high interpretability.

[0009] When dealing with time series on a classifier such as a CNN network, it is very weak in capturing long-range dependencies and usually requires a fixed-size input, being not flexible enough for variable-length sequences. When processing with an LSTM network, due to the dependence of sequence data, it is difficult for the LSTM to be parallelized, and the training and inference speeds are slow. Moreover, although the LSTM alleviates the problem of vanishing gradients to a certain extent, it may still encounter this problem when dealing with very long sequences. Summary of the Invention

[0010] To solve the technical problem of low modulation recognition rate under small sample conditions due to insufficient data sets in Internet access and related services, the present invention provides a small sample modulation recognition method based on a diffusion model and an attention mechanism. This method uses classifier-guided DDIM to generate high-quality samples under given conditions, which helps to solve the small sample problem. By generating more diverse samples to enhance the training data, the generation process is relatively stable and not prone to mode collapse. The step-by-step denoising generation process disclosed in the present invention can better control the quality of the generated samples, making the generated samples closer to real data. This step-by-step generation process provides a clear framework that can intuitively understand and analyze the changes generated at each step. This transparency helps to explain the generation mechanism of the model and the generation process of the data.

[0011] The attention mechanism was initially proposed in natural language processing to solve the problem of locating key features in long sequence information. Its core idea is to assign different "importance levels" to different positions in the input features, i.e., different frequency bands, etc., through learnable weights, so as to achieve focusing on key regions or channels and suppressing irrelevant or redundant information. In the present invention, the multi-scale Fourier decomposition module has mapped the signal to the frequency domain based on the Fourier basis function, explicitly capturing its periodicity and phase information. However, under different channel conditions, modulation formats, and noise environments, not all frequency components are equally important for classification. Through the attention layer, the more significant or discriminative frequency bands are weighted, and then the visual "attention weights" are output based on the attention mechanism, indicating which frequency bands contribute more to the final decision, improving the interpretability of the model. Finally, under Doppler conditions and multipath channel interference in complex environments, by highlighting key information points, the classification accuracy of high-order modulation signals is improved.

[0012] The object of the present invention is specifically realized through the following technical solutions:

[0013] The present invention discloses a small sample modulation recognition method based on a diffusion model and an attention mechanism, which includes:

[0014] Step 1, after converting the time-domain data in the data set into positive frequency data using the FFT algorithm, divide the data set into a training set, a test set, and a validation set according to the modulation type and a preset ratio; and divide out the training set under small sample conditions in the training set as the original data.

[0015] Step 2, input the original data into the diffusion model DDIM, and use the forward process of DDIM to gradually add noise to the original data through time steps for training to obtain pure Gaussian noise samples.

[0016] Step 3: The backward process of DDIM. An inverse denoising process is adopted to gradually denoise pure Gaussian noise samples to generate restored data;

[0017] Step 4: Train a classifier using pure Gaussian noise samples and the target modulation type corresponding to the task, calculate the logarithmic probability gradient of the classifier, and adjust the generation trajectory of the gradually denoised pure Gaussian noise samples based on the logarithmic probability gradient and the mean and covariance predicted by DDIM for the inverse denoising process under unconditional conditions to obtain synthetic data that conforms to the target modulation type and a trained DDIM;

[0018] Step 5: Construct a classifier model FATT by sequentially connecting a two-dimensional convolutional layer, a one-dimensional convolutional layer, a residual convolutional block, a multi-scale Fourier decomposition module, and an attention layer;

[0019] Take the original data, restored data, and synthetic data as input data and input them into FATT. After passing through the two-dimensional convolutional layer and the one-dimensional convolutional layer, enter the residual convolutional block. Add the input data of the residual convolutional block to the output data of the second convolution of the residual convolutional block, then enter the multi-scale Fourier decomposition module. Map the input data to the frequency domain through Fourier basis functions to explicitly capture the periodic and aperiodic features of the input data, thereby extracting comprehensive features containing frequency features and non-linear features, and splitting the comprehensive features into frequency features and non-linear features;

[0020] Through the attention layer, calculate the attention weights for each frequency band in the frequency features and perform weighting to obtain the weighted frequency features. Concatenate the weighted frequency features with the non-linear features to obtain concatenated features, and concatenate the concatenated features with the output of the residual convolutional block again to obtain a predicted modulation type vector; Visualize the predicted modulation type results of the predicted modulation type vector on the target modulation type through a confusion matrix to obtain a trained FATT;

[0021] Step 6: Input the test set and the validation set into the trained DDIM and FATT for testing and validation, and output the predicted modulation type results for small samples.

[0022] In Step 1, after converting the time-domain data in the dataset into positive-frequency data using the FFT algorithm, divide the dataset into a training set, a test set, and a validation set according to the modulation type and a preset ratio; The method of dividing a training set under small-sample conditions from the training set as the original data is as follows:

[0023] Use the FFT algorithm to convert the 128×2 time-domain data in the dataset into 256×1 frequency data, and perform absolute value processing on the frequency data to retain the positive-frequency data;

[0024] Divide the dataset into an 80% training set and a 20% test and validation set according to the modulation type and a preset ratio; among them, the training set under the small-sample condition accounts for 10% of the training set and serves as the original data.

[0025] In step two, the method of inputting the original data into the diffusion model DDIM and using the forward process of DDIM to gradually add noise to the original data through time steps for training to obtain a pure Gaussian noise sample includes:

[0026]

[0027] In the formula, X T is the pure Gaussian noise sample, T is the time step, X0 is the original data, ε is the Gaussian noise, and it follows the standard Gaussian distribution (0, I); is the product of the initial values of the noise attenuation coefficients, is the product of the noise attenuation coefficients, α t is the noise attenuation coefficient, and t is the time.

[0028] In step three, the backward process of DDIM adopts the reverse denoising process to gradually denoise the pure Gaussian noise sample to generate the restored data. The method includes:

[0029]

[0030] In the formula, X T-1 is the pure Gaussian noise sample X T gradually denoised to generate the restored data, α t-1 is the previous value of the noise attenuation coefficient α t ; ε Θ (X T , T) is the noise predicted by the diffusion model DDIM, is the variance.

[0031] In step four, train a classifier through the pure Gaussian noise sample and the target modulation type corresponding to the task, calculate the logarithmic probability gradient of the classifier, and adjust the generation trajectory of the pure Gaussian noise sample gradually denoised based on the logarithmic probability gradient and the mean and covariance predicted by DDIM for the reverse denoising process under the unconditional condition to obtain the synthetic data that meets the target modulation type and the trained DDIM. The method is:

[0032]

[0033] In the formula, is the mean predicted by the trained DDIM for the reverse denoising process under the unconditional condition, X′ T is the synthetic data, y is the target modulation type; μ θ (X T, T) is the mean predicted by DDIM for the reverse denoising process under unconditional conditions; ∑ θ (X T , T) is the covariance predicted by DDIM for the reverse denoising process under unconditional conditions; s is the guidance scale factor, used to control the influence degree of the classifier guidance on the generation trajectory; when s = 0, it is unconditional generation; when s > 0, it promotes the generated synthetic data to tend to the target modulation type; is the logarithmic probability gradient of the classifier, P φ (y∣X T ) is the classifier.

[0034] In step five, the Fourier basis function is:

[0035] Φ(x) = [cos(W P x) || sin(W P x) || σ(B P +W P x)]

[0036] In the formula, Φ(x) is the comprehensive feature extracted including frequency features and non-linear features, cos(W P x) || sin(W P x) is used to capture the periodic features of the input data x. Among them, cos(W P x) and sin(W P x) respectively capture the cosine component and sine component of the input data, directly embedding the periodicity in the Fourier series. W P is the projection matrix, used to adjust the frequency characteristics of the input data; σ(B P +W P x) is the non-linear activation function, used to capture the non-periodic features of the input data. B P is the bias term, σ represents the activation function, and || represents the concatenation operation, concatenating the sine component, cosine component and non-periodic features together.

[0037] In step five, the method of splitting the comprehensive feature into frequency features and non-linear features is:

[0038] F frep +F non =FanLayer(Φ(x))

[0039] In the formula, F frep is the frequency feature; F non is the non-linear feature; FanLayer(Φ(x)) is the comprehensive feature passing through the multi-scale Fourier decomposition module.

[0040] In step five, the method of calculating the attention weight for each frequency band in the frequency feature and performing weighting to obtain the weighted frequency feature is:

[0041]

[0042] Wherein, is the frequency feature of the i-th frequency band after weighting, and α i is the attention weight corresponding to the i-th frequency band in the frequency feature; is the frequency feature of the i-th frequency band, where i represents the number of the frequency band, and ω i is the attention score parameter corresponding to the i-th frequency band; using ω i to obtain the attention score corresponding to each frequency band; N is the total number of frequency features, and j is the number of the frequency feature; ω j is the attention score parameter corresponding to the j-th frequency feature, is the j-th frequency feature in each frequency band.

[0043] In step five, the method for splicing the weighted frequency feature and the non-linear feature to obtain the splicing feature is:

[0044]

[0045] Wherein, F out is the splicing feature; is the weighted frequency feature; F non is the non-linear feature.

[0046] In step five, the confusion matrix is:

[0047]

[0048] Wherein, C ab is the comparison between the true class and the predicted class in the confusion matrix, that is, the predicted modulation type vector, and is used to visually display the predicted modulation type result of the predicted modulation type vector on the target modulation type; wherein, a, b ∈ {1, 2,..., k}, k is the number of classes, that is, the number of samples with the true class of a predicted as class b; N is the total number of samples, n is the number of the sample, 1(.) is the indicator function, which takes the value of 1 when the condition in the parentheses is established, otherwise 0; y n is the true class of each sample; is the predicted class.

[0049] The beneficial effects of the present invention are:

[0050] 1. The technical solution disclosed by the present invention solves the technical problem of low recognition rate of modulation recognition under small sample conditions due to insufficient data set in Internet access and related services.

[0051] 2. The FFT algorithm is used to convert the time-domain data in the dataset into positive-frequency data, obtaining a 256x1-dimensional data vector, and the absolute value of the frequency is processed to retain the positive frequency. Such a preprocessing method does not destroy the original data features and has stronger expressiveness than the time-domain representation;

[0052] 3. The diffusion model DDIM is used as the generative model. Its forward process is transparent, and this transparency helps to explain the model's generation mechanism and the data generation process. The backward process uses the reverse denoising process to gradually denoise the pure Gaussian noise samples to generate restored data. This method accelerates the sampling generation, and the generated quality is not much different from that of DDPM; the DDIM backward process is more efficient and only requires fewer steps to achieve denoising and data generation; the generation process is stable, and the generated data quality is high;

[0053] Compared with DDPM, DDIM reduces the randomness in the generation process and improves the generation efficiency;

[0054] Compared with GAN, DDIM avoids the mode collapse problem that may occur during training, and the generated data has better diversity;

[0055] Compared with flipping and rotation, DDIM improves the diversity of the generated data and can simulate the complex distribution of the data;

[0056] The diffusion model DDIM enhances the training data by generating more diverse samples. The generation process is relatively stable and not prone to the problem of mode collapse. The step-by-step denoising generation process can better control the quality of the generated samples, making the generated samples closer to the real data. This method uses classifier-guided DDIM to generate high-quality samples under given conditions, which helps to solve the problem of insufficient small sample datasets.

[0057] 4. The classifier is trained through pure Gaussian noise samples and the target modulation type corresponding to the task, introducing classifier guidance for the generation process to provide guidance information on the target modulation type and adjust the generation process towards the target category;

[0058] Calculate the logarithmic probability gradient of the classifier, which is used to guide the reverse denoising process, making the generated synthetic data conform to the target category, and the generated data is closer to the real data in distribution;

[0059] Classifier guidance provides guidance information on the target conditions during the generation process, making the generation process more precisely adjusted towards the target category. Compared with ordinary conditional control, classifier guidance can dynamically adjust the generation process, improve the quality and category matching degree of the generated data, and the provided logarithmic probability gradient is directly used to guide the reverse denoising process, making the generated data more conform to the target category and improving the generation accuracy;

[0060] 5. The constructed classifier model FATT performs frequency-domain conversion on the input data through a multi-scale Fourier decomposition module, explicitly captures the periodic and aperiodic characteristics of the input data, extracts comprehensive features containing frequency features and non-linear features, and splits the comprehensive features into frequency features and non-linear features. Subsequently, the attention mechanism embedded in the attention layer automatically calculates attention weights for each frequency band feature and performs weighting, concatenates the weighted frequency features and non-linear features to obtain concatenated features, and concatenates the concatenated features with the output of the residual convolution block again to obtain a predicted modulation type vector, effectively focusing on the most discriminative frequency components, improving the discriminative ability of the signal under complex channel conditions and the accuracy of modulation recognition.

[0061] 6. This method not only alleviates the underfitting or overfitting phenomena that are prone to occur in deep learning models under small sample conditions, but also overcomes the phase interference and redundant information problems that may be brought by traditional rotation and flipping data augmentation methods, and at the same time avoids the mode collapse and gradient disappearance problems existing in the GAN generation process.

[0062] 7. Through a stable and efficient generation process and a significantly enhanced frequency-domain feature extraction ability, the present invention provides a solution with high accuracy, high robustness and good interpretability for wireless signal modulation recognition in Internet access and related services under insufficient data conditions. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] The present invention will be further described in detail below with reference to the drawings and embodiments.

[0064] Figure 1 It is a schematic block diagram of the small sample modulation recognition process provided by the present invention.

[0065] Figure 2 It is a schematic diagram of the general flowchart of small sample modulation recognition provided by the present invention.

[0066] Figure 3 It is a schematic diagram of the GAN small sample modulation recognition process provided by the present invention.

[0067] Figure 4 It is a schematic diagram of the noise addition process of the diffusion model provided by the present invention.

[0068] Figure 5 It is a schematic diagram of the denoising process of the diffusion model provided by the present invention.

[0069] Figure 6 It is a schematic diagram of the overall flowchart of the small sample modulation recognition model of the diffusion model and the attention mechanism provided by the present invention.

[0070] Figure 7 It is a schematic diagram of the modulation recognition model flowchart input to the classifier model after data augmentation by the diffusion model provided by the present invention.

[0071] Figure 8 It is a schematic diagram of the 8PSK generated signal at 16 dB provided by the present invention.

[0072] Figure 9 It is a schematic diagram of the AM-DSB generated signal at 16 dB provided by the present invention.

[0073] Figure 10 It is a schematic diagram of the BPSK generated signal at 16 dB provided by the present invention.

[0074] Figure 11 It is a schematic diagram of the CPFSK generated signal at 16 dB provided by the present invention.

[0075] Figure 12 It is a schematic diagram of the QPSK generated signal at 16 dB provided by the present invention.

[0076] Figure 13 It is a schematic diagram of the GFSK generated signal at 16 dB provided by the present invention.

[0077] Figure 14 It is a schematic diagram of the PAM4 generated signal at 16 dB provided by the present invention.

[0078] Figure 15 It is a schematic diagram of the QAM16 generated signal at 16 dB provided by the present invention.

[0079] Figure 16 It is a schematic diagram of the QAM64 generated signal at 16 dB provided by the present invention.

[0080] Figure 17 It is a schematic diagram of the WBFM generated signal at 16 dB provided by the present invention.

[0081] Figure 18 It is a schematic diagram of the classification recognition rate of the modulated signals of the validation set and the test set after data augmentation under the deep learning model provided by the present invention.

[0082] Figure 19 It is a schematic diagram of the confusion matrix of the validation set and the test set at 10 dB after data augmentation provided by the present invention. Detailed implementation manners

[0083] An embodiment of the present invention discloses a few-shot modulation recognition method based on a diffusion model and an attention mechanism, and the method includes:

[0084] Step 1: After converting the time-domain data in the dataset into positive-frequency data using the FFT algorithm, divide the dataset into a training set, a test set, and a validation set according to the modulation type and a preset ratio; and divide the training set under the small-sample condition in the training set as the original data;

[0085] For example, the dataset is RML2016.10b, including 1,200,000 modulated radio signal samples, evenly distributed in 10 different modulation categories. Each modulation category in the dataset is represented by 20 signal-to-noise ratio (SNR) levels, ranging from -20 dB to 18 dB, with an increment of 2 dB.

[0086] Step 2: Input the original data into the diffusion model DDIM, and use the forward process of DDIM to gradually add noise to the original data through time steps for training to obtain pure Gaussian noise samples;

[0087] Step 3: In the backward process of DDIM, adopt the reverse denoising process to gradually denoise the pure Gaussian noise samples to generate restored data; in the backward process of DDIM, start from the pure Gaussian noise samples and gradually predict and remove the noise, so that the noise data gradually restores to samples that conform to the real data distribution. This generation process can generate data with high quality and strong realism.

[0088] Step 4: Train a classifier through the pure Gaussian noise samples and the target modulation type corresponding to the task, calculate the logarithmic probability gradient of the classifier, and adjust the generation trajectory of the pure Gaussian noise samples gradually denoising based on the logarithmic probability gradient and the mean and covariance predicted by DDIM for the reverse denoising process under the unconditional condition to obtain synthetic data that conforms to the target modulation type and the trained DDIM;

[0089] This step needs to train a classifier P T for samples with different noise levels: pure Gaussian noise samples X φ (y∣X T ). This classifier learns how to judge whether a sample conforms to the target modulation type on the noise data. In the reverse denoising process, when generating samples at each step, it is hoped that the generated samples can be more inclined to the target condition. For this reason, it is necessary to calculate the logarithmic probability gradient of the current sample X T with respect to the target condition: which indicates how to adjust X T so that the classifier thinks it is more in line with the target modulation type y. Based on the logarithmic probability gradient and the mean and covariance predicted by DDIM for the reverse denoising process under the unconditional condition, adjust the generation trajectory of the pure Gaussian noise samples gradually denoising. In each step of the reverse denoising process, the gradient signal of the logarithmic probability gradient from the classifier is utilized, so that the finally generated synthetic data is more in line with the target modulation type.

[0090] Step 5: Sequentially connect a two-dimensional convolutional layer, a one-dimensional convolutional layer, a residual convolutional block, a multi-scale Fourier decomposition module, and an attention layer to construct a classifier model FATT;

[0091] Take the original data, restored data, and synthesized data as input data and input them into FATT. After passing through the two-dimensional convolutional layer and the one-dimensional convolutional layer, enter the residual convolutional block. After adding the input data of the residual convolutional block to the output data of the second convolution of the residual convolutional block, enter the multi-scale Fourier decomposition module. Map the input data to the frequency domain through Fourier basis functions to explicitly capture the periodic and aperiodic characteristics of the input data, thereby extracting comprehensive features containing frequency features and non-linear features, and splitting the comprehensive features into frequency features and non-linear features;

[0092] Through the attention layer, calculate the attention weights for each frequency band in the frequency features and perform weighting to obtain the weighted frequency features. Concatenate the weighted frequency features with the non-linear features to obtain the concatenated features, and concatenate the concatenated features with the output of the residual convolutional block again to obtain the predicted modulation type vector; Visualize the predicted modulation type results of the predicted modulation type vector on the target modulation type through a confusion matrix to obtain the trained FATT;

[0093] Step 6: Input the test set and the validation set into the trained DDIM and FATT for testing and validation, and output the small-sample predicted modulation type results.

[0094] In Step 1, after converting the time-domain data in the dataset into positive-frequency data using the FFT algorithm, divide the dataset into a training set, a test set, and a validation set according to the modulation type and a preset ratio; The method of dividing a training set under small-sample conditions from the training set as the original data is as follows:

[0095] Use the FFT algorithm to convert the 128×2 time-domain data in the dataset into 256×1 frequency data, and perform absolute value processing on the frequency data to retain the positive-frequency data;

[0096] Divide the dataset into an 80% training set and a 20% test set and validation set according to the modulation type and a preset ratio; Among them, the training set under small-sample conditions accounts for 10% of the training set and is used as the original data.

[0097] As Figure 4 shown, in Step 2, the method of inputting the original data into the diffusion model DDIM and using the forward process of DDIM to gradually add noise to the original data through time steps to obtain a pure Gaussian noise sample includes:

[0098]

[0099] where X Tis a pure Gaussian noise sample, T is the time step, for example, T = 1000, X0 is the original data, ε is the Gaussian noise, following the standard Gaussian distribution (0, I); is the product of the initial values of the noise attenuation coefficients, is the product of the noise attenuation coefficients, α t is the noise attenuation coefficient, similar to a hyperparameter, only sampling the noise once, X can be directly obtained from X0 T , and t is the time.

[0100] As Figure 5 shown, in step three, for the backward process of DDIM, an inverse denoising process is adopted to gradually denoise the pure Gaussian noise sample to generate the restored data. The method includes:

[0101]

[0102] In the formula, X T-1 is the pure Gaussian noise sample X T The restored result of gradual denoising. After gradual denoising through the time step T, the finally generated restored data is the original data X0, and α t-1 is the previous value of the noise attenuation coefficient α t ; ε Θ (X T , T) is the noise predicted by the diffusion model DDIM, is the variance.

[0103] In step four, a classifier is trained through the pure Gaussian noise sample and the target modulation type corresponding to the task, the logarithmic probability gradient of the classifier is calculated, and based on the logarithmic probability gradient and the mean and covariance predicted by DDIM for the inverse denoising process under unconditional conditions, the generation trajectory of the gradually denoising pure Gaussian noise sample is adjusted to obtain the synthetic data that conforms to the target modulation type and the trained DDIM. The method is:

[0104]

[0105] In the formula, is the mean predicted by the trained DDIM for the inverse denoising process under unconditional conditions, X' T is the synthetic data, y is the target modulation type; μ θ (X T , T) is the mean predicted by DDIM for the inverse denoising process under unconditional conditions; ∑ θ (X T , T) is the covariance predicted by DDIM for the inverse denoising process under unconditional conditions; s is the guidance scale factor, used to control the influence degree of the classifier to guide the generation trajectory; when s = 0, it is unconditional generation; when s > 0, it promotes the generated synthetic data to tend to the target modulation type; is the logarithmic probability gradient of the classifier, P φ (y|X T ) is the classifier.

[0106] In step five, the Fourier basis function is:

[0107] Φ(x) = [cos(W P x) || sin(W P x) || σ(B P +W P x)]

[0108] In the formula, Φ(x) is the comprehensive feature extracted, including frequency features and non-linear features. cos(W P x) || sin(W P x) is used to capture the periodic features of the input data x. Among them, cos(W P x) and sin(W P x) capture the cosine component and sine component of the input data respectively, directly embedding the periodicity in the Fourier series. W P is the projection matrix, used to adjust the frequency characteristics of the input data. By flattening the input data and multiplying it with the fixed projection matrix W P , the cosine component and sine component are calculated; σ(B P +W P x) is the non-linear activation function, used to capture the non-periodic features of the input data. B P is the bias term, σ represents the activation function, and || represents the concatenation operation, concatenating the sine component, cosine component and non-periodic features together.

[0109] In step five, the method of splitting the comprehensive feature into frequency features and non-linear features is:

[0110] F frep +F non = FanLayer(Φ(x))

[0111] In the formula, F frep is the frequency feature; F non is the non-linear feature; FanLayer(Φ(x)) is the comprehensive feature passing through the multi-scale Fourier decomposition module.

[0112] In step five, the attention weights are calculated and weighted for each frequency band in the frequency feature. The purpose is to assign different "importance levels" to different positions (different frequency bands, etc.) in the input features through learnable weights, so as to achieve focusing on key regions or channels and suppressing irrelevant or redundant information. The method is:

[0113]

[0114]

[0115] In the formula, is the frequency feature of the i-th frequency band after weighting, and α i is the attention weight corresponding to the i-th frequency band in the frequency feature, and its value reflects the importance of this frequency band; is the frequency feature of the i-th frequency band, where i represents the number of the frequency band, and ω i is the attention score parameter corresponding to the i-th frequency band; using ω i to obtain the attention score corresponding to each frequency band; N is the total number of frequency features, and j is the number of the frequency feature; ω j is the attention score parameter corresponding to the j-th frequency feature, is the j-th frequency feature in each frequency band.

[0116] In step five, the method of splicing the weighted frequency feature and the non-linear feature to obtain the spliced feature is as follows:

[0117]

[0118] In the formula, F out is the spliced feature; is the weighted frequency feature; F non is the non-linear feature. It overcomes the deficiency of the existing network model in extracting time-frequency features.

[0119] In step five, the confusion matrix is a visualization tool for evaluating the performance of a classification model. It shows the prediction results of the model on each category in the form of a table, so as to help analyze the correct classification and misclassification situations of the model on different categories. The specific formula is:

[0120]

[0121] In the formula, C ab is the comparison between the true category and the predicted category in the confusion matrix, that is, the predicted modulation type vector, which is used to visually display the predicted modulation type result of the predicted modulation type vector on the target modulation type; where a, b ∈ {1, 2,..., k}, k is the number of categories, that is, the number of samples with the true category a that are predicted as category b; N is the total number of samples, n is the number of the sample, 1(.) is the indicator function, which takes the value of 1 when the condition in the parentheses holds, otherwise 0; y n is the true category of each sample; is the predicted category.

[0122] Such as Figures 6 to 7As shown, a classifier model FATT is constructed by sequentially connecting a two-dimensional convolutional layer, a one-dimensional convolutional layer, a residual convolutional block, a multi-scale Fourier decomposition module, and an attention layer. The original data, restored data, and synthetic data are used as input data. After preprocessing such as data augmentation, chunking, and linear embedding, the input format required by the FATT network is obtained. The input signal dimension of the FATT network is [batch_size, 2, 1024]. First, the input data passes through the two-dimensional convolutional layer Conv1 with a convolutional kernel size of (2, 7) to extract features in the local region of the time series, mapping the two input channels to 1024 feature channels. Its output dimension becomes [batch_size, 1024, 1, 64]. Subsequently, by removing the third dimension (size 1), the output dimension is obtained as [batch_size, 1024, 64]. A batch normalization BatchNorm2d and a LeakyReLU activation function are appended after the convolutional layer to enhance the non-linear expression ability and numerical stability. In the residual convolutional block, the input signal [batch_size, 1024, 64] passes through the first Conv1D layer with 64 filters, a convolutional kernel size of 3, no bias is used, and L2 regularization is applied. After convolution, a BatchNormalization batch normalization layer and a LeakyReLU activation function are connected. Subsequently, through a Dropout layer, some neurons are randomly discarded at a ratio of 0.3 to prevent overfitting. Passing through the Conv1D layer again, the number of filters remains 64, the convolutional kernel size is 3, no bias is used, and L2 regularization is applied. The input is added to the output of the second convolution (skip connection) and passed through the LeakyReLU activation function, and the output dimension remains [batch_size, 1024, 64]. After passing through the residual convolutional block, the signal enters the multi-scale Fourier decomposition module. The input signal [batch_size, 1024, T] first receives the input feature [batch_size, 1024, 64], and through a random projection matrix and a trainable linear mapping, a comprehensive feature containing frequency features and non-linear features is generated. The specific operations include flattening the input and multiplying it by a fixed projection matrix, calculating the cosine component and the sine component, and then generating non-linear features through a trainable linear mapping. The output shape of the FanLayer is [batch_size, 1024, 128], where 128 = 2 * proj_dim(32) + fan_out(64).Split the output of FanLayer into frequency features [batch_size, 1024, 64] and non - linear features [batch_size, 1024, 64]. Weight the frequency features through the FrequencyAttention layer to generate attention - enhanced frequency features [batch_size, 1024, 64]. Concatenate the weighted frequency features with the non - linear features to obtain [batch_size, 1024, 128]. Subsequently, concatenate this concatenated feature with the output [batch_size, 1024, 64] of the residual convolution block to form a comprehensive feature of [batch_size, 1024, 192]. Generate an attention gate with a shape of [batch_size, 1024, 64] through a 1x1 convolution (Conv1D). Then, fuse the features through element - wise multiplication and addition operations to obtain the fused feature [batch_size, 1024, 64]. Perform global average pooling on the fused feature to obtain a feature vector with a shape of [batch_size, 64]. Use a Dense layer to expand the feature dimension from 64 to 256, with the activation function being LeakyReLU. Subsequently, connect a Dropout layer to randomly discard some neurons at a ratio of 0.5. Use a Dense layer to map the features to the target classification dimension num_classes, with the activation function being softmax, and the output dimension being [batch_size, num_classes].

[0123] The confusion matrix is a visualization tool for evaluating the performance of a classification model. It shows the prediction results of the model for each class in the form of a table, thus helping to analyze the correct and incorrect classification situations of FATT for different classes. Use the confusion matrix to represent the comparison between the true class and the predicted class, show the classification results of the FATT model for each class, and obtain the trained FATT.

[0124] To illustrate the technical solution of the present invention in detail, specific examples are provided for elaboration:

[0125] Pre - processing of the dataset in Example 1

[0126] 1. Use the RML2016.10a dataset, which contains 220,000 modulated radio signal samples, distributed in 11 different modulation classes. Each modulation class has 20 signal-to-noise ratio (SNR) levels, ranging from -20 dB to 18 dB, with an increment of 2 dB. The dataset has a time series dimension of 128x2. Or use the RML2016.10b dataset, which includes 120,000 modulated radio signal samples, evenly distributed in 10 different modulation classes. Each modulation class in the dataset is represented by 20 signal-to-noise ratio (SNR) levels, ranging from -20 dB to 18 dB, with an increment of 2 dB.

[0127] 2. Through the FFT algorithm, convert the data in the dataset from the time domain to the frequency domain, obtaining data with a dimension of 256x1. Take the absolute value of the frequency domain data and only retain the positive frequency part.

[0128] 3. Divide the dataset such that the training set under the small sample condition accounts for 10% of the training set, and the test set and validation set together account for 20%.

[0129] Example 2 Generator Training and Results

[0130] 1. During the training process of the DDIM model, introduce classifier guidance to provide guiding information for the target condition. The classifier guides the reverse denoising process through gradient information, making the generated data conform to the target category. This classifier has been trained with the complete dataset.

[0131] 2. The DDIM forward process performs noise addition training with a total number of steps T = 1000 on a segment of the signal. After adding Gaussian white noise, the result is pure Gaussian noise, and ε follows the standard Gaussian distribution (0, I). Through the linear Gaussian property and the Markov chain property, the representation formula of X can be obtained. T representation formula.

[0132] 3. In the DDIM reverse process, the Unet network predicts the noise learned from the denoising model that gradually denoises from pure noise to recover the original data. Using the noise attenuation coefficient, the final generated data X0 is obtained after T steps of denoising. Among them, the Batchsize is set to 512, the learning rate is 0.0001, and running on the Pycharm platform can obtain the IQ signal result graph of the generated signal of 8PSK at 16 dB as shown in Figure 8 where the vertical axis is the amplitude of the IQ signal and the horizontal axis is the time step of 128; Figure 9 the IQ signal result graph of the generated signal of AM-DSB at 16 dB as shown in, where the vertical axis is the amplitude of the IQ signal and the horizontal axis is the time step of 128; Figure 10 the IQ signal result graph of the generated signal of BPSK at 16 dB as shown in, where the vertical axis is the amplitude of the IQ signal and the horizontal axis is the time step of 128; Figure 11IQ signal result diagram of CPFSK generated signal at 16 dB, where the vertical axis is the amplitude of the IQ signal and the horizontal axis is the time step of 128; Figure 12 IQ signal result diagram of QPSK generated signal at 16 dB, where the vertical axis is the amplitude of the IQ signal and the horizontal axis is the time step of 128; Figure 13 IQ signal result diagram of GFSK generated signal at 16 dB, where the vertical axis is the amplitude of the IQ signal and the horizontal axis is the time step of 128; Figure 14 IQ signal result diagram of PAM4 generated signal at 16 dB, where the vertical axis is the amplitude of the IQ signal and the horizontal axis is the time step of 128; Figure 15 IQ signal result diagram of QAM16 generated signal at 16 dB, where the vertical axis is the amplitude of the IQ signal and the horizontal axis is the time step of 128; Figure 16 IQ signal result diagram of QAM64 generated signal at 16 dB, where the vertical axis is the amplitude of the IQ signal and the horizontal axis is the time step of 128; Figure 17 IQ signal result diagram of WBFM generated signal at 16 dB, where the vertical axis is the amplitude of the IQ signal and the horizontal axis is the time step of 128.

[0133] Classifier training and effect in Example 3

[0134] 1. Preprocess the original data, restored data, and synthesized data, and convert them into an input format suitable for FATT (256x1).

[0135] 2. With the input data of 256x1, after preprocessing steps such as data augmentation, chunking, and encoding, the Fourier analysis network of the multi-scale Fourier decomposition module explicitly performs frequency domain conversion on the input data to extract comprehensive features containing frequency features and non-linear features; subsequently, the attention mechanism embedded in the attention layer automatically calculates the importance weights for each frequency band feature, effectively focusing on the most discriminative frequency components.

[0136] 3. Use the confusion matrix to represent the comparison between the true class and the predicted class, and show the performance of FATT on each class. As Figure 18 shown, at a signal-to-noise ratio of 16 dB, the recognition rate of the Transformer classifier model without data augmentation is lower than that of the classifier model after data augmentation, while the classifier model FATT proposed in the present invention is superior to other classifier models after data augmentation by the diffusion model. As Figure 19As shown in the confusion matrix at 10 dB, except for the confusion of WBFM modulation types, the recognition results of other modulation type signals are all above 90%, indicating that the diffusion model has a very poor generation effect on WBFM modulation type signals, resulting in the inability to generate signals similar to the original dataset, leading to poor training effect of the classifier model, and causing confusion of WBFM modulation type signals.

[0137] In summary, the experimental results in the examples show that the small-sample modulation recognition method based on the diffusion model and attention mechanism proposed by the present invention effectively solves the problem of difficult wireless signal modulation recognition caused by insufficient datasets in the field of modulation carrier systems.

[0138] In this method, by introducing a classifier-guided DDIM reverse denoising generation process, the generated samples are gradually adjusted towards the target modulation type during the noise removal process. The gradient information calculated by the pre-trained classifier is used to make the generated data closer to the real data in distribution. At the same time, the time-domain signal is converted into a 256×1-dimensional frequency-domain representation through the FFT algorithm transformation, and the absolute value of this representation is taken to retain the positive frequency information. Then, combined with the embedded attention mechanism, the importance of each frequency band in the frequency domain is adaptively weighted, thus significantly capturing key feature information including amplitude, phase, and periodicity, and improving the discrimination ability of signals under complex channel conditions.

[0139] Experimental data shows that under the condition of 16 dB signal-to-noise ratio, after data augmentation by the diffusion model, the modulation recognition accuracy of the classifier model is significantly better than that of the non-augmented data. For all modulation signals except the WBFM modulation type, the accuracy rate reaches more than 90%, and the confusion matrix at 10 dB further proves the advantage of this method in a low signal-to-noise ratio environment.

[0140] Generally speaking, this method not only alleviates the underfitting or overfitting phenomena that are prone to occur in deep learning models under small-sample conditions, but also overcomes the phase interference and redundant information problems that may be brought by traditional rotation and flipping data augmentation methods, and at the same time avoids the mode collapse and gradient disappearance problems existing in the GAN generation process.

[0141] Generally speaking, the present invention provides a solution with high accuracy, high robustness, and good interpretability for wireless signal modulation recognition in the case of insufficient data through a stable and efficient generation process and significantly enhanced frequency-domain feature extraction ability.

[0142] The beneficial effects of the embodiments of the present invention are:

[0143] 1. The technical solution disclosed by the present invention solves the technical problem of low modulation recognition rate under small-sample conditions due to insufficient datasets in Internet access and related services.

[0144] 2. Use the FFT algorithm to convert the time-domain data in the dataset into positive-frequency data, obtaining a 256x1-dimensional data vector, and perform absolute value processing on the frequency, retaining the positive frequency. Such a preprocessing method does not damage the original data features and has stronger expressiveness than the time-domain representation;

[0145] 3. Use the diffusion model DDIM as the generative model. Its forward process is transparent, and this transparency helps to explain the model's generation mechanism and the data generation process. The backward process adopts the reverse denoising process, gradually denoising the pure Gaussian noise samples to generate the restored data. This method accelerates the sampling generation, and the generated quality is not much different from that of DDPM; the DDIM backward process is more efficient and can achieve denoising and data generation with fewer steps; the generation process is stable, and the generated data has high quality;

[0146] Compared with DDPM, DDIM reduces the randomness in the generation process and improves the generation efficiency;

[0147] Compared with GAN, DDIM avoids the mode collapse problem that may occur during the training process, and the generated data has better diversity;

[0148] Compared with flipping and rotation, DDIM improves the diversity of the generated data and can simulate the complex distribution of the data;

[0149] The diffusion model DDIM enhances the training data by generating more diverse samples. The generation process is relatively stable and not prone to the problem of mode collapse. The step-by-step denoising generation process can better control the quality of the generated samples, making the generated samples closer to the real data. This method uses classifier-guided DDIM to generate high-quality samples under given conditions, which helps to solve the problem of insufficient small-sample datasets.

[0150] 4. Train a classifier through pure Gaussian noise samples and the target modulation type corresponding to the task, introduce classifier guidance for the generation process, provide guidance information on the target modulation type, and adjust the generation process towards the target category;

[0151] Calculate the logarithmic probability gradient of the classifier, which is used to guide the reverse denoising process, so that the generated synthetic data conforms to the target category, and the generated data is closer to the real data in distribution;

[0152] Classifier guidance provides guidance information on the target conditions during the generation process, making the generation process more precisely adjusted towards the target category. Compared with ordinary conditional control, classifier guidance can dynamically adjust the generation process, improve the quality and category matching degree of the generated data, and the provided logarithmic probability gradient is directly used to guide the reverse denoising process, making the generated data more conform to the target category and improving the generation accuracy;

[0153] 5. The constructed classifier model FATT performs frequency-domain conversion on the input data through a multi-scale Fourier decomposition module, explicitly captures the periodic and non-periodic features of the input data, extracts comprehensive features containing frequency features and non-linear features, and splits the comprehensive features into frequency features and non-linear features. Subsequently, the attention mechanism embedded in the attention layer automatically calculates attention weights for each frequency band feature and performs weighting, concatenates the weighted frequency features and non-linear features to obtain a concatenated feature, and concatenates the concatenated feature with the output of the residual convolution block again to obtain a predicted modulation type vector, effectively focusing on the most discriminative frequency components, and improving the discriminative ability of the signal under complex channel conditions and the accuracy of modulation recognition.

[0154] 6. This method not only alleviates the underfitting or overfitting phenomena that are prone to occur in deep learning models under small sample conditions, but also overcomes the phase interference and redundant information problems that may be brought by traditional rotation and flipping data augmentation methods, and at the same time avoids the mode collapse and gradient disappearance problems existing in the GAN generation process.

[0155] 7. Through a stable and efficient generation process and a significantly enhanced frequency-domain feature extraction ability, the present invention provides a solution with high accuracy, high robustness and good interpretability for wireless signal modulation recognition in Internet access and related services under insufficient data conditions.

[0156] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A small sample modulation recognition method based on diffusion model and attention mechanism, characterized in that: The method includes: Step 1: After converting the time domain data in the data set into positive frequency data using the FFT algorithm, the data set is divided into a training set, a test set, and a validation set by modulation type and preset ratio; and a training set under a small sample condition is divided out from the training set as the original data; Step 2: Input the original data into the diffusion model DDIM, and use the forward process of DDIM to gradually add noise to the original data through time steps to obtain pure Gaussian noise samples; Step 3, the backward process of DDIM uses the inverse denoising process to gradually denoise the pure Gaussian noise samples to generate recovered data; Step 4: Train the classifier using pure Gaussian noise samples and the target modulation type corresponding to the task, calculate the log probability gradient of the classifier, and adjust the generation trajectory of the stepwise denoising of the pure Gaussian noise samples based on the log probability gradient and the mean and covariance predicted by DDIM for the inverse denoising process without conditions, to obtain synthetic data that meets the target modulation type and the trained DDIM; Step 5: The classifier model FATT is constructed by sequentially connecting a two-dimensional convolutional layer, a one-dimensional convolutional layer, a residual convolution block, a multi-scale Fourier decomposition module, and an attention layer; The original data, restored data and synthesized data are taken as input data and input into FATT. After passing through the two-dimensional convolution layer and the one-dimensional convolution layer, they enter the residual convolution block. After adding the input data of the residual convolution block and the output data of the second convolution layer of the residual convolution block, they enter the multi-scale Fourier decomposition module. The input data is mapped to the frequency domain through the Fourier basis function, and the periodicity and non-periodicity characteristics of the input data are explicitly captured, thereby extracting the comprehensive features including frequency features and nonlinear features, and splitting the comprehensive features into frequency features and nonlinear features. Through the attention layer, the attention weight is calculated for each frequency band in the frequency feature and weighted to obtain the weighted frequency feature. The weighted frequency feature is spliced ​​with the nonlinear feature to obtain the spliced ​​feature, and the spliced ​​feature is spliced ​​with the output of the residual convolution block again to obtain the predicted modulation type vector; the predicted modulation type result of the predicted modulation type vector on the target modulation type is visualized through the confusion matrix to obtain the trained FATT; Step 6: Input the test set and validation set into the trained DDIM and FATT for testing and validation, and output the small sample prediction modulation type results.

2. The method according to claim 1, characterized in that In step 1, after converting the time domain data in the data set into positive frequency data using the FFT algorithm, the data set is divided into a training set, a test set, and a validation set by modulation type and preset ratio; and the method of dividing the training set under the small sample condition as the original data in the training set is: The 128×2 time domain data in the data set is converted into 256×1 frequency data using the FFT algorithm, and the frequency data is processed with absolute values ​​to retain the positive frequency data; The data set is divided into 80% training set and 20% test set and validation set according to the modulation type and preset ratio; among them, the training set under small sample conditions accounts for 10% of the training set and serves as the original data.

3. The method according to claim 2, characterized in that In step 2, the original data is input into the diffusion model DDIM, and the forward process of DDIM is used to gradually add noise to the original data through time steps to obtain a pure Gaussian noise sample. The method includes: ; In the formula, is a pure Gaussian noise sample, T is the time step, is the original data, ε is Gaussian noise, which obeys the standard Gaussian distribution (0, I); is the multiplication of the initial value of the noise attenuation coefficient, is the product of the noise attenuation coefficient, is the noise attenuation coefficient, and t is the time.

4. The method according to claim 3, characterized in that In step 3, the backward process of DDIM uses an inverse denoising process to gradually denoise the pure Gaussian noise samples, and the method for generating restored data includes: )+ ; In the formula, is a pure Gaussian noise sample The restored data generated by stepwise denoising, is the noise attenuation coefficient The previous value of ; is the noise predicted by the diffusion model DDIM, is the variance.

5. The method according to claim 4, characterized in that In step 4, the classifier is trained by pure Gaussian noise samples and the target modulation type corresponding to the task, the log probability gradient of the classifier is calculated, and the generation trajectory of the stepwise denoising of the pure Gaussian noise samples is adjusted based on the log probability gradient and the mean and covariance predicted by DDIM for the inverse denoising process without conditions. The method to obtain synthetic data that meets the target modulation type and the trained DDIM is as follows: ,and)= +s ; In the formula, ,y) is the mean value of the unconditional prediction of the trained DDIM for the inverse denoising process, For synthetic data, is the target modulation type; is the mean value of DDIM's prediction of the inverse denoising process without conditions; is the covariance of DDIM's prediction of the inverse denoising process without conditions; is the guiding scale factor, which is used to control the influence of the classifier in guiding the generated trajectory; when When , it is generated unconditionally; When , the generated synthetic data is pushed toward the target modulation type; is the log probability gradient of the classifier, For the classifier.

6. The method according to claim 5, characterized in that In step 5, the Fourier basis function is: ; In the formula, is the extracted comprehensive feature including frequency feature and nonlinear feature, x represents input data, and is the input IQ signal; Used to capture the periodic characteristics of the input data x, where and The cosine and sine components of the input data are captured separately, directly embedding the periodicity in the Fourier series. is the projection matrix, which is used to adjust the frequency characteristics of the input data; is a nonlinear activation function used to capture the non-periodic characteristics of the input data. is the bias term, σ represents the activation function, and || represents the concatenation operation, which concatenates the sine component, cosine component and non-periodic features together.

7. The method according to claim 6, characterized in that In step 5, the method of splitting the comprehensive features into frequency features and nonlinear features is: ; In the formula, is the frequency characteristic; It is a nonlinear feature; To integrate features, a multi-scale Fourier decomposition module is used.

8. The method according to claim 7, characterized in that In step 5, the attention weight is calculated for each frequency band in the frequency feature and weighted. The method for obtaining the weighted frequency feature is: * ; ; In the formula, is the frequency characteristic of the i-th frequency band after weighting, is the attention weight corresponding to the i-th frequency band in the frequency feature; is the frequency characteristic of the i-th frequency band, where i represents the number of the frequency band, is the attention score parameter corresponding to the i-th frequency band; using Get the attention score corresponding to each frequency band; N is the total number of frequency features, j is the number of the frequency feature; is the attention score parameter corresponding to the j-th frequency feature, is the jth frequency feature in each frequency band.

9. The method according to claim 8, characterized in that In step 5, the method of concatenating the weighted frequency feature and the nonlinear feature to obtain the concatenated feature is: ; In the formula, is the splicing feature; is a feature concatenation function, used to concatenate weighted frequency features and nonlinear features; is the weighted frequency feature; It is a non-linear feature.

10. The method according to claim 9, characterized in that In step 5, the confusion matrix is: ; In the formula, is the comparison between the real category and the predicted category in the confusion matrix, that is, the predicted modulation type vector, which is used to visualize the predicted modulation type result of the predicted modulation type vector on the target modulation type; among them, , k is the number of categories, that is, the number of samples whose true category is a that are predicted to be category b; N is the total number of samples, n is the sample number, 1(.) is the indicator function, which takes the value of 1 when the condition in the brackets is met, otherwise it takes the value of 0; is the true category of each sample; is the predicted category.

Citation Information

Patent Citations

  • Automatic identification method for radar signal modulation type

    CN114564982A

  • Small sample communication signal automatic modulation identification method based on incremental learning

    CN114580484A