Small sample modulation identification method based on diffusion model and attention mechanism

By adopting a small sample modulation recognition method based on diffusion model and attention mechanism in wireless signal modulation recognition, the problem of low recognition rate caused by insufficient data set is solved, and the modulation recognition effect with high accuracy and robustness is achieved.

CN119996133AActive Publication Date: 2025-05-13PLA PEOPLES LIBERATION ARMY OF CHINA STRATEGIC SUPPORT FORCE AEROSPACE ENG UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510221985.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-05-13
Estimated Expiration
2045-02-26

AI Technical Summary

Technical Problem

In Internet access and related services, due to insufficient data sets, the wireless signal modulation recognition recognition rate under small sample conditions is low.

Method used

A small sample modulation recognition method based on diffusion model and attention mechanism is adopted to generate high-quality samples through classifier-guided DDIM, enhance training data, and extract frequency domain features through multi-scale Fourier decomposition module and attention layer.

Benefits of technology

The modulation recognition accuracy under small sample conditions is improved, underfitting or overfitting is alleviated, and the phase interference and redundant information problems of traditional data augmentation methods are avoided, as well as the problem of pattern crashes and gradient disappearance in the GAN generation process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119996133A_ABST
    Figure CN119996133A_ABST
Patent Text Reader

Abstract

The invention discloses a small sample modulation recognition method based on a diffusion model and an attention mechanism, belongs to the field of modulation carrier systems, solves the problem of low modulation recognition rate under the condition of small samples, and comprises the following steps: inputting original data into DDIM for gradual noise adding training, and gradually denoising by adopting a reverse denoising process to generate recovery data; based on logarithm probability gradient adjustment of a classifier, generating a track to obtain synthetic data and a trained DDIM; the original data, the recovery data and the synthetic data are input into FATT and mapped to a frequency domain through a Fourier basis function, periodic and non-periodic features are explicitly captured, and comprehensive features are extracted; splitting the comprehensive feature into a frequency feature and a nonlinear feature through an attention layer, splicing the weighted frequency feature and the nonlinear feature, and splicing the frequency feature and the nonlinear feature with the output of the residual convolution block again to obtain a prediction modulation type vector; a modulation type prediction result is visually displayed through the confusion matrix, and a trained FATT is obtained; according to the invention, the modulation identification accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of modulation carrier systems, in the technical field of small sample modulation recognition of wireless signals, and relates to a small sample modulation recognition method based on a diffusion model and an attention mechanism. Background Art

[0002] In the field of modulation carrier systems, especially in the research on small sample modulation recognition of wireless signals, in Internet access and related services, due to insufficient data sets, the neural network in deep learning is underfitted and the recognition effect is poor, resulting in the problem of low modulation recognition rate under small sample conditions. How to solve the problem of insufficient data sets is a major difficulty in the field of small sample modulation recognition of wireless signals.

[0003] In the prior art, some scholars believe that the use of classifiers in modulated carrier systems, such as SVM, Bayesian reasoning, sparse recovery and other machine learning methods, can to a certain extent solve the problem of low modulation recognition rate under small sample conditions in Internet access and related services. However, machine learning methods rely heavily on a large amount of data to find the optimal solution. This approach is highly interpretable. However, when data is insufficient, using machine learning for optimization often fails to find the optimal solution and the effect is poor.

[0004] Although the existing AMC technology performs well in certain specific scenarios, in actual applications, the modulated carrier system often faces challenges from complex channel environments and multiple modulation methods, especially under small sample conditions. These problems are particularly prominent. On the one hand, channel defects such as incomplete channel state information, carrier frequency offset, symbol timing offset, and phase offset in Internet access and related services significantly reduce the classification performance of modulated signals. On the other hand, under small sample conditions, existing deep learning models, such as convolutional neural networks (CNN) and long short-term memory networks (LSTM), find it difficult to obtain sufficient labeled training data to capture the feature differences of different modulation formats, resulting in insufficient model generalization capabilities and low classification accuracy. Figure 1 Shown is the basic method for automatic modulation recognition under small sample conditions.

[0005] like Figure 2As shown in the figure, a general flow chart of small sample modulation recognition in the field of modulated carrier systems, a flow chart of small sample modulation recognition using data augmentation methods such as rotation and flipping, Liang et al. published L. Huang, W. Pan, Y. Zhang, L. Qian, N. Gao and Y. Wu, "Data Augmentation for Deep Learning-Based Radio Modulation Classification," in IEEE Access, vol.8, pp.1498-1506, 2020, doi:10.1109 / ACCESS.2019.2960775. The data amplification technology in the image field is applied to wireless signal modulation recognition. The open radio signal dataset RadioML2016.10a is used, which contains signal samples of 11 different modulation categories. According to the characteristics of the modulation signal, three enhancement methods, rotation, flipping and Gaussian noise, are used for preprocessing to enhance the dataset. The experiment is based on the LSTM network architecture, and the network structure includes two LSTM layers and one fully connected layer for classification. The experimental results show that the three enhancement methods can improve the classification accuracy. Data enhancement also allows the successful classification of radio signals with fewer sampling points, thereby simplifying the deep learning model and shortening the classification response time.

[0006] Later, when the GAN network achieved great success in the field of image generation, many scholars converted the modulated signal into a constellation diagram and directly generated the constellation diagram for modulation recognition like image generation. However, this is not essential. Figure 3Flowchart for small sample modulation recognition using GAN, Tang et al. published Z. Tang, M. Tao, J. Su, Y. Gong, Y. Fan and T. Li, "Data Augmentation for Signal Modulation Classification using Generative Adversity Network," 2021IEEE 4th International Conference on Electronic Information and Communication Technology (ICEICT), Xi'an, China, 2021, pp. 450-453, doi: 10.1109 / ICEICT53123.2021.9531296. A signal modulation classification data enhancement method based on a generative inverse network is used. The signal does not need to be converted, but is directly processed, taking into account various signal-to-noise ratio situations. The open radio signal dataset RadioML2016.10a is used, which contains signal samples of 11 different modulation categories. The original small sample data is augmented using a generative inverse network, and then classified and recognized using a CNN network.

[0007] However, rotation can cause some modulation types to be very sensitive to phase changes. The rotation operation may change the phase information of the signal, thus affecting the classification results. The rotation operation will generate multiple similar samples, which may increase the redundant information in the data set, thereby increasing the computational overhead and training time. Flipping, such as horizontal flipping and vertical flipping, may change the physical meaning of the signal, causing the model to learn incorrect features.

[0008] The training process of the GAN network is unstable and prone to problems such as mode collapse and gradient vanishing, which leads to a lack of diversity in the generated data and requires a lot of computing resources and time. Especially when the network structure of the generator and discriminator is complex, the generation process is complicated and it is difficult to explain the internal mechanism of the generated data, which is a disadvantage for some application scenarios that require high interpretability.

[0009] When dealing with time series, classifiers such as CNN networks are weak at capturing long-distance dependencies, and usually require fixed-size inputs, which are not flexible enough for processing variable-length sequences. When dealing with LSTM networks, due to the dependencies of sequence data, LSTM is difficult to parallelize, and the training and inference speeds are slow. Although LSTM alleviates the gradient vanishing problem to a certain extent, it may still encounter this problem when dealing with very long sequences. Summary of the invention

[0010] In order to solve the technical problem of low recognition rate of modulation recognition under small sample conditions due to insufficient data set in Internet access and related services, the present invention provides a small sample modulation recognition method based on diffusion model and attention mechanism. The method uses DDIM guided by classifier to generate high-quality samples under given conditions, which helps to solve the small sample problem. By generating more samples to enhance the training data, the generation process is relatively stable and is not prone to mode collapse. The step-by-step denoising generation process disclosed in the present invention can better control the quality of generated samples and make the generated samples closer to real data. This step-by-step generation process provides a clear framework, which can intuitively understand and analyze the changes in each step of generation. This transparency helps to explain the generation mechanism of the model and the generation process of the data.

[0011] The attention mechanism was originally proposed in natural language processing to solve the problem of locating key features in long sequence information. Its core idea is to assign different "importances" to different positions in the input features, that is, different frequency bands, etc. through learnable weights, so as to achieve the focus on key areas or channels and suppress irrelevant or redundant information. In the present invention, the multi-scale Fourier decomposition module has mapped the signal to the frequency domain based on the Fourier basis function, explicitly capturing its periodicity and phase information. However, under different channel conditions, modulation formats and noise environments, not all frequency components are equally important for classification. Through the attention layer, the frequency bands with higher significance or discrimination are weighted, and then the visual "attention weights" are output based on the attention mechanism to indicate which frequency bands contribute more to the final decision, improve the interpretability of the model, and finally, under Doppler conditions in complex environments and multipath channel interference, by highlighting key information points, the classification accuracy of high-order modulated signals is improved.

[0012] The purpose of the present invention is specifically achieved through the following technical solutions:

[0013] The present invention discloses a small sample modulation recognition method based on a diffusion model and an attention mechanism, the method comprising:

[0014] Step 1: After converting the time domain data in the data set into positive frequency data using the FFT algorithm, the data set is divided into a training set, a test set, and a validation set by modulation type and preset ratio; and a training set under a small sample condition is divided out from the training set as the original data;

[0015] Step 2: Input the original data into the diffusion model DDIM, and use the forward process of DDIM to gradually add noise to the original data through time steps to obtain pure Gaussian noise samples;

[0016] Step 3, the backward process of DDIM uses the inverse denoising process to gradually denoise the pure Gaussian noise samples to generate recovered data;

[0017] Step 4: Train the classifier using pure Gaussian noise samples and the target modulation type corresponding to the task, calculate the log probability gradient of the classifier, and adjust the generation trajectory of the stepwise denoising of the pure Gaussian noise samples based on the log probability gradient and the mean and covariance predicted by DDIM for the inverse denoising process without conditions, to obtain synthetic data that meets the target modulation type and the trained DDIM;

[0018] Step 5: The classifier model FATT is constructed by sequentially connecting a two-dimensional convolutional layer, a one-dimensional convolutional layer, a residual convolution block, a multi-scale Fourier decomposition module, and an attention layer;

[0019] The original data, restored data and synthesized data are taken as input data and input into FATT. After passing through the two-dimensional convolution layer and the one-dimensional convolution layer, they enter the residual convolution block. After adding the input data of the residual convolution block and the output data of the second convolution layer of the residual convolution block, they enter the multi-scale Fourier decomposition module. The input data is mapped to the frequency domain through the Fourier basis function, and the periodicity and non-periodicity characteristics of the input data are explicitly captured, thereby extracting the comprehensive features including frequency features and nonlinear features, and splitting the comprehensive features into frequency features and nonlinear features.

[0020] Through the attention layer, the attention weight is calculated for each frequency band in the frequency feature and weighted to obtain the weighted frequency feature. The weighted frequency feature is spliced ​​with the nonlinear feature to obtain the spliced ​​feature, and the spliced ​​feature is spliced ​​with the output of the residual convolution block again to obtain the predicted modulation type vector; the predicted modulation type result of the predicted modulation type vector on the target modulation type is visualized through the confusion matrix to obtain the trained FATT;

[0021] Step 6: Input the test set and validation set into the trained DDIM and FATT for testing and validation, and output the small sample prediction modulation type results.

[0022] In step 1, after converting the time domain data in the data set into positive frequency data using the FFT algorithm, the data set is divided into a training set, a test set, and a validation set by modulation type and preset ratio; and the method of dividing the training set under the small sample condition as the original data in the training set is:

[0023] The 128×2 time domain data in the data set is converted into 256×1 frequency data using the FFT algorithm, and the frequency data is processed with absolute values ​​to retain the positive frequency data;

[0024] The data set is divided into 80% training set and 20% test set and validation set according to the modulation type and preset ratio; among them, the training set under the small sample condition accounts for 10% of the training set and is used as the original data.

[0025] In step 2, the original data is input into the diffusion model DDIM, and the forward process of DDIM is used to gradually add noise to the original data through time steps to obtain a pure Gaussian noise sample. The method includes:

[0026]

[0027] Where, X T is a pure Gaussian noise sample, T is the time step, X0 is the original data, ε is Gaussian noise, and it obeys the standard Gaussian distribution (0, I); is the multiplication of the initial value of the noise attenuation coefficient, is the product of the noise attenuation coefficient, α t is the noise attenuation coefficient, and t is the time.

[0028] In step 3, the backward process of DDIM uses an inverse denoising process to gradually denoise the pure Gaussian noise samples, and the method for generating restored data includes:

[0029]

[0030] Where, X T-1 is a pure Gaussian noise sample X T The restored data generated by stepwise denoising, α t-1 is the noise attenuation coefficient α t The previous value of Θ (X T ,T) is the noise predicted by the diffusion model DDIM, is the variance.

[0031] In step 4, the classifier is trained by pure Gaussian noise samples and the target modulation type corresponding to the task, the log probability gradient of the classifier is calculated, and the generation trajectory of the stepwise denoising of the pure Gaussian noise samples is adjusted based on the log probability gradient and the mean and covariance predicted by DDIM for the inverse denoising process without conditions. The method to obtain synthetic data that meets the target modulation type and the trained DDIM is as follows:

[0032]

[0033] In the formula, is the mean value of the unconditional prediction of the inverse denoising process by the trained DDIM, X′ T is the synthetic data, y is the target modulation type; μ θ (X T,T) is the mean value of DDIM’s prediction of the inverse denoising process without any conditions;∑ θ (X T ,T) is the covariance of DDIM's prediction of the inverse denoising process without conditions; s is the guidance scale factor, which is used to control the influence of the classifier's guidance on the generated trajectory; when s=0, it is unconditional generation; when s>0, the generated synthetic data is pushed towards the target modulation type; is the log probability gradient of the classifier, P φ (y|X T ) is a classifier.

[0034] In step 5, the Fourier basis function is:

[0035] Φ(x)=[cos(W P x)||sin(W P x)||σ(B P +W P x)]

[0036] In the formula, Φ(x) is the extracted comprehensive feature including frequency feature and nonlinear feature, cos(W P x)||sin(W P x) is used to capture the periodic characteristics of the input data x, where cos(W P x) and sin(W P x) captures the cosine and sine components of the input data respectively, directly embedding the periodicity in the Fourier series, W P is the projection matrix, which is used to adjust the frequency characteristics of the input data; σ(B P +W P x) is a nonlinear activation function used to capture the non-periodic characteristics of the input data, B P is the bias term, σ represents the activation function, and || represents the concatenation operation, which concatenates the sine component, cosine component and non-periodic features together.

[0037] In step 5, the method of splitting the comprehensive features into frequency features and nonlinear features is:

[0038] F frep +F non =FanLayer(Φ(x))

[0039] In the formula, F frep is the frequency characteristic; F non is a nonlinear feature; FanLayer(Φ(x)) is a comprehensive feature through a multi-scale Fourier decomposition module.

[0040] In step 5, the attention weight is calculated for each frequency band in the frequency feature and weighted. The method for obtaining the weighted frequency feature is:

[0041]

[0042] In the formula, is the frequency characteristic of the i-th frequency band after weighting, α i is the attention weight corresponding to the i-th frequency band in the frequency feature; is the frequency characteristic of the i-th frequency band, where i represents the number of the frequency band, ω i is the attention score parameter corresponding to the i-th frequency band; using ω i Get the attention score corresponding to each frequency band; N is the total number of frequency features, j is the number of frequency features; ω j is the attention score parameter corresponding to the j-th frequency feature, is the jth frequency feature in each frequency band.

[0043] In step 5, the method of concatenating the weighted frequency feature and the nonlinear feature to obtain the concatenated feature is:

[0044]

[0045] In the formula, F out is the splicing feature; is the weighted frequency feature; F non It is a non-linear feature.

[0046] In step 5, the confusion matrix is:

[0047]

[0048] In the formula, C ab is the comparison between the true category and the predicted category in the confusion matrix, that is, the predicted modulation type vector, which is used to visualize the predicted modulation type result of the predicted modulation type vector on the target modulation type; where a,b∈{1,2,…,k}, k is the number of categories, that is, the number of samples with the true category a predicted as category b; N is the total number of samples, n is the sample number, 1(.) is the indicator function, which takes the value of 1 when the condition in the brackets is met, otherwise it is 0; y n is the true category of each sample; is the predicted category.

[0049] The beneficial effects of the present invention are:

[0050] 1. The technical solution disclosed in the present invention solves the technical problem of low modulation recognition rate under small sample conditions due to insufficient data set in Internet access and related services.

[0051] 2. Use the FFT algorithm to convert the time domain data in the data set into positive frequency data to obtain a data vector of 256x1 dimension, and perform absolute value processing on the frequency to retain the positive frequency. This preprocessing method will not destroy the original data characteristics and has stronger expressiveness than the time domain representation;

[0052] 3. Using the diffusion model DDIM as a generative model, the forward process is transparent, which helps to explain the model generation mechanism and data generation process. The backward process uses the reverse denoising process to gradually denoise the pure Gaussian noise samples. The method of generating restored data accelerates the sampling generation and the generation quality is not much different from DDPM. The DDIM reverse process is more efficient and only requires fewer steps to achieve denoising and data generation. The generation process is stable and the generated data is of high quality.

[0053] Compared with DDPM, DDIM reduces the randomness in the generation process and improves the generation efficiency;

[0054] Compared with GAN, DDIM avoids the problem of mode collapse that may occur during training and generates better diversity of data;

[0055] Compared with flipping and rotation, DDIM improves the diversity of generated data and can simulate complex distribution of data;

[0056] The diffusion model DDIM enhances the training data by generating more samples. The generation process is relatively stable and not prone to mode collapse. The step-by-step denoising generation process can better control the quality of the generated samples and make the generated samples closer to the real data. This method uses the classifier-guided DDIM to generate high-quality samples under given conditions, which helps solve the problem of insufficient small sample data sets.

[0057] 4. Train the classifier by using pure Gaussian noise samples and the target modulation type corresponding to the task, introduce the classifier to guide the generation process, provide guidance information of the target modulation type, and adjust the generation process towards the target category;

[0058] Calculate the logarithmic probability gradient of the classifier to guide the reverse denoising process so that the generated synthetic data conforms to the target category and the generated data is closer to the real data in distribution;

[0059] Classifier guidance provides guidance information of target conditions during the generation process, so that the generation process can be adjusted towards the target category more accurately. Compared with ordinary condition control, classifier guidance can dynamically adjust the generation process, improve the quality of generated data and category matching, and provide logarithmic probability gradients to directly guide the reverse denoising process, so that the generated data is more consistent with the target category and improve the accuracy of generation;

[0060] 5. The constructed classifier model FATT converts the input data into frequency domain through a multi-scale Fourier decomposition module, explicitly captures the periodic and non-periodic features of the input data, and thus extracts comprehensive features including frequency features and nonlinear features, and splits the comprehensive features into frequency features and nonlinear features; then, the attention mechanism embedded in the attention layer automatically calculates the attention weights for each frequency band feature and weights them, concatenates the weighted frequency features with the nonlinear features to obtain the concatenated features, and concatenates the concatenated features with the output of the residual convolution block again to obtain the predicted modulation type vector; it effectively focuses on the most discriminative frequency components, thereby improving the signal discrimination ability and modulation recognition accuracy under complex channel conditions.

[0061] 6. This method not only alleviates the underfitting or overfitting phenomenon that is prone to occur in deep learning models under small sample conditions, but also overcomes the phase interference and redundant information problems that may be caused by traditional rotation and flipping data enhancement methods, while avoiding the pattern collapse and gradient disappearance problems that exist in the GAN generation process.

[0062] 7. The present invention provides a highly accurate, robust and well-interpretable solution for wireless signal modulation recognition in Internet access and related services in the case of insufficient data through a stable and efficient generation process and significantly enhanced frequency domain feature extraction capabilities. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] The present invention is further described in detail below based on the accompanying drawings and embodiments.

[0064] Figure 1 It is a schematic diagram of a small sample modulation recognition process block diagram provided by the present invention.

[0065] Figure 2 It is a schematic diagram of a rough flow chart of small sample modulation recognition provided by the present invention.

[0066] Figure 3 It is a schematic diagram of the flow chart of small sample modulation recognition performed by GAN provided by the present invention.

[0067] Figure 4 It is a schematic diagram of the noise adding process of the diffusion model provided by the present invention.

[0068] Figure 5 It is a schematic diagram of the diffusion model denoising process provided by the present invention.

[0069] Figure 6 It is a schematic diagram of the overall flow chart of the small sample modulation recognition model using the diffusion model and attention mechanism provided by the present invention.

[0070] Figure 7 It is a schematic diagram of a flow chart of a modulation recognition model provided by the present invention, which is input into a classifier model after being enhanced by diffusion model data.

[0071] Figure 8 It is a schematic diagram of a signal diagram generated by 8PSK at 16db provided by the present invention.

[0072] Fig. 9 It is a schematic diagram of a signal diagram generated by AM-DSB at 16db provided by the present invention.

[0073] Fig.10 It is a schematic diagram of a BPSK generated signal diagram under 16db provided by the present invention.

[0074] Fig.11 It is a schematic diagram of a CPFSK signal generation diagram under 16db provided by the present invention.

[0075] Fig.12 It is a schematic diagram of a QPSK generated signal diagram under 16db provided by the present invention.

[0076] Fig.13 It is a schematic diagram of a GFSK generated signal diagram under 16db provided by the present invention.

[0077] Fig.14 This is a schematic diagram of a PAM4 generated signal diagram under 16db provided by the present invention.

[0078] Fig.15 It is a schematic diagram of a QAM16 generated signal diagram under 16db provided by the present invention.

[0079] Fig.16 It is a schematic diagram of a QAM64 generated signal diagram under 16db provided by the present invention.

[0080] Fig.17 It is a schematic diagram of a WBFM signal generated at 16db provided by the present invention.

[0081] Fig.18 It is a schematic diagram of the classification recognition rate of the validation set and test set modulated signals under the deep learning model after data enhancement provided by the present invention.

[0082] Fig.19 It is a schematic diagram of the confusion matrix of the validation set and the test set at 10db after data enhancement provided by the present invention. DETAILED DESCRIPTION

[0083] The embodiment of the present invention discloses a small sample modulation recognition method based on a diffusion model and an attention mechanism, the method comprising:

[0084] Step 1: After converting the time domain data in the data set into positive frequency data using the FFT algorithm, the data set is divided into a training set, a test set, and a validation set by modulation type and preset ratio; and a training set under a small sample condition is divided out from the training set as the original data;

[0085] For example, the dataset is RML2016.10b, which includes 120,0000 samples of modulated radio signals, evenly distributed in 10 different modulation categories. Each modulation category in the dataset is represented by 20 signal-to-noise ratio (SNR) levels, ranging from -20dB to 18dB, with an increment of 2dB.

[0086] Step 2: Input the original data into the diffusion model DDIM, and use the forward process of DDIM to gradually add noise to the original data through time steps to obtain pure Gaussian noise samples;

[0087] Step three, the backward process of DDIM uses the inverse denoising process to gradually denoise the pure Gaussian noise samples and generate restored data; the backward process of DDIM starts from the pure Gaussian noise samples and gradually predicts and removes noise, so that the noise data is gradually restored to samples that conform to the real data distribution. This generation process can produce high-quality and realistic data.

[0088] Step 4: Train the classifier using pure Gaussian noise samples and the target modulation type corresponding to the task, calculate the log probability gradient of the classifier, and adjust the generation trajectory of the stepwise denoising of the pure Gaussian noise samples based on the log probability gradient and the mean and covariance predicted by DDIM for the inverse denoising process without conditions, to obtain synthetic data that meets the target modulation type and the trained DDIM;

[0089] This step needs to target samples with different noise levels: pure Gaussian noise sample X T Train a classifier P with the target modulation type corresponding to the task φ (y|X T ). The classifier learns how to judge whether the sample meets the target modulation type on the noisy data. In the reverse denoising process, when generating samples at each step, it is hoped that the generated samples will be more inclined to the target condition. To this end, it is necessary to calculate the current sample X T Log-probability gradient conditioned on the target: Instructions on how to adjust X T In order to make the classifier think that it is more consistent with the target modulation type y, the generation trajectory of the stepwise denoising of pure Gaussian noise samples is adjusted based on the logarithmic probability gradient and the mean and covariance predicted by DDIM in the unconditional reverse denoising process. In each step of the reverse denoising process, the gradient signal of the logarithmic probability gradient from the classifier is used, so that the final generated synthetic data is more consistent with the target modulation type.

[0090] Step 5: The classifier model FATT is constructed by sequentially connecting a two-dimensional convolutional layer, a one-dimensional convolutional layer, a residual convolution block, a multi-scale Fourier decomposition module, and an attention layer;

[0091] The original data, restored data and synthesized data are taken as input data and input into FATT. After passing through the two-dimensional convolution layer and the one-dimensional convolution layer, they enter the residual convolution block. After adding the input data of the residual convolution block and the output data of the second convolution layer of the residual convolution block, they enter the multi-scale Fourier decomposition module. The input data is mapped to the frequency domain through the Fourier basis function, and the periodicity and non-periodicity characteristics of the input data are explicitly captured, thereby extracting the comprehensive features including frequency features and nonlinear features, and splitting the comprehensive features into frequency features and nonlinear features.

[0092] Through the attention layer, the attention weight is calculated for each frequency band in the frequency feature and weighted to obtain the weighted frequency feature. The weighted frequency feature is spliced ​​with the nonlinear feature to obtain the spliced ​​feature, and the spliced ​​feature is spliced ​​with the output of the residual convolution block again to obtain the predicted modulation type vector; the predicted modulation type result of the predicted modulation type vector on the target modulation type is visualized through the confusion matrix to obtain the trained FATT;

[0093] Step 6: Input the test set and validation set into the trained DDIM and FATT for testing and validation, and output the small sample prediction modulation type results.

[0094] In step 1, after converting the time domain data in the data set into positive frequency data using the FFT algorithm, the data set is divided into a training set, a test set, and a validation set by modulation type and preset ratio; and the method of dividing the training set under the small sample condition as the original data in the training set is:

[0095] The 128×2 time domain data in the data set is converted into 256×1 frequency data using the FFT algorithm, and the frequency data is processed with absolute values ​​to retain the positive frequency data;

[0096] The data set is divided into 80% training set and 20% test set and validation set according to the modulation type and preset ratio; among them, the training set under the small sample condition accounts for 10% of the training set and is used as the original data.

[0097] like Figure 4 As shown, in step 2, the original data is input into the diffusion model DDIM, and the forward process of DDIM is used to gradually add noise to the original data through time steps to obtain a pure Gaussian noise sample. The method includes:

[0098]

[0099] Where, X Tis a pure Gaussian noise sample, T is the time step, for example, T = 1000, X0 is the original data, ε is Gaussian noise, and it obeys the standard Gaussian distribution (0, I); is the multiplication of the initial value of the noise attenuation coefficient, is the product of the noise attenuation coefficient, α t is the noise attenuation coefficient, similar to a hyperparameter, which only needs to sample the noise once and can be obtained directly from X0 T , t is the time.

[0100] like Figure 5 As shown, in step 3, the backward process of DDIM uses a reverse denoising process to gradually denoise the pure Gaussian noise samples, and the method for generating restored data includes:

[0101]

[0102] Where, X T-1 is a pure Gaussian noise sample X T The restoration result of step-by-step denoising, after step-by-step denoising for time step T, the final restored data generated is the original data X0, α t-1 is the noise attenuation coefficient α t The previous value of Θ (X T ,T) is the noise predicted by the diffusion model DDIM, is the variance.

[0103] In step 4, the classifier is trained by pure Gaussian noise samples and the target modulation type corresponding to the task, the log probability gradient of the classifier is calculated, and the generation trajectory of the stepwise denoising of the pure Gaussian noise samples is adjusted based on the log probability gradient and the mean and covariance predicted by DDIM for the inverse denoising process without conditions. The method to obtain synthetic data that meets the target modulation type and the trained DDIM is as follows:

[0104]

[0105] In the formula, is the mean value of the unconditional prediction of the inverse denoising process by the trained DDIM, X′ T is the synthetic data, y is the target modulation type; μ θ (X T ,T) is the mean value of DDIM’s prediction of the inverse denoising process without any conditions;∑ θ (X T ,T) is the covariance of DDIM's prediction of the inverse denoising process without conditions; s is the guidance scale factor, which is used to control the influence of the classifier's guidance on the generated trajectory; when s=0, it is unconditional generation; when s>0, the generated synthetic data is pushed towards the target modulation type; is the log probability gradient of the classifier, P φ (y|X T ) is a classifier.

[0106] In step 5, the Fourier basis function is:

[0107] Φ(x)=[cos(W P x)||sin(W P x)||σ(B P +W P x)]

[0108] In the formula, Φ(x) is the extracted comprehensive feature including frequency feature and nonlinear feature, cos(W P x)||sin(W P x) is used to capture the periodic characteristics of the input data x, where cos(W P x) and sin(W P x) captures the cosine and sine components of the input data respectively, directly embedding the periodicity in the Fourier series, W P is the projection matrix, which is used to adjust the frequency characteristics of the input data by flattening the input data and then comparing it with the fixed projection matrix W P Multiply to calculate the cosine and sine components; σ(B P +W P x) is a nonlinear activation function used to capture the non-periodic characteristics of the input data, B P is the bias term, σ represents the activation function, and || represents the concatenation operation, which concatenates the sine component, cosine component and non-periodic features together.

[0109] In step 5, the method of splitting the comprehensive features into frequency features and nonlinear features is:

[0110] F frep +F non =FanLayer(Φ(x))

[0111] In the formula, F frep is the frequency characteristic; F non is a nonlinear feature; FanLayer(Φ(x)) is a comprehensive feature through a multi-scale Fourier decomposition module.

[0112] In step 5, the attention weight is calculated for each frequency band in the frequency feature and weighted to obtain the weighted frequency feature. The purpose is to give different "importances" to different positions (different frequency bands, etc.) in the input feature through learnable weights, so as to focus on key areas or channels and suppress irrelevant or redundant information. The method is:

[0113]

[0114]

[0115] In the formula, is the frequency characteristic of the i-th frequency band after weighting, α i is the attention weight corresponding to the i-th frequency band in the frequency feature, and its value reflects the importance of the frequency band; is the frequency characteristic of the i-th frequency band, where i represents the number of the frequency band, ω i is the attention score parameter corresponding to the i-th frequency band; using ω i Get the attention score corresponding to each frequency band; N is the total number of frequency features, j is the number of frequency features; ω j is the attention score parameter corresponding to the j-th frequency feature, is the jth frequency feature in each frequency band.

[0116] In step 5, the method of concatenating the weighted frequency feature and the nonlinear feature to obtain the concatenated feature is:

[0117]

[0118] In the formula, F out is the splicing feature; is the weighted frequency feature; F non It is a nonlinear feature. It overcomes the shortcomings of the existing network model in extracting time-frequency features.

[0119] In step 5, the confusion matrix is ​​a visualization tool used to evaluate the performance of the classification model. It displays the prediction results of the model in each category in a table, thereby helping to analyze the correct classification and misclassification of the model in different categories. The specific formula is:

[0120]

[0121] In the formula, C ab is the comparison between the true category and the predicted category in the confusion matrix, that is, the predicted modulation type vector, which is used to visualize the predicted modulation type result of the predicted modulation type vector on the target modulation type; where a,b∈{1,2,…,k}, k is the number of categories, that is, the number of samples with the true category a predicted as category b; N is the total number of samples, n is the sample number, 1(.) is the indicator function, which takes the value of 1 when the condition in the brackets is met, otherwise it is 0; y n is the true category of each sample; is the predicted category.

[0122] like Figure 6 to Figure 7As shown in the figure, the classifier model FATT is constructed by sequentially connecting a two-dimensional convolution layer, a one-dimensional convolution layer, a residual convolution block, a multi-scale Fourier decomposition module and an attention layer; the original data, the restored data and the synthetic data are used as input data, and after preprocessing such as data enhancement, slicing and linear embedding, the input format required by the FATT network is obtained; the input signal dimension of the FATT network is [batch_size, 2, 1024]. First, the input data passes through the two-dimensional convolution layer Conv1, and the convolution kernel size is (2, 7). Features are extracted in the local area of ​​the time series, and the two input channels are mapped to 1024 feature channels. Its output dimension becomes [batch_size, 1024, 1, 64]. Then, by removing the third dimension (size is 1), the output dimension is obtained to be [batch_size, 1024, 64]. Batch normalization BatchNorm2d and LeakyReLU activation functions are added after the convolution layer to enhance nonlinear expression ability and numerical stability. In the residual convolution block, the input signal [batch_size, 1024, 64] passes through the first layer of Conv1D, with 64 filters, 3 convolution kernels, no bias, and L2 regularization. The convolution is followed by a BatchNormalization layer and a LeakyReLU activation function, followed by a Dropout layer that randomly discards some neurons at a ratio of 0.3 to prevent overfitting. It passes through the Conv1D layer again, with 64 filters, 3 convolution kernels, no bias, and L2 regularization. The input is added to the output of the second layer of convolution (skip connection), and the LeakyReLU activation function is used, and the output dimension is kept as [batch_size, 1024, 64]. After passing through the residual convolution block, the signal enters the multi-scale Fourier decomposition module. The input signal [batch_size, 1024, T] first accepts the input feature [batch_size, 1024, 64], and generates a comprehensive feature containing frequency features and nonlinear features through a random projection matrix and a trainable linear mapping. The specific operation includes flattening the input and multiplying it with a fixed projection matrix, calculating the cosine component and the sine component, and then generating nonlinear features through a trainable linear mapping. The output shape of FanLayer is [batch_size, 1024, 128], where 128 = 2*proj_dim(32)+fan_out(64).The output of FanLayer is split into frequency features [batch_size, 1024, 64] and nonlinear features [batch_size, 1024, 64]. The frequency features are weighted by the FrequencyAttention layer to generate the attention-enhanced frequency features [batch_size, 1024, 64]. The weighted frequency features are concatenated with the nonlinear features to obtain [batch_size, 1024, 128]. Subsequently, this concatenated feature is concatenated again with the output of the residual convolution block [batch_size, 1024, 64] to form a comprehensive feature of [batch_size, 1024, 192]. The attention gate is generated by 1x1 convolution (Conv1D), and the shape is [batch_size, 1024, 64]. Then, the features are fused through element-wise multiplication and addition operations to obtain the fused features [batch_size, 1024, 64]. The fused features are globally average pooled to obtain a feature vector with a shape of [batch_size, 64]. The Dense layer is used to expand the feature dimension from 64 to 256, and the activation function is LeakyReLU. Then, the Dropout layer is used to randomly discard some neurons at a ratio of 0.5. The Dense layer is used to map the features to the target classification dimension num_classes, the activation function is softmax, and the output dimension is [batch_size, num_classes].

[0123] The confusion matrix is ​​a visualization tool used to evaluate the performance of classification models. It shows the prediction results of the model in each category in a table, which helps analyze the correct classification and misclassification of FATT in different categories. The confusion matrix is ​​used to show the comparison between the true category and the predicted category, and the classification results of the FATT model in each category are displayed to obtain the trained FATT.

[0124] In order to explain the technical solution of the present invention in detail, specific examples are provided for illustration:

[0125] Example 1: Preprocessing of the dataset

[0126] 1. Use the RML2016.10a dataset, which contains 220,000 modulated radio signal samples, distributed in 11 different modulation categories, each modulation category has 20 signal-to-noise ratio (SNR) levels, ranging from -20dB to 18dB, with an increment of 2dB, and the dataset is a time series dimension of 128x2. Or use the RML2016.10b dataset, which includes 120,0000 modulated radio signal samples, evenly distributed in 10 different modulation categories. Each modulation category in the dataset is represented by 20 signal-to-noise ratio (SNR) levels, ranging from -20dB to 18dB, with an increment of 2dB.

[0127] 2. Use the FFT algorithm to convert the data in the data set from the time domain to the frequency domain to obtain data of 256x1 dimension. Take the absolute value of the frequency domain data and only retain the positive frequency part.

[0128] 3. Divide the data set into a training set under small sample conditions, which accounts for 10% of the training set, and a test set and a validation set, which account for a total of 20%.

[0129] Example 2 Generator training and results

[0130] 1. During the DDIM model training process, classifier guidance is introduced to provide guidance information of the target conditions. The classifier guides the reverse denoising process through gradient information so that the generated data conforms to the target category. This classifier has been trained with the complete data set.

[0131] 2. The DDIM forward process performs noise training on a signal with a total number of steps T = 1000. After adding Gaussian white noise, it becomes pure Gaussian noise, and ε obeys the standard Gaussian distribution (0, I). Through the linear Gaussian properties and Markov chain properties, we can get X T The expression formula of .

[0132] 3. DDIM reverse process Unet network gradually denoises the original data from pure noise, and the denoising model learned predicts the noise. Using the noise attenuation coefficient, after T steps of denoising, the final generated data X0 is obtained. Among them, Batchsize is set to 512, the learning rate is 0.0001, and the following can be obtained by running on the Pycharm platform: Figure 8 The IQ signal result diagram of the 8PSK generated signal under 16db is shown, where the ordinate is the amplitude of the IQ signal and the abscissa is the time step of 128; Fig. 9 The IQ signal result diagram of the AM-DSB generated signal under 16db is shown, where the ordinate is the amplitude of the IQ signal and the abscissa is the time step of 128; Fig.10 The IQ signal result diagram of the BPSK generated signal under 16db is shown, where the ordinate is the amplitude of the IQ signal and the abscissa is the time step of 128; Fig.11The IQ signal result diagram of the CPFSK generated signal under 16db is shown, where the ordinate is the amplitude of the IQ signal and the abscissa is the time step of 128; Fig.12 The IQ signal result diagram of the QPSK generated signal under 16db is shown, where the ordinate is the amplitude of the IQ signal and the abscissa is the time step of 128; Fig.13 The IQ signal result diagram of the GFSK generated signal under 16db is shown, where the ordinate is the amplitude of the IQ signal and the abscissa is the time step of 128; Fig.14 The IQ signal result diagram of the PAM4 generated signal at 16db is shown, where the ordinate is the amplitude of the IQ signal and the abscissa is the time step of 128; Fig.15 The IQ signal result diagram of the QAM16 generated signal under 16db is shown, where the ordinate is the amplitude of the IQ signal and the abscissa is the time step of 128; Fig.16 The IQ signal result diagram of the QAM64 generated signal under 16db is shown, where the ordinate is the amplitude of the IQ signal and the abscissa is the time step of 128; Fig.17 The IQ signal result diagram of the WBFM generated signal at 16db is shown, where the ordinate is the amplitude of the IQ signal and the abscissa is the time step of 128.

[0133] Example 3 Classifier training and results

[0134] 1. Preprocess the original data, restored data and synthesized data and convert them into the input format suitable for FATT (256x1).

[0135] 2. The input data is 256x1. After preprocessing steps such as data enhancement, segmentation, and encoding, the Fourier analysis network of the multi-scale Fourier decomposition module explicitly converts the input data into the frequency domain to extract comprehensive features including frequency features and nonlinear features. Subsequently, the attention mechanism embedded in the attention layer automatically calculates the importance weights for the features of each frequency band, effectively focusing on the frequency components with the most discriminative ability.

[0136] 3. Use the confusion matrix to show the comparison between the true category and the predicted category, and show the performance of FATT in each category. Fig.18 As shown in the figure, when the signal-to-noise ratio is 16db, the recognition rate of the Transformer classifier model without data enhancement is lower than that of the classifier model after data enhancement, while the classifier model FATT proposed in the present invention is better than other classifier models after data enhancement of the diffusion model. Fig.19As shown, in the confusion matrix under 10db, except for the WBFM modulation type confusion, the recognition results of other modulation type signals are all above 90%, indicating that the diffusion model has a very poor generation effect on the WBFM modulation type signal, resulting in the inability to generate signals similar to the original data set, resulting in poor training of the classifier model, leading to WBFM modulation type signal confusion.

[0137] In summary, the experimental results in the example show that the small sample modulation recognition method based on the diffusion model and attention mechanism proposed in the present invention effectively solves the problem of difficulty in wireless signal modulation recognition caused by insufficient data sets in the field of modulated carrier systems.

[0138] In this method, by introducing the classifier-guided DDIM reverse denoising generation process, the generated samples are gradually adjusted toward the target modulation type during the noise removal process. The gradient information calculated by the pre-trained classifier is used to make the generated data closer to the real data in distribution. At the same time, the time domain signal is converted into a 256×1-dimensional frequency domain representation through FFT algorithm transformation, and the absolute value of the representation is taken to retain the positive frequency information. The importance of each frequency band in the frequency domain is adaptively weighted in combination with the embedded attention mechanism, thereby significantly capturing key feature information including amplitude, phase and periodicity, and improving the signal discrimination ability under complex channel conditions.

[0139] Experimental data show that under 16dB signal-to-noise ratio conditions, after diffusion model data enhancement, the modulation recognition accuracy of the classifier model is significantly better than that of the unenhanced data. For all modulation signals except the WBFM modulation type, the accuracy rate is over 90%. The confusion matrix under 10dB conditions further proves the advantage of this method in low signal-to-noise ratio environments.

[0140] Overall, this method not only alleviates the underfitting or overfitting phenomenon that deep learning models are prone to under small sample conditions, but also overcomes the phase interference and redundant information problems that may be caused by traditional rotation and flipping data enhancement methods, while avoiding the pattern collapse and gradient vanishing problems that exist in the GAN generation process.

[0141] In summary, the present invention provides a solution with high accuracy, high robustness and good interpretability for wireless signal modulation recognition in the case of insufficient data through a stable and efficient generation process and significantly enhanced frequency domain feature extraction capabilities.

[0142] The beneficial effects of the embodiments of the present invention are:

[0143] 1. The technical solution disclosed in the present invention solves the technical problem of low modulation recognition rate under small sample conditions due to insufficient data set in Internet access and related services.

[0144] 2. Use the FFT algorithm to convert the time domain data in the data set into positive frequency data to obtain a data vector of 256x1 dimension, and perform absolute value processing on the frequency to retain the positive frequency. This preprocessing method will not destroy the original data characteristics and has stronger expressiveness than the time domain representation;

[0145] 3. Using the diffusion model DDIM as a generative model, the forward process is transparent, which helps to explain the model generation mechanism and data generation process. The backward process uses the reverse denoising process to gradually denoise the pure Gaussian noise samples. The method of generating restored data accelerates the sampling generation and the generation quality is not much different from DDPM. The DDIM reverse process is more efficient and only requires fewer steps to achieve denoising and data generation. The generation process is stable and the generated data is of high quality.

[0146] Compared with DDPM, DDIM reduces the randomness in the generation process and improves the generation efficiency;

[0147] Compared with GAN, DDIM avoids the problem of mode collapse that may occur during training and generates better diversity of data;

[0148] Compared with flipping and rotation, DDIM improves the diversity of generated data and can simulate complex distribution of data;

[0149] The diffusion model DDIM enhances the training data by generating more samples. The generation process is relatively stable and not prone to mode collapse. The step-by-step denoising generation process can better control the quality of the generated samples and make the generated samples closer to the real data. This method uses the classifier-guided DDIM to generate high-quality samples under given conditions, which helps solve the problem of insufficient small sample data sets.

[0150] 4. Train the classifier by using pure Gaussian noise samples and the target modulation type corresponding to the task, introduce the classifier to guide the generation process, provide guidance information of the target modulation type, and adjust the generation process towards the target category;

[0151] Calculate the logarithmic probability gradient of the classifier to guide the reverse denoising process so that the generated synthetic data conforms to the target category and the generated data is closer to the real data in distribution;

[0152] Classifier guidance provides guidance information of target conditions during the generation process, so that the generation process can be adjusted towards the target category more accurately. Compared with ordinary condition control, classifier guidance can dynamically adjust the generation process, improve the quality of generated data and category matching, and provide logarithmic probability gradients to directly guide the reverse denoising process, so that the generated data is more consistent with the target category and improve the accuracy of generation;

[0153] 5. The constructed classifier model FATT converts the input data into frequency domain through a multi-scale Fourier decomposition module, explicitly captures the periodic and non-periodic features of the input data, and thus extracts comprehensive features including frequency features and nonlinear features, and splits the comprehensive features into frequency features and nonlinear features; then, the attention mechanism embedded in the attention layer automatically calculates the attention weights for each frequency band feature and weights them, concatenates the weighted frequency features with the nonlinear features to obtain the concatenated features, and concatenates the concatenated features with the output of the residual convolution block again to obtain the predicted modulation type vector; it effectively focuses on the most discriminative frequency components, thereby improving the signal discrimination ability and modulation recognition accuracy under complex channel conditions.

[0154] 6. This method not only alleviates the underfitting or overfitting phenomenon that is prone to occur in deep learning models under small sample conditions, but also overcomes the phase interference and redundant information problems that may be caused by traditional rotation and flipping data enhancement methods, while avoiding the pattern collapse and gradient disappearance problems that exist in the GAN generation process.

[0155] 7. The present invention provides a highly accurate, robust and well-interpretable solution for wireless signal modulation recognition in Internet access and related services in the case of insufficient data through a stable and efficient generation process and significantly enhanced frequency domain feature extraction capabilities.

[0156] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.

Claims

1. A small sample modulation recognition method based on diffusion model and attention mechanism, characterized in that: The method includes: Step 1: After converting the time domain data in the data set into positive frequency data using the FFT algorithm, the data set is divided into a training set, a test set, and a validation set by modulation type and preset ratio; and a training set under a small sample condition is divided out from the training set as the original data; Step 2: Input the original data into the diffusion model DDIM, and use the forward process of DDIM to gradually add noise to the original data through time steps to obtain pure Gaussian noise samples; Step 3, the backward process of DDIM uses the inverse denoising process to gradually denoise the pure Gaussian noise samples to generate recovered data; Step 4: Train the classifier using pure Gaussian noise samples and the target modulation type corresponding to the task, calculate the log probability gradient of the classifier, and adjust the generation trajectory of the stepwise denoising of the pure Gaussian noise samples based on the log probability gradient and the mean and covariance predicted by DDIM for the inverse denoising process without conditions, to obtain synthetic data that meets the target modulation type and the trained DDIM; Step 5: The classifier model FATT is constructed by sequentially connecting a two-dimensional convolutional layer, a one-dimensional convolutional layer, a residual convolution block, a multi-scale Fourier decomposition module, and an attention layer; The original data, restored data and synthesized data are taken as input data and input into FATT. After passing through the two-dimensional convolution layer and the one-dimensional convolution layer, they enter the residual convolution block. After adding the input data of the residual convolution block and the output data of the second convolution layer of the residual convolution block, they enter the multi-scale Fourier decomposition module. The input data is mapped to the frequency domain through the Fourier basis function, and the periodicity and non-periodicity characteristics of the input data are explicitly captured, thereby extracting the comprehensive features including frequency features and nonlinear features, and splitting the comprehensive features into frequency features and nonlinear features. Through the attention layer, the attention weight is calculated for each frequency band in the frequency feature and weighted to obtain the weighted frequency feature. The weighted frequency feature is spliced ​​with the nonlinear feature to obtain the spliced ​​feature, and the spliced ​​feature is spliced ​​with the output of the residual convolution block again to obtain the predicted modulation type vector; the predicted modulation type result of the predicted modulation type vector on the target modulation type is visualized through the confusion matrix to obtain the trained FATT; Step 6: Input the test set and validation set into the trained DDIM and FATT for testing and validation, and output the small sample prediction modulation type results.

2. The method according to claim 1, characterized in that In step 1, after converting the time domain data in the data set into positive frequency data using the FFT algorithm, the data set is divided into a training set, a test set, and a validation set by modulation type and preset ratio; and the method of dividing the training set under the small sample condition as the original data in the training set is: The 128×2 time domain data in the data set is converted into 256×1 frequency data using the FFT algorithm, and the frequency data is processed with absolute values ​​to retain the positive frequency data; The data set is divided into 80% training set and 20% test set and validation set according to the modulation type and preset ratio; among them, the training set under the small sample condition accounts for 10% of the training set and is used as the original data.

3. The method according to claim 2, characterized in that In step 2, the original data is input into the diffusion model DDIM, and the forward process of DDIM is used to gradually add noise to the original data through time steps to obtain a pure Gaussian noise sample. The method includes: Where, X T is a pure Gaussian noise sample, T is the time step, X0 is the original data, ε is Gaussian noise, and it obeys the standard Gaussian distribution (0, I); is the multiplication of the initial value of the noise attenuation coefficient, is the product of the noise attenuation coefficient, α t is the noise attenuation coefficient, and t is the time.

4. The method according to claim 3, characterized in that In step 3, the backward process of DDIM uses an inverse denoising process to gradually denoise the pure Gaussian noise samples, and the method for generating restored data includes: Where, X T-1 is a pure Gaussian noise sample X T The restored data generated by stepwise denoising, α t-1 is the noise attenuation coefficient α t The previous value of Θ (X T ,T) is the noise predicted by the diffusion model DDIM, is the variance.

5. The method according to claim 4, characterized in that In step 4, the classifier is trained by pure Gaussian noise samples and the target modulation type corresponding to the task, the log probability gradient of the classifier is calculated, and the generation trajectory of the stepwise denoising of the pure Gaussian noise samples is adjusted based on the log probability gradient and the mean and covariance predicted by DDIM for the inverse denoising process without conditions. The method to obtain synthetic data that meets the target modulation type and the trained DDIM is as follows: In the formula, is the mean value of the unconditional prediction of the inverse denoising process by the trained DDIM, X′ T is the synthetic data, y is the target modulation type; μ θ (X T ,T) is the mean value of DDIM’s prediction of the inverse denoising process without any conditions;∑ θ (X T ,T) is the covariance of DDIM's prediction of the inverse denoising process without conditions; s is the guidance scale factor, which is used to control the influence of the classifier's guidance on the generated trajectory; when s=0, it is unconditional generation; when s>0, the generated synthetic data is pushed towards the target modulation type; is the log probability gradient of the classifier, P φ (y|X T ) is a classifier.

6. The method according to claim 5, characterized in that In step 5, the Fourier basis function is: Φ(x)=[cos(W P x)||sin(W P x)||σ(B P +W P x)] In the formula, Φ(x) is the extracted comprehensive feature including frequency feature and nonlinear feature, cos(W P x)||sin(W P x) is used to capture the periodic characteristics of the input data x, where cos(W P x) and sin(W P x) captures the cosine and sine components of the input data respectively, directly embedding the periodicity in the Fourier series, W P is the projection matrix, which is used to adjust the frequency characteristics of the input data; σ(B P +W P x) is a nonlinear activation function used to capture the non-periodic characteristics of the input data, B P is the bias term, σ represents the activation function, and || represents the concatenation operation, which concatenates the sine component, cosine component and non-periodic features together.

7. The method according to claim 6, characterized in that In step 5, the method of splitting the comprehensive features into frequency features and nonlinear features is: F frep +F non =FanLayer(Φ(x)) In the formula, F frep is the frequency characteristic; F non is a nonlinear feature; FanLayer(Φ(x)) is a comprehensive feature through a multi-scale Fourier decomposition module.

8. The method according to claim 7, characterized in that In step 5, the attention weight is calculated for each frequency band in the frequency feature and weighted. The method for obtaining the weighted frequency feature is: In the formula, is the weighted frequency characteristic of the ith frequency band, α i is the attention weight corresponding to the i-th frequency band in the frequency feature; is the frequency characteristic of the i-th frequency band, where i represents the number of the frequency band, ω i is the attention score parameter corresponding to the i-th frequency band; using ∈ i Get the attention score corresponding to each frequency band; N is the total number of frequency features, j is the number of frequency features; ω j is the attention score parameter corresponding to the j-th frequency feature, is the jth frequency feature in each frequency band.

9. The method according to claim 8, characterized in that In step 5, the method of concatenating the weighted frequency feature and the nonlinear feature to obtain the concatenated feature is: In the formula, F out is the splicing feature; is the weighted frequency feature; F non It is a non-linear feature.

10. The method according to claim 9, characterized in that In step 5, the confusion matrix is: In the formula, C ab is the comparison between the true category and the predicted category in the confusion matrix, that is, the predicted modulation type vector, which is used to visualize the predicted modulation type result of the predicted modulation type vector on the target modulation type; where a,b∈{1,2,…,k}, k is the number of categories, that is, the number of samples with the true category a predicted as category b; N is the total number of samples, n is the sample number, 1(.) is the indicator function, which takes the value of 1 when the condition in the brackets is met, otherwise it is 0; y n is the true category of each sample; is the predicted category.

Citation Information

Patent Citations

  • Automatic identification method for radar signal modulation type

    CN114564982A

  • Small sample communication signal automatic modulation identification method based on incremental learning

    CN114580484A

  • Small sample signal automatic modulation identification method and device

    CN119167146A