A radio signal modulation classification method based on deep ensemble network

By adopting a deep integrated network in radio signal modulation classification, combining attention CNN, LSTM and Transformer models, a multi-level feature extraction module is built, which solves the limitations of a single network architecture and achieves a high accuracy and scalability modulation classification effect.

CN119669868BActive Publication Date: 2025-05-06XIHUA UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510189403.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-05-06
Estimated Expiration
2045-02-20

AI Technical Summary

Technical Problem

In the prior art, a single network architecture is limited to the modeling capabilities of local features, sequential patterns and global dependencies in radio signal modulation classification, and is difficult to achieve satisfactory performance in low signal-to-noise ratio scenarios.

Method used

Using a deep integrated network-based method, combining attention CNN, LSTM and Transformer models, a multi-level feature extraction module is built, including local feature extraction, timing information extraction and global feature extraction, and modulation classification is realized through deep integrated network.

Benefits of technology

Under different signal-to-noise ratio levels and modulation modes, the accuracy and scalability of modulation classification are significantly improved, and a good balance between model parameters and classification accuracy is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119669868B_ABST
    Figure CN119669868B_ABST
Patent Text Reader

Abstract

The present application provides a radio signal modulation classification method based on a deep integrated network, including: S1: constructing an attention CNN module, inputting IQ signal samples into the attention CNN module and converting them into multi-channel feature maps; S2: constructing a timing information extraction module, using the timing information extraction module to extract sequential patterns from local features; S3: constructing a visual transformer module, inputting the output of the timing information extraction module into a global feature extractor module to extract global features to obtain a feature vector representing the modulation category; S4: constructing a classifier, mapping the feature vector representing the modulation category from a high-dimensional space to a low-dimensional space representing the signal type, and then using softmax for activation to obtain the probability distribution of the modulation type of the signal. The method proposed in the present application shows excellent accuracy and scalability at different signal-to-noise ratio levels and modulation modes, providing favorable support for cognitive radio and other wireless communication systems in practical applications.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of wireless communication technology, and in particular to a radio signal modulation classification method based on a deep integrated network. Background Art

[0002] Radio modulation classification is a key link in the effective management and utilization of radio spectrum resources, and aims to identify radio signals in non-communication environments. As a key technology in cognitive radio, Automatic Modulation Classification (AMC) provides the ability to automatically identify modulation modes for non-cooperative communications and has been widely used in military and civilian communications. However, since signals are interfered by noise, multipath effects, frequency offset and other factors during transmission, the task of modulation classification is full of challenges.

[0003] Before the emergence of deep learning technology, traditional modulation classification methods relied on manually designed features and expert knowledge to achieve classification by extracting features such as high-order cumulants. However, these methods have poor adaptability in complex environments, especially in low signal-to-noise ratio (SNR) scenarios, and it is difficult to achieve satisfactory performance.

[0004] With the development of deep learning, data-driven deep learning methods have gradually become a research hotspot in the field of AMC. The introduction of models such as convolutional neural networks (CNN) and recurrent neural networks (RNN) has made it possible to automatically extract features, significantly improving the accuracy of modulation classification. At the same time, deep learning models have shown strong advantages in feature extraction and pattern recognition, and can better cope with complex channel conditions.

[0005] In recent years, attention mechanisms and Transformer architectures have achieved remarkable results in natural language processing, image recognition and other fields, and their ability to process long-term dependencies and capture global information has been widely recognized. The AMC method that introduces the attention mechanism improves the recognition accuracy and robustness of the model by focusing on key information areas. On the other hand, the Transformer architecture shows strong adaptability in processing complex modulated signals due to its non-sequential dependencies and self-attention mechanism. For modulation recognition tasks in different environments, many researchers have proposed a variety of innovative AMC methods by combining convolutional neural networks, attention mechanisms and Transformer models.

[0006] Although deep learning technology has shown significant advantages in the field of modulation classification due to its powerful feature representation capabilities, a single network architecture (such as convolutional neural networks or long short-term memory networks) has limitations in its ability to model local features, sequential patterns, and global dependencies. Summary of the invention

[0007] In view of this, the present application provides a radio signal modulation classification method based on a deep integrated network to overcome the limitations of the single network architecture in the prior art.

[0008] To achieve the above objectives, the technical solutions adopted in this application are as follows:

[0009] A radio signal modulation classification method based on a deep integration network, comprising:

[0010] S1: Construct an attention CNN module and input the IQ signal samples into the attention CNN module to convert them into multi-channel feature maps. The feature maps obtained at this time contain local features extracted from the IQ signal;

[0011] S2: construct a time series information extraction module (LSTM), use the time series information extraction module to process the feature map containing the local features of the IQ signal output by the attention CNN module, and then extract the sequential pattern from the feature map;

[0012] S3: construct a vision transformer module, use the output of the temporal information extraction module to input into the global features extractor module to extract global features, so as to obtain a feature vector representing the modulation category;

[0013] S4: Construct a classifier, use the classifier to map the feature vector representing the modulation category from a high-dimensional space to a low-dimensional space representing the signal type, and then use softmax to activate it to obtain the probability distribution of the modulation type of the signal.

[0014] Furthermore, the step S1 is specifically as follows:

[0015] S11: Input specification in the attention CNN module is IQ signal samples S, B represents the input batch size, and L represents the length of the signal sample;

[0016] S12: The input signal sample passes through a 1D (one-dimensional) convolutional layer to extract a feature matrix containing local features in multiple dimensions;

[0017] S13: Through the channel attention module, feature maps with multiple channels are weighted to enhance the expressiveness of feature maps that are beneficial to the target task and suppress feature maps that are not significantly effective for the target task;

[0018] S14: Apply batch normalization layer to make all features have the same mean and variance, thereby removing the correlation between features and reducing the training time of the model;

[0019] S15: Use the nonlinear Sigmoid activation function to obtain nonlinear features and further enhance the expressiveness of the feature map;

[0020] S16: Use a size of , the maximum pooling layer with a step size of 2 pools the feature map to obtain a more compact feature representation to simplify the complexity of the CNN model.

[0021] Furthermore, the step S2 is specifically as follows:

[0022] S21: Map the output feature of the attention CNN module Pass it to the LSTM module to process the timing information;

[0023] S22: Forget gate calculation, the output calculation process of the forget gate is shown in formula (5);

[0024]

[0025] in is the activation function, represents the output of the forget gate, represents the forget gate weight matrix, represents the feature matrix of the hidden layer of the forget gate, represents the bias vector of the forget gate, represents the attention CNN module at time step t The extracted feature maps, h t-1 Indicates that the hidden layer is t -1 moment output; the forget gate is used to determine which positions in the input sequence are retained;

[0026] S23: Memory Gate Computation, Memory Gate The calculation formula is as follows:

[0027]

[0028] In formula (6), t The calculation process of the moment memory gate, Indicates that the input gate is at time step t The output, The role of is to retain the important information in the IQ signal after being processed by the attention convolutional neural network module and determine the amount of new information to be retained, where is the activation function; and is the input gate parameter, represents the input gate weight matrix, represents the feature matrix of the hidden layer of the input gate, represents the bias of the input gate; h t-1 Indicates that the hidden layer is t -1 output at time, represents the attention CNN module at time step t Extracted feature maps;

[0029] In formula (7), Represents the candidate cell state value, by the time step t The output of the hidden layer at time -1 and t The input at the moment is calculated using the tanh function and is used to update the hidden state of the LSTM network at the current moment; and is a parameter representing the cell state, represents the weight matrix of the input state, The feature matrix of the hidden layer representing the input state, represents the bias of the input state, tanh represents the hyperbolic tangent function, h t-1 Indicates that the hidden layer is t -1 output at time represents the attention CNN module at time step t Extracted feature maps;

[0030] In formula (8), represents the output of the forget gate, C t express t The cell state at a given moment is given by t Cell state at -1 C t-1 and t Candidate cell state values ​​at time Calculated, Indicates that the input gate is at time step t Output:

[0031] In the long short-term memory network (LSTM), candidate cells (Candidate Cell), also called candidate cell states, are intermediate variables used in LSTM to temporarily store and update cell states; "cell" usually refers to cell states (CellState), which is the core component of LSTM for storing and transmitting long-term information;

[0032] S24: Output gate calculation, output gate The features used to determine the final output are calculated as follows:

[0033]

[0034] Output Gate The detailed calculation process of is shown in formula (9), where is the activation function, represents the output gate weight matrix, represents the feature matrix of the hidden layer of the output gate, represents the bias of the output gate, h t-1 Indicates that the hidden layer is t -1 output at time represents the attention CNN module at time step t Extract feature maps; Output of LSTM at all times As a new signal feature, it is denoted as , It contains both the local characteristics of the signal and the timing information of the signal;

[0035] S25: Hidden layer state calculation, hidden layer output of LSTM Calculated by the following formula:

[0036]

[0037] in is the hidden state of LSTM at the current moment, which is used as the final output of the network to provide the signal characteristics at the current moment; Depend on t Cell state at a moment Through non-linear activation function After processing, and with the output of the output gate Calculated by performing a dot multiplication operation.

[0038] Furthermore, the step S3 specifically includes:

[0039] S31: Sequence output from the timing information extraction module A learnable class vector classtoken is added before the input, and the learnable class vector class token is responsible for capturing the semantic information of the entire sequence; then, a learnable position encoding vector is added to the entire sequence, and the position encoding vector actively learns the position information of each element in the sequence through the back propagation algorithm; finally, after adding the position encoding and class vector classtoken to the data input to the Transformer module, a new sequence containing position information is obtained, which is recorded as ;

[0040] S32: Construct the self-attention mechanism in the Transformer encoder to give the given feature vector, i.e. the new sequence Transformed into three new matrices, the three new matrices are respectively realized by multiplying the given eigenvectors by different weight matrices, as shown in formulas (11)-(13).

[0041]

[0042] in, 、 and denote the query matrix, key matrix, and value matrix respectively; 、 and are the weight matrices used to generate the query matrix, key matrix, and value matrix respectively; represents a given feature vector;

[0043] S33: Using what you get Q, K and V Matrix is ​​used to calculate the attention weight of the feature vector at each position. The specific calculation process is shown in formula (14);

[0044]

[0045] In formula (14), Indicates the dimension of the key, which is used to control the result of the dot product to prevent the network gradient from being too small, resulting in unstable network training; Indicates that the attention score is normalized to obtain the query and key Attention score;

[0046] S34: Calculate the attention output. In order to enhance the representation ability of the model, the Transformer layer adopts a multi-head attention mechanism. Through multiple parallel attention heads, the model can calculate the attention output from multiple angles. Suppose there is heads, the output of the multi-head attention mechanism is shown in formulas (15) and (16);

[0047]

[0048] In formulas (15) and (16), express The attention output of the head, It means concatenating the outputs of all attention heads and multiplying them by a projection matrix , the output of the multi-head attention is projected into the same dimensional space as the input; the semantic information of the entire sequence is captured through the Transformer network.

[0049] Furthermore, the step S4 specifically includes:

[0050] S41: The probability process of calculating the modulation type is shown in formulas (17)-(19);

[0051]

[0052] in, class_token It is a tag containing global information after being processed by the Transformer network, which is used to represent the semantic information of the entire input sequence; Dropout 1 represents a regularization method, Dropout In order to improve the generalization ability of the model and avoid overfitting, a structure of two fully connected layers is adopted. and Represent the outputs of the first and second fully connected layers, respectively. and are two fully connected layers, Prob The probabilities corresponding to different modulation types;

[0053] S42: After obtaining the probability distribution of each modulation type, determining the modulation type of the received signal according to the maximum probability, the calculation method is shown in formula (20);

[0054]

[0055] in, ; represents the probability of each modulation type predicted by the model;

[0056] S43: In order to optimize the entire model, the cross entropy is used as the objective function to train the entire network; the objective loss function is shown in formula (21);

[0057]

[0058] in, Represents the unique hot encoding of the true label, the vector length is the number of categories ; Represents the true label vector No. elements, with a value of 1 or 0. When the class is =1, otherwise =0; Represents the probability distribution vector predicted by the model, and its length is also the number of categories ; Each element in Indicates that the model predicts that the sample is The probability of the class.

[0059] Compared with the prior art, the beneficial effects of this application are:

[0060] 1. The method proposed in this application demonstrates excellent accuracy and scalability under different signal-to-noise ratio levels and modulation modes, providing favorable support for cognitive radio and other wireless communication systems in practical applications.

[0061] 2. The method proposed in this application achieves a good balance between model parameters and classification accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0063] Figure 1 A flow chart of a radio signal modulation classification method based on a deep integrated network for this application;

[0064] Figure 2 This is the architecture diagram of the deep integration network model for this application;

[0065] Figure 3 This is a schematic diagram of the structure of the channel attention module of this application;

[0066] Figure 4 This is a schematic diagram of the structure of the timing information extraction module of this application;

[0067] Figure 5 This is a schematic diagram of the Transformer encoder unit module structure of this application;

[0068] Figure 6 A line graph showing the performance of this application in the RML2016.10a dataset classified by signal-to-noise ratio and modulation type;

[0069] Figure 7 A line graph showing the performance of this application on the RML2016.10b dataset classified by signal-to-noise ratio and modulation type. DETAILED DESCRIPTION

[0070] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments.

[0071] like Figure 1-2 As shown, a radio signal modulation classification method based on a deep integration network includes:

[0072] S1: Construct an attention CNN module and input the IQ signal samples into the attention CNN module to convert them into multi-channel feature maps. The feature maps obtained at this time contain local features extracted from the IQ signal;

[0073] S2: construct a time series information extraction module (LSTM), use the time series information extraction module to process the feature map containing the local features of the IQ signal output by the attention CNN module, and then extract the sequential pattern from the feature map;

[0074] S3: construct a Vision Transformer module, use the output of the temporal information extraction module to input into the Global Features Extractor module to extract global features, so as to obtain a feature vector representing the modulation category;

[0075] S4: Construct a classifier, use the classifier to map the feature vector representing the modulation category from a high-dimensional space to a low-dimensional space representing the signal type, and then use softmax to activate it to obtain the probability distribution of the modulation type of the signal.

[0076] Furthermore, the step S1 is specifically as follows:

[0077] S11: Input specification in the attention CNN module is IQ signal samples S, B represents the input batch size, and L represents the length of the signal sample;

[0078] An IQ signal is a time series signal composed of an in-phase component and a quadrature component. The signal has two channels with equal lengths. Each IQ signal is usually described as The neural network can simultaneously receive and process a certain batch of IQ signals, which can significantly accelerate the convergence process of the network; in the attention convolutional neural network (CNN) module, the input IQ signal sample is denoted as S, and its input specification is Among them, B represents the input batch size, which determines the number of IQ signals processed simultaneously each time; L represents the length of the signal sample. In practical applications, the length is determined by the specific application scenario.

[0079] S12: The input signal sample passes through a 1D (one-dimensional) convolutional layer to extract a feature matrix containing local features in multiple dimensions;

[0080] Specifically, the method of this application uses a size of , a convolution kernel with a stride of 1, and the number of output channels is set to 256.

[0081] S13: Through the channel attention module, feature maps with multiple channels are weighted to enhance the expressiveness of feature maps that are beneficial to the target task and suppress feature maps that are not significantly effective for the target task;

[0082] like Figure 3 As shown in Figure 1, the channel attention is calculated by adaptive max pooling and average pooling. First, the input feature map is processed by an adaptive max pooling layer and an average pooling layer respectively, and the size of each feature map output is , Indicates the number of channels of the feature map. In order to enable the feature map to be input into the fully connected layer for processing, it is necessary to cancel the last dimension of the feature map. That is, the feature map after canceling the last dimension becomes ; Then the feature maps processed by the adaptive maximum pooling layer and the adaptive average pooling layer are input into two fully connected layers respectively to project them into the same vector space to obtain two weight vectors, and then the information of the two weight vectors is fused by vector addition, and finally a Sigmoid activation layer is used to obtain the weight of the feature map of each channel;

[0083] S14: Apply batch normalization layer to make all features have the same mean and variance, thereby removing the correlation between features and reducing the training time of the model;

[0084] S15: Use the nonlinear Sigmoid activation function to obtain nonlinear features and further enhance the expressiveness of the feature map;

[0085] S16: Use a size of , the maximum pooling layer with a step size of 2 pools the feature map to obtain a more compact feature representation to simplify the complexity of the attention CNN model.

[0086] The formula for extracting feature representation by the attention CNN module is:

[0087]

[0088] Where S is the original IQ signal sample; Represents a one-dimensional convolution operation; represents the channel attention module; Represents a batch normalization operation; Represents the Sigmoid activation function; represents the maximum pooling operation; Represents a more compact feature representation, which has higher dimension and richer semantic information than S.

[0089] Further, such as Figure 4 As shown, the step S2 is specifically as follows:

[0090] S21: Map the output feature of the attention CNN module Passed to the LSTM module for timing

[0091] Processing of information;

[0092] In order to make the best use of the time series information contained in the signal itself, the LSTM network is used to further capture the timing characteristics of the signal to enhance the accuracy of modulation recognition.

[0093] S22: Forget gate calculation, the output calculation process of the forget gate is shown in formula (5);

[0094]

[0095] in is the activation function, represents the output of the forget gate, represents the forget gate weight matrix, represents the feature matrix of the hidden layer of the forget gate, represents the bias vector of the forget gate, represents the attention CNN module at time step t The extracted feature maps, h t-1 Indicates that the hidden layer is t The output at time -1; the forget gate is used to determine which positions in the input sequence are retained.

[0096] S23: Memory Gate Computation, Memory Gate The calculation formula is as follows:

[0097]

[0098] In formula (6), t The calculation process of the moment memory gate, Indicates that the input gate is at time step t The output, The role of is to retain the important information in the IQ signal after being processed by the attention convolutional neural network module and determine the amount of new information to be retained, where is the activation function; and is the input gate parameter, represents the input gate weight matrix, represents the feature matrix of the hidden layer of the input gate, represents the bias of the input gate; h t-1 Indicates that the hidden layer is t -1 output at time represents the attention CNN module at time step t Extracted feature maps;

[0099] In formula (7), Represents the candidate cell state value, by the time step t The output of the hidden layer at time -1 and t The input at the moment is calculated using the tanh function and is used to update the hidden state of the LSTM network at the current moment; and is a parameter representing the cell state, represents the weight matrix of the input state, The feature matrix of the hidden layer representing the input state, represents the bias of the input state, tanh represents the hyperbolic tangent function, h t-1 Indicates that the hidden layer is t -1 output at time represents the attention CNN module at time step t Extracted feature maps;

[0100] In formula (8), represents the output of the forget gate, C t express t The cell state at a given moment is given by t Cell state at -1 C t-1 and t Candidate cell state values ​​at time Calculated, Indicates that the input gate is at time step t Output:

[0101] In the long short-term memory network (LSTM), candidate cells (Candidate Cell), also called candidate cell states, are intermediate variables used in LSTM to temporarily store and update cell states. "Cell" usually refers to cell state, which is the core component of LSTM for storing and transmitting long-term information. It plays a vital role in the process of LSTM processing sequence data. Both cells (cell state) and candidate cells (candidate cell state) play an important role in the information processing process, but there are obvious differences between the two: cell state is used to store and transmit long-term information. C t It is the core information storage unit in LSTM, like a "long-term memory bank", which can continuously save and transmit information throughout the entire sequence processing process, running through each time step of the time series, and is responsible for remembering all important information from the beginning of the sequence to the current moment, and using and updating this information as needed in subsequent time steps. Candidate cell states are denoted by Represents an intermediate variable temporarily calculated based on the current input and the hidden state of the previous moment. It represents the new information that may be added to the cell state at the current moment. It can be regarded as a kind of "encoding" of the input information at the current moment, providing new candidate content for updating the cell state.

[0102] S24: Output gate calculation, output gate The features used to determine the final output are calculated as follows:

[0103]

[0104] Output Gate The detailed calculation process of is shown in formula (9), where is the activation function, represents the output gate weight matrix, represents the feature matrix of the hidden layer of the output gate, represents the bias of the output gate, h t-1 Indicates that the hidden layer is t -1 output at time represents the attention CNN module at time step t Extract feature maps; Output of LSTM at all times As a new signal feature, it is denoted as , It contains both the local characteristics of the signal and the timing information of the signal;

[0105] S25: Hidden layer state calculation, hidden layer output of LSTM Calculated by the following formula:

[0106]

[0107] in is the hidden state of LSTM at the current moment, which is used as the final output of the network to provide the signal characteristics at the current moment; Depend on t Cell state at a moment Through non-linear activation function After processing, and with the output of the output gate Calculated by performing a dot multiplication operation.

[0108] Furthermore, the step S3 specifically includes:

[0109] S31: Before using the global feature extractor to extract global features, it is necessary to extract the sequence output by the temporal information extraction module. A learnable class vector class token is added before the input, and the learnable class vector class token is responsible for capturing the semantic information of the entire sequence; then, a learnable position encoding vector is added to the entire sequence, and the position encoding vector actively learns the position information of each element in the sequence through the back propagation algorithm; finally, after adding the position encoding and class vector class token to the data input to the Transformer module, a new sequence with position information is obtained, which is recorded as ;

[0110] S32: Build the Transformer encoder (structure as Figure 5 The self-attention mechanism in (shown) is used to Transformed into three new matrices, the three new matrices are respectively realized by multiplying the given eigenvectors by different weight matrices, as shown in formulas (11)-(13).

[0111]

[0112] in, 、 and denote the query matrix, key matrix and value matrix respectively, 、 and They are weight matrices used to generate query matrix, key matrix and value matrix respectively. These weight matrices are parameters learned during model training; represents a given feature vector;

[0113] S33: Using what you get Q, K andV Matrix is ​​used to calculate the attention weight of the feature vector at each position. The specific calculation process is shown in formula (14);

[0114]

[0115] In formula (14), firstly Q Matrix and K Matrix to calculate the similarity matrix of the input sequence elements, Indicates the dimension of the key, which is used to control the result of the dot product to prevent the network gradient from being too small, resulting in unstable network training; Indicates that the attention score is normalized to obtain the query and key The attention score is finally obtained according to the attention score Perform weighted summation to obtain the final attention output, which contains information about elements at other positions.

[0116] S34: Calculate the attention output. In order to enhance the representation ability of the model, the Transformer layer adopts a multi-head attention mechanism. Through multiple parallel attention heads, the model can calculate the attention output from multiple angles. Assume that there is heads, the output of the multi-head attention mechanism is shown in formulas (15) and (16);

[0117]

[0118] In formulas (15) and (16), express The attention output of the head, It means concatenating the outputs of all attention heads and multiplying them by a projection matrix , the output of the multi-head attention is projected into the same dimensional space as the input; the semantic information of the entire sequence is captured through the Transformer network;

[0119] Furthermore, the step S4 specifically includes:

[0120] S41: The probability process of calculating the modulation type is shown in formulas (17)-(19);

[0121]

[0122] in, class_token is a tag containing global information after being processed by the Transformer network, which is used to represent the semantic information of the entire input sequence; 1 represents a regularization method that randomly sets the output of neurons to zero with a certain probability, thereby enhancing the generalization ability of the model. Dropout In order to improve the generalization ability of the model and avoid overfitting, a structure of two fully connected layers is adopted. and Represent the outputs of the first and second fully connected layers, respectively. and are two fully connected layers, Prob The probabilities corresponding to different modulation types;

[0123] S42: After obtaining the probability distribution of each modulation type, determining the modulation type of the received signal according to the maximum probability, the calculation method is shown in formula (20);

[0124]

[0125] in, ; represents the probability of each modulation type predicted by the model;

[0126] S43: In order to optimize the entire model, the cross entropy is used as the objective function to train the entire network; the objective loss function is shown in formula (21);

[0127]

[0128] in, Represents the unique hot encoding of the true label, the vector length is the number of categories ; Represents the true label vector No. elements, with a value of 1 or 0. When the class is =1, otherwise =0; Represents the probability distribution vector predicted by the model, and its length is also the number of categories ; Each element in Indicates that the model predicts that the sample is The probability of the class.

[0129] In order to demonstrate the superiority of the method proposed in this application, the following experimental settings are used to further illustrate the method proposed in this application and compare it with a series of advanced methods in modulation classification:

[0130] Before the comparative experiment, in order to ensure the fairness of the experiment, the two datasets RML2016.10a and RML2016.10b were randomly divided into a ratio of 6:2:2. Then the method proposed in this application and the baseline method (CNN1, CNN2, CLDNN, DenseNet, ResNet, GRU and LSTM) were trained, verified and tested on the same training set, validation set and test set.

[0131] All experiments were completed on the same server. The hardware environment for training was AMD EPYC 754332-CoreProcessor and NVIDIA A100 Tensor Core GPU, and the software environment was pytorch and python3.7. In the SGD optimizer, the initial learning rate was set to 10-3, the momentum was set to 0.9, and the weight_decay was set to 1e-4. At the same time, in order to compare each method more fairly, the batch size and epoch of the training network were uniformly set to 400 and 2000. When training all networks, the cross entropy loss function was used as the objective function, and the SGD optimizer was used to optimize the network.

[0132] The classification performance is evaluated by the indicator parameter Accuracy coefficient and Params model parameters.

[0133] First, the method proposed in this application is compared with the baseline method on the two datasets RML2016.10a and RML2016.10b. Among these methods, CNN1 and CNN2 both use convolution kernels and maximum pooling layers to extract local features of the signal, and then use a fully connected layer to project these features into a low-dimensional vector space representing the signal type. The two models, CLDNN recurrent convolutional neural network and DenseNet densely connected neural network, introduce jump connections to enable information to flow to deeper network layers to avoid information loss. In addition, CLDNN combines the characteristics of two different architecture networks, convolution and LSTM, to perform modulation classification, showing its uniqueness. LSTM and GRU gated recurrent units start from the perspective that the input signal has the characteristics of the time dimension, and summarize the information of different time steps and input them into a fully connected layer for modulation classification. Among them, CNN1, CNN2, CLDNN, DenseNet densely connected neural network and ResNet residual network are all modulation classification networks based on convolution architecture. The results are shown in Table 1. CNN1, CNN2, CLDNN, DenseNet, ResNet, GRU and LSTM represent baseline methods, Accuracy(%) represents accuracy, Params(K) represents model parameters, and the ones in bold font in the table are the optimal results.

[0134] Table 1 Comparison results between the proposed method and the baseline method on two datasets

[0135]

[0136] According to the data in Table 1, the method of this application has demonstrated its unique advantages on both datasets. On the RML2016.10a dataset and the RML2016.10b dataset, the classification accuracy of the method proposed in this application is the highest, increasing by 3.9% to 10.46% on the RML2016.10a dataset and by 3.62% to 7.71% on the RML2016b dataset. In terms of model parameters, the number of parameters of the method of this application is much smaller than that of the network parameters of the modulation classification based on the convolutional architecture, reducing by 395.59K to 2820.07K.

[0137] In summary, the method of this application achieves a good balance between model parameters and classification accuracy.

[0138] Secondly, in order to further evaluate the classification performance of the model under different signal-to-noise ratios, the classification performance of the model under different signal-to-noise ratios was statistically analyzed. The statistical results are as follows: Figure 6 and Figure 7As shown. In these two data sets, when the signal-to-noise ratio is lower than -8dB, the classification results of the method of the present application and the baseline method are similar. This is because under low signal-to-noise ratio, the signal distortion is more serious, and it is difficult for all models to extract applicable features from the original signal for classification. At the same time, when the signal-to-noise ratio is higher than 0dB, the classification accuracy of the method of the present application is always higher than that of the baseline model, with an average accuracy of 90.5% on the RML2016.10a data set and an average accuracy of 92.5% on the RML2016.10b data set. On the RML2016.10a data set, when the signal-to-noise ratio is higher than 4dB, the accuracy of most baseline models is maintained at 83.5%, while the accuracy of the method proposed in this application is maintained at 91%, which is about 7.3% higher than the baseline. This shows that the method proposed in this application has unique advantages in extracting signal features. The performance of the method proposed in this application on the RML2016.10b data set further proves this concept. When the signal-to-noise ratio is lower than -5dB, the performance of the proposed method and the baseline method on the RML2016.10b and RML2016.10a datasets is basically the same. However, when the signal-to-noise ratio is higher than 0dB, the classification accuracy of the proposed method is significantly higher than the baseline model, maintaining at around 92.5%, an average improvement of 6.8% over the baseline.

[0139] Then, ablation experiments were conducted on each module on the RML2016.10a and RML2016.10b datasets to prove the effectiveness of each module. The method proposed in this application is divided into three small modules, namely CNN Backbone convolutional neural network backbone, Temporal Features Extractor temporal feature extractor and Global Features Extractor global feature extractor. Three groups of experiments were set up here. The first group of experiments only contained CNN Backbone and classifier; the second group added Temporal Features Extractor; the third group of experiments used the proposed model, which contained all modules. The specific details are shown in Table 2.

[0140] Table 2 Ablation study results on RML2016.10a and RML2016.10b

[0141]

[0142] Table 2 shows the degree of improvement of different modules of the deep integrated network model of this application for the task. It can be seen that after adding the temporal features extractor and the global features extractor, the classification accuracy of the model has been improved on both datasets, indicating that these two modules play an important role in the modulation classification task. In particular, after the introduction of the temporal features extractor, the classification accuracy of the model has been improved the most, increasing by 11.55% and 11.93% on the two datasets respectively, proving that temporal features play an important role in modulation recognition. Finally, after adding the global features extractor, the recognition accuracy of the model on the RML2016.10a and RML2016.10b datasets has been further improved by 6.44% and 4.91%, which further proves that the extraction of global features plays an important role in improving the accuracy of modulation classification. In short, the results of Table 2 show the importance of the two modules, the temporal features extractor and the global features extractor, for the modulation classification task.

[0143] The above are only specific implementations of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A radio signal modulation classification method based on deep integrated network, characterized in that: include: S1: Construct an attention CNN module and input the IQ signal samples into the attention CNN module to convert them into multi-channel feature maps. The feature maps obtained at this time contain local features extracted from the IQ signal; S2: construct a temporal information extraction module, and use the temporal information extraction module to process the feature map containing the local features of the IQ signal output by the attention CNN module, and then extract the sequential pattern from the feature map; S3: construct a visual transformer module, input the output result of the temporal information extraction module into the global feature extractor module to extract global features to obtain a feature vector representing the modulation category; S4: construct a classifier, use the classifier to map the obtained feature vector representing the modulation category from a high-dimensional space to a low-dimensional space representing the signal type, and then use softmax to activate it to obtain the probability distribution of the modulation type of the signal; The step S3 specifically includes: S31: Sequence output from the timing information extraction module A learnable class vector class token is added before the input, and the learnable class vector class token is responsible for capturing the semantic information of the entire sequence; then, a learnable position encoding vector is added to the entire sequence, and the position encoding vector actively learns the position information of each element in the sequence through the back propagation algorithm; finally, after adding the position encoding and class vector class token to the data input to the Transformer module, a new sequence with position information is obtained, which is recorded as ; S32: Construct the self-attention mechanism in the Transformer encoder to give the given feature vector, i.e. the new sequence Transformed into three new matrices, the three new matrices are respectively realized by multiplying the given eigenvectors by different weight matrices, as shown in formulas (11)-(13); ; in, , and denote the query matrix, key matrix, and value matrix respectively; , and are the weight matrices used to generate the query matrix, key matrix, and value matrix respectively; represents a given feature vector; S33: Using what you get , and Matrix is ​​used to calculate the attention weight of the feature vector at each position. The specific calculation process is shown in formula (14); ; In formula (14), Indicates the dimension of the key, which is used to control the result of the dot product to prevent the network gradient from being too small, resulting in unstable network training; Indicates that the attention score is normalized to obtain the query and key Attention score; S34: Calculate the attention output. In order to enhance the representation ability of the model, the Transformer layer adopts a multi-head attention mechanism. Through multiple parallel attention heads, the model can calculate the attention output from multiple angles. Suppose there is heads, the output of the multi-head attention mechanism is shown in formulas (15) and (16); ; In formulas (15) and (16), express The attention output of the head, It means concatenating the outputs of all attention heads and multiplying them by a projection matrix , the output of the multi-head attention is projected into the same dimensional space as the input; the semantic information of the entire sequence is captured through the Transformer network.

2. A radio signal modulation classification method based on a deep integrated network as claimed in claim 1, characterized in that: The step S1 is specifically as follows: S11: Input specification in the attention CNN module is IQ signal samples S, B represents the input batch size, and L represents the length of the signal sample; S12: The input signal sample passes through a 1D convolutional layer to extract a feature matrix containing local features in multiple dimensions; S13: Through the channel attention module, feature maps with multiple channels are weighted to enhance the expressiveness of feature maps that are beneficial to the target task and suppress feature maps that are not significantly effective for the target task; S14: Apply batch normalization layer to make all features have the same mean and variance, thereby removing the correlation between features and reducing the training time of the model; S15: Use the nonlinear Sigmoid activation function to obtain nonlinear features and further enhance the expressiveness of the feature map; S16: Use a size of , the maximum pooling layer with a step size of 2 pools the feature map to obtain a more compact feature representation to simplify the complexity of the CNN model.

3. A radio signal modulation classification method based on a deep integrated network as claimed in claim 2, characterized in that: The step S2 is specifically as follows: S21: Map the output feature of the attention CNN module Pass it to the LSTM module to process the timing information; S22: Forget gate calculation. The calculation process of the forget gate is shown in formula (5); ; in is the activation function, , and is the parameter of the forget gate, represents the attention CNN module at time step t The extracted feature maps, Indicates that the hidden layer is t -1 moment output; the forget gate is used to determine which positions in the input sequence are retained; S23: Memory Gate Computation, Memory Gate The calculation formula is as follows: ; In formula (6), t The calculation process of the moment memory gate, The role of is to retain the important information in the IQ signal after being processed by the attention convolutional neural network module and determine the amount of new information to be retained, where is the activation function; and is the input gate parameter, represents the input gate weight matrix, represents the feature matrix of the hidden layer of the input gate, represents the bias of the input gate; Indicates that the hidden layer is t -1 Output situation at time; In formula (7), From the time step t The output of the hidden layer at time -1 and t The input at the moment is used to update the hidden state of the LSTM network at the current moment; and is the parameter representing the input state, represents the weight matrix of the input state, The feature matrix of the hidden layer representing the input state, Indicates the bias of the input state; The internal state of the module is C t , will be t -1, the internal state is updated to t The internal state at the time is as shown in formula (8); S24: Output gate calculation, output gate The features used to determine the final output are calculated as follows: ; Output Gate The detailed calculation process of is shown in formula (9). represents the output gate weight matrix, represents the feature matrix of the hidden layer of the output gate, Represents the bias of the output gate; the output of LSTM at all times As a new signal feature, the It contains both the local characteristics of the signal and the timing information of the signal; S25: Hidden layer state calculation, hidden layer output of LSTM Calculated by the following formula: ; in is the hidden state of LSTM at the current moment, which is used as the final output of the network to provide the signal characteristics at the current moment; the internal state Through non-linear activation function Processing generates hidden states, which are recorded as .

4. A radio signal modulation classification method based on a deep integrated network as claimed in claim 3, characterized in that: The step S4 specifically includes: S41: The probability process of calculating the modulation type is shown in formulas (17)-(19); ; in, class_token It is a tag containing global information after being processed by the Transformer network, which is used to represent the semantic information of the entire input sequence; Dropout 1 represents a regularization method, Dropout In order to improve the generalization ability of the model and avoid overfitting, a structure of two fully connected layers is adopted. and Represent the outputs of the first and second fully connected layers, respectively. and are two fully connected layers, Prob The probabilities corresponding to different modulation types; S42: After obtaining the probability distribution of each modulation type, determine the modulation type of the received signal according to the maximum probability, and the calculation method is shown in formula (20); ; in, ; represents the probability of each modulation type predicted by the model; S43: In order to optimize the entire model, the cross entropy is used as the objective function to train the entire network; the objective loss function is shown in formula (21); ; in, Represents the unique hot encoding of the true label, the vector length is the number of categories ; Represents the true label vector No. elements, with a value of 1 or 0. When the class is ,otherwise ; Represents the probability distribution vector predicted by the model, and its length is also the number of categories ; Each element in Indicates that the model predicts that the sample is The probability of the class.

Citation Information

Patent Citations

  • Wheat cold resistance identification method based on hybrid neural network fused with Attention mechanism

    CN111651980A

  • Linear self-attention lightweight facial expression recognition method

    CN117912083A