Automatic modulation identification method based on full space-time dense neural network
Through the combined combination of the Generative Adversarial Network (GAN) and Bidirectional LSTM, the problem of insufficient feature extraction under low signal-to-noise ratio of automatic radio signal modulation recognition is solved, and higher recognition accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510274960.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-07-22
AI Technical Summary
In the automatic modulation recognition of radio signals, especially under low signal-to-noise ratio, it is difficult to fully mine the multi-domain characteristics of the space-time frequency energy of the signal, resulting in insufficient recognition accuracy and robustness.
Full-time and space-intensive neural network (FSTDNN) is used, combining generative adversarial network (GAN) and Bidirectional LSTM, and through dense connections, jump connections and attention mechanisms, feature extraction and timing data processing capabilities are enhanced to capture the spatial and temporal dynamic characteristics of the signal.
Under low signal-to-noise ratio conditions, the accuracy and robustness of automatic modulation recognition are significantly improved, and the generalization ability and feature recognition effect of the model are improved.
Smart Images

Figure CN120354093A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of automatic modulation recognition, and more specifically, to an automatic modulation recognition method based on a full spatio-temporal dense neural network. Background Art
[0002] Automatic modulation recognition technology (AMR) is a key task in the field of modern wireless communication, especially playing an important role in signal processing and spectrum sensing, and is crucial in fields such as spectrum monitoring, cognitive radio, and distributed signal processing. Driven by the continuous breakthrough of computing power, deep learning has become the focus of common attention in the academic and industrial circles. Similarly, DL-based AMR technology has also received extensive attention in the field of wireless communication due to its advantages in feature extraction and classification performance.
[0003] DL-based AMR technology generally includes three main steps: signal preprocessing, feature extraction, and modulation classification. In the preprocessing stage, the signal will undergo operations such as noise reduction and filtering to adjust the data format for subsequent processing. Feature extraction and classification can be completed end-to-end by a deep neural network, or divided into two steps. First, features are extracted by traditional methods in the preprocessing part, and then used to train a classification model. Existing deep learning models, such as convolutional neural networks (CNNs), recurrent neural networks (RNNs), long short-term memory (LSTMs), and graph convolutional networks (GCNs), have shown great potential in the field of AMR. These models can directly learn features automatically from signal data without manual design, which is especially valuable in complex electromagnetic environments. Although deep learning models have achieved remarkable success in the automatic modulation recognition (AMR) task, due to the characteristics of radio signals having multiple domains of space, time, frequency, and energy, existing feature extraction methods often cannot fully exploit the useful information therein, especially in complex electromagnetic environments with low signal-to-noise ratios. Therefore, a generative adversarial network (GAN) structure is specifically designed to enhance the robustness and feature extraction ability of the model under low signal-to-noise ratio conditions. Summary of the Invention
[0004] The content of the present invention is to provide an automatic modulation recognition method based on a full spatio-temporal dense neural network, which can better perform automatic modulation recognition.
[0005] An automatic modulation recognition method based on a full spatio-temporal dense neural network according to the present invention includes the following steps:
[0006] Step 1: Establish a full spatio-temporal dense neural network FSTDNN;
[0007] The Full-Space-Time Dense Neural Network (FSTDNN) includes an input layer for receiving data with a shape of input_shape, and the data is processed by a Generative Adversarial Network (GAN). Then there are a BatchNormalization layer and an Activation layer, introducing non-linearity through the ReLU activation function. After that, the data flows through two dense_blocks, each dense_block including multiple conv_blocks, and these conv_blocks achieve dense connection of features through identity_blocks. After each dense_block, there is a transition_block for reducing the dimension of the feature map and the number of channels, followed by a SEN_self_att layer to enhance the model's attention to important features through the self-attention mechanism.
[0008] After passing through two dense_blocks and one transition_block, there is a BidirectionalLSTM layer for capturing the temporal dynamic features of sequential data. If the include_top parameter is True, a GlobalAveragePooling2D layer is added after the BidirectionalLSTM layer, followed by a fully-connected layer and a softmax activation function for multi-class classification. If include_top is False, a GlobalAveragePooling2D or GlobalMaxPooling2D layer is selected according to the pooling parameter for feature extraction or transfer learning.
[0009] Step 2: Perform automatic modulation recognition through the Full-Space-Time Dense Neural Network (FSTDNN).
[0010] Preferably, the Generative Adversarial Network (GAN) includes a generator and a discriminator. The generator is responsible for generating feature representations, and the discriminator evaluates the quality of these features.
[0011] Preferably, in the dense_block, skip connections are introduced to solve the problem of gradient vanishing in deep networks. Each conv_block connects its output to the input and then passes it to the next conv_block. Between the dense_blocks, the output of the current Dense Block is reduced in dimension through 1x1 convolution and pooling operations and then passed to the next Dense Block.
[0012] Preferably, the BidirectionalLSTM layer solves the problems of gradient vanishing and gradient explosion by introducing a gating mechanism, and the gating mechanism includes an input gate, a forget gate, and an output gate.
[0013] The present invention designs a Full Spatio-Temporal Dense Neural Network (FSTDNN). By fully extracting temporal features and spatial features, enhancing the information extraction ability based on skip connections, and combining an attention mechanism to screen important and necessary information, this problem is further explored. The model will generate two types of feature representations. The first will more comprehensively capture the overall dynamic changes in time and space, and the second will more meticulously capture local features. FSTDNN is a full spatio-temporal dense neural network that can alleviate the overfitting problem on the premise of fully and effectively extracting the spatio-temporal multi-domain features of signals. Through experiments, the Full Spatio-Temporal Dense Neural Network (FSTDNN) has better results compared with other mainstream methods. Description of the Drawings
[0014] Figure 1 Schematic diagram of the Full Spatio-Temporal Dense Neural Network FSTDNN in the embodiment;
[0015] Figure 2 Schematic diagram of the signal of the RML2016.10a standard data set in the embodiment;
[0016] Figure 3 Schematic diagram of the comparison of the recognition rates of four models, namely CNN2, CLDNN, Resnet, and FSTDNN, with the signal-to-noise ratio ranging from -20 dB to +20 dB in the embodiment;
[0017] Figure 4 Schematic diagram of the confusion matrix of the recognition rates of each model when the signal-to-noise ratio is +18 dB in the RML2016.10a data set in the embodiment;
[0018] Figure 5 Schematic diagram of the learning rates of FSTDNN with and without BiLSTM under RML2016.10a in the embodiment;
[0019] Figure 6 Schematic diagram of the learning rates of FSTDNN with and without GAN under RML2016.10a in the embodiment;
[0020] Figure 7 Schematic diagram of the comparison of the recognition rates under different model depths in the embodiment;
[0021] Figure 8 Schematic diagram of the comparison of the recognition rates under different numbers of convolutional kernels in the embodiment;
[0022] Figure 9 Curve graph of the comparison of the recognition rates of four models, namely CNN2, CLDNN, Resnet, and FSTDNN, with the signal-to-noise ratio ranging from -20 dB to +20 dB in the RML2016.10b data set in the embodiment;
[0023] Figure 10Schematic diagram of the confusion matrix of the recognition rates of each model when the signal-to-noise ratio is +18 dB in the dataset RML2016.10b in the embodiment. Detailed implementation manners
[0024] To further understand the content of the present invention, the present invention will be described in detail in conjunction with the accompanying drawings and embodiments. It should be understood that the embodiments are only for explaining the present invention rather than limiting it.
[0025] Embodiment
[0026] This embodiment provides an automatic modulation recognition method based on a full spatio-temporal dense neural network, which includes the following steps:
[0027] Step 1, establish a full spatio-temporal dense neural network FSTDNN;
[0028] As Figure 1 shown, the full spatio-temporal dense neural network FSTDNN includes an input layer for receiving data with a shape of input_shape, and the data is processed by a generative adversarial network GAN; then there are a BatchNormalization layer and an Activation layer, introducing non-linearity through the ReLU activation function; after that, the data flows through two dense_blocks, each dense_block includes multiple conv_blocks, and these conv_blocks achieve dense connection of features through identity_blocks; after each dense_block, there is a transition_block for reducing the dimension of the feature map and reducing the number of channels, followed by a SEN_self_att layer to enhance the model's attention to important features through the self-attention mechanism;
[0029] After passing through two dense_blocks and a transition_block, there is a BidirectionalLSTM layer for capturing the temporal dynamic features of sequence data; if the include_top parameter is True, a GlobalAveragePooling2D layer is added after the BidirectionalLSTM layer, followed by a fully connected layer and a softmax activation function for multi-class classification; if include_top is False, a GlobalAveragePooling2D or GlobalMaxPooling2D layer is selected according to the pooling parameter for feature extraction or transfer learning;
[0030] Step 2, perform automatic modulation recognition through the full spatio-temporal dense neural network FSTDNN.
[0031] Optimize in the automatic modulation recognition task by drawing on the advantages and disadvantages of DenseNet. DenseNet is an efficient deep learning architecture. Its core idea is feature reuse, that is, each layer is connected to all previous layers. Such a design reduces the number of parameters and enhances feature transmission. In DenseNet, the output of each convolutional layer is directly connected to subsequent layers, which not only helps the direct propagation of the information flow but also alleviates the problem of vanishing gradients. The Transition layer plays an important role in DenseNet. They reduce the number and size of feature maps through 1x1 convolution and average pooling operations to make transitions between different DenseBlocks. In DenseNet, the number and size of feature maps are reduced through 2x2 convolution and average pooling operations. Such a design can effectively reduce the number of parameters and computational complexity of the model.
[0032] The full-space-time dense neural network proposed in this embodiment is a composite model integrating various deep learning techniques. At the beginning of the model, the input of the model is first processed by a generative adversarial network (GAN). In terms of the network architecture, it draws on the dense connection characteristics of DenseNet, and also introduces skip connections to enhance the gradient transmission of the network when dealing with deep structures, increases the depth of the network, and strengthens the direct transmission of features. At the same time, BiLSTM is incorporated to enhance the model's ability to process time-series data. In addition, the attention mechanism integrated in the network enables the model to focus on key information and further improve the accuracy of feature recognition. This makes the network perform excellently in complex space-time tasks such as automatic modulation recognition.
[0033] The generative adversarial network GAN includes a generator and a discriminator. The generator is responsible for generating feature representations, and the discriminator evaluates the quality of these features. GAN enhances the diversity of the dataset by generating additional training data, thereby improving the generalization ability of FSTDNN. This is particularly important in the case of scarce data because GAN can generate samples similar to real data but with different features, expanding the training set and potentially improving the model performance. In addition, the generator of GAN can learn the complex distribution of the input data and provide effective pre-training features for FSTDNN, helping it converge faster and improve the accuracy of feature recognition. The original GAN is divided into two components. One is the discriminator, which is used to distinguish real samples from generated samples; the other is the generator, which is used to generate fake samples to deceive the discriminator. Given a distribution, the probability distribution is defined as the distribution of the samples. The purpose of GAN is to learn the generator distribution that approximates the actual data distribution.
[0034] The Bidirectional LSTM layer can learn and remember long-term dependencies. By introducing a gating mechanism, it solves the problems of vanishing gradients and exploding gradients. The gating mechanism includes an input gate, a forget gate, and an output gate. These gating units can control the flow of information, enabling the network to retain or forget information when needed, and thus performing well in fields such as time series prediction and natural language processing. Skip connections solve the problem of vanishing gradients in deep networks by introducing cross-layer connections (or skip connections). In a residual network, a residual block contains two or more convolutional layers, and the input can directly skip some layers and be added to the subsequent layers. This helps the gradients flow directly to the previous layers, making the training of deep networks possible. The attention mechanism allows the model to focus on the most important parts of the input data. This mechanism is achieved by assigning different weights to each element in the sequence, enabling the model to pay more attention to the information that is more crucial for the current task. The attention mechanism has a wide range of applications in fields such as natural language processing and speech recognition and can significantly improve the performance of the model. The attention mechanism can be regarded as a tool for making full use of limited resources. By highlighting the important parts of the input signal, it makes the features more effectively transmitted. In modulation recognition, in order to retain the information in the I and Q channels of the signal, average pooling is performed, compressing each row into a value, and then a fully connected layer is used to obtain the weights for each channel.
[0035] The basic structure of FSTDNN adopts the dense connection method of DenseNet, which means that each layer is directly connected to all previous layers, thus realizing the reuse of feature maps. This design helps the propagation of gradients, alleviates the problem of gradient disappearance, and improves parameter efficiency. At the beginning of the model, a generative adversarial network (GAN) is used to process the input of data to help FSTDNN better adapt to various different input situations. In the stacked blocks, 2x2 convolutional kernels are used for feature extraction, and skip connections are introduced to facilitate the training of deep networks. The transition blocks are used for downsampling, reducing the scale of the feature maps through average pooling and reducing the computational amount. At the same time, an attention mechanism is introduced to learn the weights of each layer of the transferred feature layers, enabling the model to distinguish the importance of different feature layers and thus more effectively perform feature transfer. Inside FSTDNN, there are not only dense connections of DenseNet but also a large number of skip connections. In each dense_block, each conv_block connects its output to the input and then passes it to the next conv_block. Between dense_blocks, the output of the current Dense Block reduces the dimension through 1x1 convolution and pooling operations and then passes it to the next Dense Block. These skip connections allow information to flow more directly in the network, helping with the reuse of features and the propagation of gradients. At the end of FSTDNN, Bidirectional LSTM is added to process time series data. BiLSTM can capture long-term dependencies in time series, thus providing a more comprehensive representation of time series features.
[0036] Experiments and Analysis
[0037] The radio signal standard dataset RML2016.10a is adopted in this embodiment. This dataset modulates real signal sources and considers channel transmission effects such as sampling rate offset, center frequency offset, and multipath fading during the generation of IQ modulation signals, having high research value. It contains a total of 11 modulation signals, including analog modulation and digital modulation. The entire dataset has a total of 220,000 modulation signal samples, with the signal-to-noise ratio range from -20dB to 18dB, at an interval of 2dB. The number of modulation signal samples corresponding to each signal-to-noise ratio is 1000, and the data format is 2*128, where 2 corresponds to the I / Q two channels and 128 represents the number of sampling points. The visualization of the modulation signals in this dataset is as Figure 2 shown, where the red curve represents the I-channel signal and the blue curve represents the Q-channel signal.
[0038] In this embodiment, the RML2016.10b dataset is further used for experiments to further test and analyze the performance of the proposed method. Both RML2016.10a and RML2016.10b are datasets for modulation recognition, which contain signal data of different modulation types under different channel conditions. The RML2016.10a dataset contains 11 modulation methods, of which 8 are digital modulation methods and 3 are analog modulation methods, while the RML2016.10b dataset contains a total of 10 modulation methods. The total number of samples in the RML2016.10a dataset is 220,000, while the sample size of the RML2016.10b dataset is more abundant, reaching 1,200,000. The main difference between these two datasets is that they contain different channel models. RML2016.10a is a dataset based on the Rician channel model, while RML2016.10b is a dataset based on the Rayleigh channel model. This means that there are differences in parameters such as channel transmission characteristics, signal amplitude, and signal-to-noise ratio between the two datasets, thus affecting the modulation recognition of the datasets.
[0039] Recognition Rate and Comparative Analysis of Full Space-Time Dense Neural Network (FSTDNN)
[0040] This embodiment presents the experimental results of the proposed full space-time dense neural network (FSTDNN). The data samples are divided into a training set, a validation set, and a test set in a ratio of 7:1:2. The experimental platform uses the Windows 11 operating system, the server is configured with an NVIDIA GeForce RTX 3060Ti graphics card, the programming language is python 3.8, and the CUDA 11.0 toolkit is used to improve the parallel computing efficiency of the GPU. The network model is constructed using the keras architecture in TensorFlow 1.14.
[0041] RML2016.10a is a widely used signal modulation recognition dataset that can be used as a test benchmark. It contains samples of various modulation methods, with 20,000 samples for each modulation method. These samples are distributed under different signal-to-noise ratio (SNR) conditions, and each SNR condition contains 1,000 samples. To further enhance the diversity and complexity of the dataset and improve the generalization ability of the model, this embodiment introduces a generative adversarial network (GAN) to process the input data of the model. The generative adversarial network is a powerful deep learning model that generates new data samples through adversarial training between a generator and a discriminator. In the experiments of this embodiment, GAN will be used to process the input of the model. Specifically, for each modulation method, the goal is to increase the number of samples under each SNR to 2,000 after GAN processing. This means that after GAN processing, the number of samples for each modulation method will be twice that of the original dataset, thus providing a richer data resource for model training and testing. This process can not only increase the number of samples but also improve the model's recognition ability for different signals by introducing new data features, which is of great significance for improving the accuracy and robustness of signal modulation recognition.
[0042] The recognition effects of the proposed FSTDNN are experimentally compared with those of the models CNN2, CLDNN, and ResNet respectively, and the results are as Figure 3 shown. Among them, the CNN2 model is a variant of the convolutional neural network (CNN), which contains multiple convolutional layers and pooling layers for extracting features from the input data. CLDNN is a combination of the convolutional neural network (CNN), long short-term memory network (LSTM), and deep fully connected network. ResNet is designed with 34 layers, and FSTDNN is designed with 25 layers. In this experiment, FSTDNN has 4 stacked modules, and the number of residual blocks in each stacked module is: 3, 4, 3, 0, Figure 3 shows the recognition rates of each model at each SNR of the RML2016.10a standard dataset. Among them, the recognition rate of CNN2 is significantly lower than that of FSTDNN. When the SNR is very low, FSTDNN shows its powerful effect and is also 20% higher than the other three models at low SNR. When the SNR is greater than 0, the recognition rate of FSTDNN is approximately 7% - 8% higher than that of the CLDNN and ResNet models. At an SNR of 14 dB, the recognition rate of FSTDNN is approximately 8% higher than that of CLDNN. This result shows that FSTDNN has significant advantages compared with similar mainstream models in the problem of modulation recognition, especially between low SNRs of -14 dB to -4 dB.
[0043] Table 1 shows the average recognition rate, average time for one training cycle, and average time for testing one signal of each model in the range of 0 dB to +20 dB. It can be seen from the table that the average recognition rate of FSTDNN almost reaches 94%, while compared with other models, the highest recognition rate of other models, CLDNN, is only 87.2%. However, the training time of FSTDNN is longer. Its training time is longer than that of CLDNN. During the testing process, the time consumed by FSTDNN is slightly longer than that of ResNet-34. Although FSTDNN and ResNet-34 have similar numbers of layers, there is a nearly 4-fold difference in testing time. The main reason is that FSTDNN and ResNet have both skip connections and a feature reuse mechanism, which makes the feature maps that require convolution operations in each dense layer more. Coupled with the incorporation of BiLSTM and the attention mechanism. In addition, FSTDNN has more layers than CLDNN and shorter testing time, but FSTDNN has a higher recognition rate. The reason for the long time of CLDNN is due to the relatively large computational difficulty of BiLSTM.
[0044] Table 1 Average recognition rate, training time, and testing time of four models: CNN2, Resnet-34, CLDNN, and FDNN (The training time and testing time refer to the time required for one cycle and the time required for testing one signal, respectively)
[0045]
[0046]
[0047] The stronger the feature extraction of the deep learning model, the higher the recognition rate for each modulation type. Figure 4 Shows the confusion matrix of the recognition rates of CNN2, Resnet-34, CLDNN, and FSTDNN at a signal-to-noise ratio of +18 dB in the dataset RML2016.10a. In the confusion matrix, the recognition rates of each modulation type in the corresponding model can be clearly reflected. The darker the color, the higher the recognition rate. Overall, FSTDNN has a better recognition effect.
[0048] In order to verify the ability of BiLSTM in signal processing, this embodiment designs an experiment to evaluate its effect by comparing the learning rate of FSTDNN and the accuracy change after removing BiLSTM at a signal-to-noise ratio of -20 dB to +18 dB of the original model. The purpose of the experiment is to test the role of BiLSTM in enhancing the recognition ability of signals. As Figure 5 can be seen, as the last layer before signal output, BiLSTM can verify the signal processing ability of FSTDNN through the learning rate with and without BiLSTM in the model. Figure 5Shows the accuracy performance of two models under different signal-to-noise ratio (SNR) conditions. The red curve represents the K1 model, which is the complete FSTDNN, while the blue curve represents the K2 model, which is the performance of FSTDNN after removing BiLSTM.
[0049] As can be seen from the figure, under the condition of low signal-to-noise ratio (SNR less than -10dB), the accuracies of both models are relatively low, but the K1 model (red) is always slightly better than the K2 model (blue). As the signal-to-noise ratio increases, the accuracies of both models increase significantly, especially in the range of -10dB to 0dB, where the accuracy improvement is the most obvious. When the signal-to-noise ratio reaches 0dB and above, the accuracy of the K1 model tends to stabilize, approaching 0.925, showing superior performance under high signal-to-noise ratio conditions. While the K2 model also shows stability under high signal-to-noise ratio conditions, its accuracy is slightly lower than that of the K1 model, indicating that BiLSTM is beneficial for improving the performance of the model under high signal-to-noise ratio conditions.
[0050] In addition, the figure also contains an enlarged inset, focusing on showing the accuracies of the two models in the high signal-to-noise ratio range (0dB to 18dB). In this range, the accuracy of the K1 model always remains above 0.91, and as the signal-to-noise ratio increases, the accuracy slightly improves. In contrast, although the accuracy of the K2 model is also good under high signal-to-noise ratio conditions, it is always lower than that of the K1 model, which further confirms the contribution of BiLSTM in improving the model performance.
[0051] In summary, although the K2 model without BiLSTM performs similarly to the K1 model under low signal-to-noise ratio conditions, under high signal-to-noise ratio conditions, the K1 model containing BiLSTM can provide higher accuracy, showing the importance of BiLSTM for improving the model performance.
[0052] After the above comparison, it can be found that: BiLSTM in FSTDNN can effectively improve the learning ability of the model. This performance has been verified in the experiments with and without BiLSTM, where FSTDNN can achieve a higher learning rate. BiLSTM in FSTDNN can make full use of long-term temporal information when processing sequence data, automatically capture and learn the non-linear dynamic characteristics and temporal information in the system, thereby improving the effect of signal processing.
[0053] To verify the ability of GAN in FSTDNN, a series of experiments are designed in this embodiment to evaluate its effect by comparing the learning rates of the original model in FSTDNN at a signal-to-noise ratio of -20dB to +18dB and the learning rate after removing GAN. The purpose of the experiment is to test the role of GAN in enhancing the signal recognition ability.
[0054] In this embodiment, a comparative analysis of the performance of the Full Space-Time Dense Neural Network (FSTDNN) with and without the assistance of a Generative Adversarial Network (GAN) was carried out. The experimental results are as Figure 6 shown. The two curves in the figure respectively correspond to the accuracy changes of FSTDNN combined with GAN (K3) and without GAN (K4).
[0055] Analysis Figure 6 It can be found that in a low signal-to-noise ratio environment (SNR less than -10 dB), although the accuracies of both models are not high, the FSTDNN combined with GAN shows slightly superior performance compared to the model without GAN. As the signal-to-noise ratio increases, especially in the range of -10 dB to 0 dB, the accuracies of both models increase significantly, and this trend is particularly obvious in the model of FSTDNN combined with GAN. Under high signal-to-noise ratio conditions, the accuracy of the FSTDNN combined with GAN not only remains stable but is also slightly higher than that of the model without GAN as a whole. This result indicates that the introduction of GAN has a positive effect on improving the performance of FSTDNN in high and low signal-to-noise ratio environments.
[0056] Further observing the inset in the chart, it can be seen that in the high signal-to-noise ratio range, the accuracy of the FSTDNN combined with GAN basically remains above 0.92, and as the signal-to-noise ratio increases, the accuracy shows a slight upward trend. In contrast, although the model without GAN also shows a high accuracy under high signal-to-noise ratio conditions, it has never exceeded the model combined with GAN. This finding further confirms the significant contribution of GAN in improving the performance of FSTDNN, especially in enhancing the model's ability to identify high signal-to-noise ratio signals. At the same time, due to GAN, especially the GAN designed for automatic modulation format of radio data in this study, it enriches the diversity of input data, which is also helpful for improving the robustness of GSTDNN.
[0057] Comparison of the recognition rates of the Full Space-Time Dense Neural Network (FSTDNN) under different hyperparameters
[0058] In FSTDNN, the number of building blocks in each dense block is determined by the blocks parameter. In order to explore the influence of network depth on the recognition rate, Figure 7 the recognition rates of the model with 4 different model depths at each signal-to-noise ratio are shown. The parameters for the v1, v2, v3, and v4 cases are: (4, 6, 12, 12), (6, 12, 24, 16), (8, 16, 32, 32), (6, 12, 48, 32). As Figure 7As shown, as the model depth increases, the recognition rate further improves. As the network depth increases, the recognition rate increases from 90.8% to 92.3% at high signal-to-noise ratios. This experiment fully demonstrates that the increase in network depth has a certain improvement on the recognition rate of this modulation signal dataset. However, it can be seen from the figure that: excessive depth will instead cause the model to be difficult to converge, resulting in the inability to obtain the optimal model depth. A network with a moderate depth can ensure the recognition rate while also ensuring that the parameter scale of the network is easy to train.
[0059] To explore the impact of the change in the number of convolutional kernels on the model. In this experiment, models v1, v2, v3, and v4 were used for the experiment, and the numbers of convolutional kernels were 12, 15, 32, and 48 respectively. Figure 8 The figure shows the recognition rate curves of models with different numbers of convolutional kernels. It can be seen from the figure that as the number of convolutional kernels increases, the recognition rate has a certain improvement. However, as the model complexity increases, the model may overfit the training data, resulting in good performance on the training set but poor performance on unseen test data. This is because the model learns the noise and details in the training data without capturing the general features of the data. The change in the number of convolutional kernels has a greater impact on the size of the model. Selecting an appropriate number of convolutional kernels can ensure the recognition rate while also ensuring the model scale. The selection of network depth and network width is related to the software and hardware conditions in the actual application process.
[0060] The foregoing Figure 3 The figure shows and compares the recognition rates of the model at various signal-to-noise ratios in the RML2016.10a standard dataset. It can be seen that FSTDNN has certain advantages. Similarly, to explore the generalization ability of the model, we also conducted experimental verification on another dataset, the RML2016.10b dataset, and found that the accuracy of the model is generally high under high signal-to-noise conditions ( Figure 9 ), which may be related to the larger sample size of the RML2016.10b dataset, which provides more data to learn features. Among them, the accuracy of FSTDNN is close to 0.94, further verifying the effectiveness of the FSTDNN method.
[0061] Figure 10 The figure shows the confusion matrix of the recognition rates of CNN2, Resnet-34, CLDNN, and FSTDNN at a signal-to-noise ratio of +18dB in the RML2016.10b dataset. Similarly, by observing the depth of the diagonal color and combining the color distribution in the non-diagonal area, FSTDNN has a better recognition effect.
[0062] In this embodiment, a new deep learning model called Fully Spatio-Temporal Dense Neural Network (FSTDNN for short) is proposed. This model simultaneously arranges skip connections within and between dense modules to enhance the gradient transfer of the network when processing deep structures, and combines to capture the long-term dependencies in time series data, comprehensively improving the model's ability to extract spatio-temporal multi-domain features of radio data. The generative adversarial network GAN is introduced to improve the robustness of the model. In addition, the attention mechanism integrated in the network enables the model to focus on key information, further improving the accuracy of feature recognition. Experimental results on multiple radio datasets show that FSTDNN is more competitive than existing mainstream algorithms in performance and can better perform automatic modulation recognition.
[0063] The above schematically describes the present invention and its implementation manners. This description is not restrictive. What is shown in the drawings is only one of the implementation manners of the present invention, and the actual structure is not limited thereto. Therefore, if those of ordinary skill in the art are inspired by it and design similar structural manners and embodiments to this technical solution without creative efforts without departing from the purpose of the present invention, they shall fall within the protection scope of the present invention.
Claims
1. An automatic modulation recognition method based on a full-time and space dense neural network, characterized in that: It includes the following steps: Step 1, establish a full - space - time dense neural network FSTDNN; The full - space - time dense neural network FSTDNN includes an input layer for receiving data with the shape of input_shape, and the data is processed by a generative adversarial network GAN. Then comes the BatchNormalization layer and the Activation layer, introducing non - linearity through the ReLU activation function. After that, the data flows through two dense_blocks, each dense_block includes multiple conv_blocks, and these conv_blocks achieve dense connection of features through identity_blocks. After each dense_block, there is a transition_block for reducing the dimension of the feature map and the number of channels, followed by the SEN_self_att layer to enhance the model's attention to important features through the self - attention mechanism; After passing through two dense_blocks and a transition_block, there is a BidirectionalLSTM layer for capturing the temporal dynamic features of sequential data; If the include_top parameter is True, a GlobalAveragePooling2D layer is added after the BidirectionalLSTM layer, followed by a fully - connected layer and a softmax activation function for multi - class classification; if include_top is False, a GlobalAveragePooling2D or GlobalMaxPooling2D layer is selected according to the pooling parameter for feature extraction or transfer learning; Step 2, perform automatic modulation recognition through the full - space - time dense neural network FSTDNN.
2. The automatic modulation recognition method based on a full spatio-temporal dense neural network according to claim 1, characterized in that: The generative adversarial network GAN includes a generator and a discriminator. The generator is responsible for generating feature representations, and the discriminator evaluates the quality of these features.
3. The automatic modulation recognition method based on the full-time and full-space dense neural network according to claim 2, characterized in that: In the dense_block, skip connections are introduced to solve the vanishing gradient problem in deep neural networks. Each conv_block connects its output to the input and then passes it to the next conv_block; between dense_blocks, the output of the current DenseBlock is reduced in dimension through 1x1 convolution and pooling operations and then passed to the next Dense Block.
4. The automatic modulation recognition method based on the full spatio-temporal dense neural network according to claim 3, characterized in that: The BidirectionalLSTM layer solves the vanishing gradient and gradient explosion problems by introducing a gating mechanism, and the gating mechanism includes an input gate, a forget gate, and an output gate.