Automatic modulation recognition method based on three channels

Through the three-channel automatic modulation recognition method, the Fourier time-frequency domain transformation and space-time frequency fusion network model are used to solve the problem of high computational complexity and similar modulation types in the prior art, and achieve higher recognition accuracy and training speed.

CN118158044BActive Publication Date: 2025-08-12SOUTHWEST JIAOTONG UNIV

Patent Information

Application Number
CN202410433120.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-11
Publication Date
2025-08-12
Estimated Expiration
2044-04-11

AI Technical Summary

Technical Problem

The existing automatic modulation identification technology has the problems of high computational complexity, a lot of prior knowledge, poor generalization and robustness, and only dealing with time domain data leads to confusion between similar modulation types.

Method used

The automatic modulation recognition method based on three channels is adopted, and frequency information is extracted using Fourier time-frequency domain transformation, combined with the space-time frequency fusion network model, including convolutional neural network, recurrent neural network and fully connected deep neural network, to process signal characteristics of time, space and frequency dimensions, and separate in-phase/orthogonal data for feature extraction and fusion.

Benefits of technology

In the positive signal-to-noise ratio environment, the recognition accuracy is improved by 1% to 4%, and the training speed is increased by 2.26 times, which can better solve the confusion problem between high-dimensional modulated signals and improve the recognition accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118158044B_ABST
    Figure CN118158044B_ABST
Patent Text Reader

Abstract

The present invention relates to a three-channel automatic modulation recognition method. When a single-input single-output communication system receives an unknown signal, the time domain signal is discretely sampled in the time domain to form a time domain discrete signal, and one-dimensional complex data is formed. The Fourier time-frequency domain transform obtains frequency domain data. The frequency domain data is combined with the time domain discrete signal to obtain time-frequency domain discrete signal data. The frequency data in the time-frequency domain discrete signal data is fed into the two-dimensional convolution layer of the time-space frequency fusion network model. The in-phase data and orthogonal data in the time domain data are respectively fed into the corresponding one-dimensional convolution layer. The model determines the modulation format type of the unknown signal based on the characteristics of various modulation formats. In a positive signal-to-noise ratio environment, the present invention greatly improves the recognition accuracy of the modulation type of unknown electromagnetic spectrum signals, and the training speed is several times faster than that of existing models. For high-dimensional modulated signals, it can better solve the confusion problem between similar signals and achieve higher recognition accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of modulation type recognition of radio spectrum signals, and in particular to an automatic modulation recognition method based on three channels. Background Art

[0002] With the rapid development of traditional and emerging wireless services, automatic modulation recognition (AMR) technology plays an important role in military and civilian fields such as cognitive radio, spectrum regulation and interference identification in modern communication systems. Traditional automatic modulation recognition methods can be roughly divided into two categories: maximum likelihood ratio recognition methods based on hypothesis likelihood, and pattern recognition methods based on feature extraction. However, both methods have problems such as high computational complexity, the need for more prior knowledge, and poor generalization and robustness. With the rapid development of deep learning technology in computer vision, natural language processing and other fields, researchers have gradually begun to introduce deep learning technology to realize automatic modulation recognition of electromagnetic signals. Deep learning-based automatic modulation recognition (DL-AMR) technology does not require prior knowledge [1][2], can automatically extract target features, and significantly reduces feature engineering overhead. Therefore, convolutional neural networks (CNNs) [3], recurrent neural networks (RNNs) [4][5], and convolutional long-term deep neural network (CLDNN) [6] hybrid architectures have been applied to AMR, and their performance has been significantly improved compared with traditional AMR methods [7][8]. However, these methods usually directly or indirectly borrow architectures or methods from fields such as image processing and natural language processing, and are not specifically designed for AMR. In recent years, research has shifted to developing representative DL-AMR models that meet the requirements of high accuracy and low complexity. For example, in [9], a lightweight network architecture (PET-CGDNN) for phase parameter estimation and transformation based on CNN and gated recurrent unit (GRU) was proposed. In

[10] , a spatiotemporal fusion network architecture (MCLDNN) based on CNNs and long short-term memory network (LSTM) was designed to achieve high-accuracy recognition on the benchmark datasets RML2016.10a and RML2016.10b [3]

[11] and is considered to be one of the state-of-the-art (SoA) models. Ma et al.

[12] used the cumulative positive coordinate feature conversion technique to convert in-phase / quadrature (I / Q) data into Gram angular field (GAF) images as the input of Swin-Transformer. In

[13] , a hybrid architecture model based on residual neural network (ResNet) and LSTM was designed. However, most models only consider I / Q time domain data as input, ignoring other potential information. Processing only time domain data can lead to confusion between modulation types with similar time domain / waveform characteristics, limiting practical applications and missing the extraction and utilization of potential modulation information.

[0003] References

[0004] [1]Peng S,Sun S,Yao Y D.A survey of modulation classification usingdeep learning:Signal representation and data preprocessing[J].IEEETransactions on Neural Networks and Learning Systems,2021,33(12):7020-7038.

[0005] [2]Zhang F,Luo C,Xu J,et al.Deep learning based automatic modulationrecognition:Models,datasets,and challenges[J].Digital Signal Processing,2022,129:103650.

[0006] [3]O’Shea T J,Corgan J,Clancy T C.Convolutional radio modulationrecognition networks[C] / / Engineering Applications of Neural Networks:17thInternational Conference,EANN 2016,Aberdeen,UK,September 2-5,2016,Proceedings17.Springer International Publishing,2016:213-226.

[0007] [4]Rajendran S,Meert W,Giustiniano D,et al.Deep learning models forwireless signal classification with distributed low-cost spectrum sensors[J].IEEE Transactions on Cognitive Communications and Networking,2018,4(3):433-445.

[0008] [5]Huang S,Dai R,Huang J,et al.Automatic modulation classificationusing gated recurrent residual network[J].IEEE Internet of Things Journal,2020,7(8):7795-7807.

[0009] [6]West N E,O′shea T.Deep architectures for modulation recognition[C] / / 2017 IEEE international symposium on dynamic spectrum access networks(DySPAN).IEEE,2017:1-6.

[0010] [7]Dulek B.Online hybrid likelihood based modulation classificationusing multiple sensors[J].IEEE Transactions on Wireless Communications,2017,16(8):4984-5000.

[0011] [8]Hazza A,Shoaib M,Alshebeili S A,et al.An overview of feature-basedmethods for digital modulation classification[C] / / 2013 1st internationalconference on communications,signal processing,and their applications(ICCSPA).IEEE,2013:1-6.

[0012] [9]Zhang F,Luo C,Xu J,et al.An efficient deep learning model forautomatic modulation recognition based on parameter estimation andtransformation[J].IEEE Communications Letters,2021,25(10):3287-3290.

[0013]

[10] Xu J,Luo C,Parr G,et al.A spatiotemporal multi-channel learningframework for automatic modulation recognition[J].IEEE WirelessCommunications Letters,2020,9(10):1629-1632.

[0014]

[11] O′shea T J,West N.Radio machine learning dataset generation withgnu radio[C] / / Proceedings of the GNU Radio Conference.2016,1(1).

[0015]

[12] Ma W,Cai Z.Deep Learning Based Cognitive Radio ModulationParameter Estimation[J].IEEE Access,2023,11:20963-20978.

[0016]

[13] Elsagheer M M,Ramzy S M.A hybrid model for automatic modulationclassification based on residual neural networks and long short term memory[J].Alexandria Engineering Journal,2023,67:117-128.

[0017]

[14] Khandelwal S, Lecouteux B, Besacier L.Comparing GRU and LSTM for automatic speech recognition[D].LIG, 2016. Summary of the Invention

[0018] The present invention provides a three-channel based automatic modulation recognition method, which improves the recognition accuracy of radio automatic modulation and reduces calculation complexity.

[0019] The present invention is based on a three-channel automatic modulation identification method. When a single-input single-output communication system receives an unknown signal, it first performs time-domain discrete sampling on the in-phase / orthogonal time domain signals to form a time-domain discrete signal. Then, the time-domain discrete signal is combined into one-dimensional complex data, and Fourier time-frequency domain transform is performed to obtain frequency domain data. The frequency domain data is combined with the time-domain discrete signal to obtain time-frequency domain discrete signal data. The in-phase data and orthogonal data of the time domain data in the time-frequency domain discrete signal data are separated. Then, the frequency data in the time-frequency domain discrete signal data is fed into the two-dimensional convolution layer of the time-space-frequency fusion network model, and the in-phase data and orthogonal data are respectively fed into the corresponding one-dimensional convolution layer. The time-space-frequency fusion network model infers and determines the type of modulation format of the received unknown signal based on the unique characteristics of the various modulation formats that have been learned.

[0020] This invention targets single-input, single-output (SISO) communication systems and processes received unknown modulated signal data along three characteristic dimensions: time, space, and frequency. It utilizes Fourier transformed data and independent in-phase / quadrature channel (I / Q channel) data as input to thoroughly extract characteristic information from the data, improving recognition accuracy. The invention also expands the types of modulation information and utilizes frequency domain information to resolve potential confusion between signals with similar time domain / waveform characteristics.

[0021] The theoretical feasibility of this invention is based on two key reasons. First, frequency information demonstrates robustness to various interferences that may occur during signal transmission, such as noise, fading, and multipath. Second, separating multiple channels facilitates extracting independent features from the I / Q channels. The subsequent merging of the I / Q channels provides complementary information from both channels.

[0022] Furthermore, the Fourier time-frequency domain transform adopts discrete Fourier time-frequency domain transform.

[0023] Furthermore, the Fourier time-frequency transform adopts fast Fourier time-frequency transform to reduce the computational complexity of the model in the preprocessing stage.

[0024] Furthermore, the space-time frequency fusion network model includes a convolutional neural network module, a recurrent neural network module and a fully connected deep neural network module, and the input consists of the frequency domain data, the in-phase data and the orthogonal data.

[0025] Furthermore, in the convolutional neural module, the first two-dimensional convolutional layer is used to extract the frequency characteristics of the time-frequency domain discrete signal data, and the first one-dimensional convolutional layer and the second one-dimensional convolutional layer are used to extract the spatial characteristics of the time-frequency domain discrete signal data through convolution operations between consecutive frames, as well as the sliding window mechanism and edge detection capabilities of the convolution kernel. That is, the first two-dimensional convolutional layer inputs the frequency domain data in the time-frequency domain discrete signal data, the first one-dimensional convolutional layer inputs the in-phase data in the time-frequency domain discrete signal data, and the second one-dimensional convolutional layer inputs the orthogonal data in the time-frequency domain discrete signal data, forming a three-channel data input.

[0026] Furthermore, after the first one-dimensional convolution layer and the second one-dimensional convolution layer, there is a second two-dimensional convolution layer, which fuses the in-phase output data and the orthogonal output data after spatial feature extraction and feeds them into the second two-dimensional convolution layer for learning higher-level abstract features in the signal, and then combines the higher-level abstract features with the frequency domain output data of the first two-dimensional convolution layer based on the channel dimension so as to be perpendicular to the spatial correlation.

[0027] Furthermore, the output data of the second two-dimensional convolutional layer is fed into the third two-dimensional convolutional layer. The data channel dimension is compressed by setting the convolution kernel to fuse the spatiotemporal and frequency information, and the collaboration and complementarity between the data are utilized to aggregate useful information to complete the feature capture in the data.

[0028] Furthermore, after the convolutional neural module extracts the spatial and frequency dimensional features of the input signal data, the data is fed into the gated recurrent unit layer of the recurrent neural network module. The gated recurrent unit layer has three layers and is used to learn the long-term dependencies and time correlations of the data sequence and extract the time domain features of the unknown signal.

[0029] Furthermore, the output data of the recurrent neural network module is fed into the fully connected deep neural network module to map the features in the data to a space that is easier to separate. The fully connected layer of the fully connected deep neural network module includes a nonlinear activation function and a dropout layer. The final model output layer is a fully connected layer with n neurons, and the number n is determined by the type of modulation format to be identified input into the space-time frequency fusion network model during pre-training.

[0030] The beneficial effects of the present invention include:

[0031] 1. Under a positive signal-to-noise ratio environment, the model recognition accuracy is improved by 1% to 4% compared with the existing technology.

[0032] 2. The model training speed is 2.26 times faster than the model with the highest recognition accuracy in existing technologies.

[0033] 3. For signals with high-dimensional modulation schemes, such as 16-QAM and 64-QAM, it can better solve the confusion problem between similar signals and achieve higher recognition accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 This is a flow chart of the three-channel automatic modulation recognition method of the present invention.

[0035] Figure 2 This is the architecture diagram of the time-space-frequency fusion network model of the present invention.

[0036] Figure 3 This is a comparison chart of the recognition accuracy of the model of the present invention (MCGDNN) based on the RML2016.10a dataset and other existing models.

[0037] Figure 4 This is a comparison chart of the recognition accuracy of the model of the present invention (MCGDNN) based on the RML2016.10b dataset and other existing models.

[0038] Figure 5 This is a two-dimensional visualization diagram of the data of the model of the present invention (MCGDNN) under a signal-to-noise ratio of 10dB.

[0039] Figure 6 This is a two-dimensional visualization comparison diagram of the model of the present invention (MCGDNN) with and without frequency domain data at a signal-to-noise ratio of 10dB. DETAILED DESCRIPTION

[0040] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application for protection, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of this application.

[0041] like Figure 1As shown, the present invention is based on a three-channel automatic modulation identification method. When a single-input single-output communication system receives an unknown signal, the in-phase / orthogonal time domain signal is first discretely sampled by 128 points in the time domain to form a time domain discrete signal with a dimension of 2×128. The time domain discrete signal is then combined into one-dimensional complex data, and a Fourier time-frequency domain transform is performed to obtain frequency domain data. The frequency domain data is combined with the time domain discrete signal to obtain time-frequency domain discrete signal data with a dimension of 4×128. The in-phase data and the orthogonal data of the time domain data in the time-frequency domain discrete signal data are separated. Then, the frequency data in the time-frequency domain discrete signal data is fed into the two-dimensional convolution layer of the time-space-frequency fusion network model, and the in-phase data and the orthogonal data are respectively fed into the corresponding one-dimensional convolution layer. The time-space-frequency fusion network model infers and determines the type of modulation format of the received unknown signal based on the unique characteristics of the various modulation formats that have been learned.

[0042] This invention targets single-input, single-output (SISO) communication systems and processes received unknown modulated signal data along three characteristic dimensions: time, space, and frequency. It utilizes Fourier transformed data and independent in-phase / quadrature channel (I / Q channel) data as input to thoroughly extract characteristic information from the data, improving recognition accuracy. The invention also expands the types of modulation information and utilizes frequency domain information to resolve potential confusion between signals with similar time domain / waveform characteristics.

[0043] The theoretical feasibility of this invention is based on two key reasons. First, frequency information demonstrates robustness to various interferences that may occur during signal transmission, such as noise, fading, and multipath. Second, separating multiple channels facilitates extracting independent features from the I / Q channels. The subsequent merging of the I / Q channels provides complementary information from both channels.

[0044] To further illustrate the present invention in detail, the following description is given:

[0045] The present invention is designed for a single-input single-output (SISO) communication system. Generally, a signal received through a channel in a SISO communication system can be expressed as:

[0046] r(t)=s(t)*h(t)+n(t) (1);

[0047] In formula (1), s(t) represents the noise-free complex baseband envelope of the received signal, n(t) is the additive white Gaussian noise (AWGN) with zero mean, and h(t) is the time-varying impulse response of the channel during wireless signal transmission. For the transmitted and received analog signals s(t) and r(t), an analog-to-digital converter (ADC) is usually used to convert the received signal into a real-time signal. It is sampled at a rate where Ts is the sampling period. The sampled received signal is usually stored in I / Q format (in-phase / quadrature). Therefore, the discrete data can be expressed as:

[0048] r(n)=r I (n)+jr Q (n) (2);

[0049] where r I (n) and r Q (n) represent the in-phase component data and quadrature component data of the received signal, respectively.

[0050] By performing discrete Fourier transform (DFT) on the received signal, the frequency domain information of r(n) can be obtained:

[0051]

[0052] In formula (3), N represents the total number of sampling points, w N =e -2jπ / N , n=0, 1, ..., N-1. In order to reduce the computational complexity of the model preprocessing stage, the FFT (Fast Fourier Transform) algorithm is used instead of DFT. The FFT algorithm converts the one-dimensional (1D) DFT into a smaller-dimensional DFT, reducing the computational complexity of the DFT from O(N 2 ) is reduced to o(Nlog(N)).

[0053] In the most commonly used Cooley-Tukey (Cooley-Tukey Fast Fourier Transform) algorithm: N is decomposed into N = N1N2, where N1 and N2 represent the smaller DFT dimension currently being calculated and the smaller DFT dimension to be calculated after N is decomposed, respectively. n = N2n1 + n2, k = k1 + N1k2. The frequency domain information of r(n) after FFT conversion is:

[0054]

[0055] Among them, r FFT (k) represents the kth frequency component in the frequency domain, k = 0, 1, ..., N-1, and are the DFT factors of dimensions N1 and N2 respectively, is the rotation factor. The FFT algorithm is used to convert the discrete time domain data in the data set into frequency domain data, preparing for the subsequent extraction of spatial, temporal and frequency features.

[0056] The proposed time-space-frequency fusion network model architecture is as follows Figure 2It is a hybrid architecture that integrates convolutional neural network module (CNN), recurrent neural network module and fully connected deep neural network module (FC). Its input consists of frequency domain data, in-phase data and orthogonal data. Figure 2 The Concat in the code represents a combination operation. The following describes the components of the architecture in turn:

[0057] Convolutional neural module: Due to the relative independence of the input, the architecture first uses the first two-dimensional (2D) convolution layer (Conv1) to extract the frequency features of the discrete signal data in the time-frequency domain, and uses the first one-dimensional (1D) convolution layer (Conv2) and the second one-dimensional (1D) convolution layer (Conv3) to extract the spatial features in the discrete signal data in the time-frequency domain. In signal processing, the convolution kernel can not only highlight the high-frequency components of the signal, but also capture spatial features through convolution between consecutive frames, as well as the sliding window mechanism and edge detection capabilities. That is, Conv1 inputs the frequency domain data of the discrete signal data in the time-frequency domain, Conv2 inputs the in-phase data (I) of the discrete signal data in the time-frequency domain, and Conv3 inputs the orthogonal data (Q) of the discrete signal data in the time-frequency domain. Different signal data are separated through three channels, minimizing interference from different dimensions and ensuring maximum emphasis on key feature information.

[0058] At the same time, it should be noted that this separate convolution setting will worsen the orthogonal relationship between I and Q caused by amplitude and phase imbalance, resulting in differences between the two channels. To this end, a second two-dimensional convolution layer (Conv4) is added after Conv2 and Conv3. The in-phase output data and the orthogonal output data extracted by spatial features are fused and fed into the Conv4 to learn higher-level abstract features in the signal through joint processing. Subsequently, they are combined with the frequency domain output data of Conv1 in Concat2 based on the channel dimension to be perpendicular to the spatial correlation (see reference

[10] in the background technology for details). In the signal transmission system in the real environment, many factors such as multipath effects will be encountered, and frequency domain data can better overcome such effects in the communication transmission process, so that the signal information can be better presented. At the same time, many modulation schemes are presented differently in the frequency domain, such as many signals of high-dimensional modulation schemes (see reference [3] in the background technology for details). Therefore, the combination of time domain and frequency domain information helps to further aggregate useful information and discard confusing information from different dimensions, providing excellent feature data that can be extracted for the recurrent neural network module. Subsequently, the data is fed into the third two-dimensional convolutional layer (Conv5), which compresses the data channel dimension through the setting of the convolution kernel to fuse the spatiotemporal and frequency information, and utilizes the collaboration and complementarity between the data to aggregate useful information to complete feature capture.

[0059] Recurrent neural network module: After the convolutional neural network module (CNN algorithm) extracts the spatial and frequency dimensional features of the input signal data, the data will be fed into the GRU layer (gated recurrent unit layer) of the recurrent neural network module. The GRU layer has the ability to learn the long-term dependencies and temporal correlations of the data sequence, which helps the architecture extract useful temporal features (see reference [5] in the background technology for details). Therefore, the architecture of the present invention sets up three layers of GRU layers in the design, with 128 cells in each layer. For the modulation format signal to be identified with an original data dimension of 2×128, each cell in the first GRU layer contains 50 neurons, and the other two GRU layers contain 128 neurons each. Each GRU layer is expanded to 124 time steps. In the training phase of the space-time frequency fusion network model, back propagation is used for training to learn the temporal correlation features in the data, so that the time domain data of the unknown signal can be extracted in the inference phase.

[0060] Fully connected deep neural network module: In order to map features to a space that is easier to separate, the output data of the recurrent neural network module is fed into a fully connected deep neural network module, which consists of two FC layers (fully connected layers) with 128 neurons to simulate the approximate modulation function. Each layer includes a nonlinear activation function (ReLU) and a dropout layer to deepen the network and prevent overfitting. The final model output layer is a fully connected layer with n neurons, and the number n is determined by the type of modulation format to be identified in the time-space frequency fusion network model during pre-training.

[0061] The above describes in detail the architecture module and data time-frequency transformation principle of the present invention. After preparing the time-frequency discrete signal data received by the SISO system and the frequency domain data after FFT (Fast Fourier Transform), it is necessary to train the time-space-frequency fusion network model based on the principle of deep learning to prepare for the subsequent recognition and inference of the present invention in a real communication environment. The specific training steps are as follows:

[0062] 1. Two open-source benchmark datasets, RML2016.10a and RML2016.10b, were selected for training and subsequent performance evaluation of the proposed method. These datasets were generated using software-defined radio (SDR) to simulate realistic and harsh communication environments. The RML datasets account for time-varying random channel effects common in most wireless systems, including center frequency offset, sampling rate offset, additive white Gaussian noise, multipath, and fading. The RML2016.10a dataset contains 220,000 input samples with a signal-to-noise ratio (SNR) ranging from -20dB to +18dB in 2dB steps. The dataset includes 11 standard digital and analog modulation formats: WBFM, AM-DSB, AM-SSB, BPSK, CPFSK, GFSK, 4-PAM, 16-QAM, 64-QAM, QPSK, and 8PSK. RML2016.10b is a more extensive dataset, containing the same ten digital and analog modulation formats, excluding AM-SSB, for a total of 1,200,000 samples.

[0063] 2. Use the Mindspore framework to build the time-space-frequency fusion network model architecture, and select the UniTeng heterogeneous computing architecture (CANN) and Ascend910 NPU for training.

[0064] 3. 80% of the data in the above open source benchmark dataset is selected as the training set and validation set. The training adopts five-fold cross-validation, one fold is used as the validation set, and the remaining four folds are used as the training set, and iterates five times.

[0065] 4. For training, select the cross-entropy loss function and use Adam as the optimizer. Set the initial learning rate to 0.001 and adjust it dynamically. Set the batch size to 128. Stop training when the validation loss stops decreasing and is higher than the training loss, and the validation accuracy converges.

[0066] 5. Save the model weight with the highest verification accuracy in each fold, and select the model weight with the highest verification set accuracy in the five folds as the training result and assign it to the architecture of the present invention for subsequent modulation format recognition and inference of unknown signals.

[0067] After the training is completed and the optimal weights are obtained, the signals of WBFM, AM-DSB, AM-SSB, BPSK, CPFSK, GFSK, 4-PAM, 16-QAM, 64-QAM, QPSK and 8PSK modulation formats can be identified and detected. Since the model training process uses eleven modulation format signals generated by simulating a real environment, the present invention can identify the above eleven modulation format signals in the practical implementation link. In order to expand the recognition range of the space-time and frequency fusion network model of the present invention, the operator can generate other modulation format types of signal data according to the specific needs of the specific scenario and feed them into the space-time and frequency fusion network model of the present invention for the above training and learning. After completing the training and learning, the operator can rely on the reasoning ability of the space-time and frequency fusion network model to identify unknown signals in specific scenarios. Specifically, the inference device can select CPU, GPU, Ascend910, Ascend310, etc.

[0068] Tests have shown that, under positive signal-to-noise ratio conditions, the proposed model's recognition accuracy improves by 1% to 4% compared to existing technologies, and its training speed is up to 2.26 times faster than the current model with the highest recognition accuracy. When targeting signals with high-dimensional modulation schemes, such as 16-QAM and 64-QAM, it can better resolve confusion between similar signals, achieving even higher recognition accuracy.

[0069] In order to more intuitively demonstrate the advantages of the present invention and explain the reasons for the good technical effects, the following will use Python tool simulation, according to Figure 2 The architecture is built and the model is trained, and the test results of the above three advantages are displayed and explained in turn to verify the correctness and feasibility of the present invention.

[0070] First, after training the spatiotemporal-frequency fusion network model (MCGDNN) on two open-source benchmark datasets, the model was used to detect unknown signals on the remaining 20% test set, excluding the training and validation sets. The recognition accuracy was compared with that of six AMR models with the most advanced architectures. Figure 3 It shows that (SNR is signal-to-noise ratio, Accuracy is accuracy) based on RML2016.10a data, the recognition accuracy of the model of the present invention (MCGDNN) is significantly improved compared with the existing six benchmark high-precision models. In an environment with a signal-to-noise ratio of -6dB or above, the performance is better than other architectures in all aspects, with the highest accuracy reaching 94%. The average recognition rate from 0dB to 18dB reaches 93%, which is 1% to 4% higher than other existing models. In order to reflect the robustness and generalization performance of the present invention, the same test was also done on the RML2016.10b dataset, as shown in the figure. Figure 4 As shown in Figure 3, the recognition accuracy of MCGDNN is still better than other models.

[0071] Next, as shown in Table 1, five key metrics are listed to evaluate the complexity of the seven models: parameters, training time, optimal weight rounds, average validation accuracy, and minimum validation loss. The number of training parameters primarily assesses model complexity, while the speed of training convergence is measured by training time and optimal weight rounds. Average validation accuracy and minimum validation loss reflect the trend of training convergence.

[0072] Table 1: Comparison of model complexity based on the RML2016.10a dataset

[0073]

[0074] Since the operating mechanism of GRU is very simple, with only update gates and reset gates set up inside, the information transmission path is very short. In contrast, the operating mechanism of LSTM (Long Short-Term Memory Network) is equipped with input gates, forget gates, and output gates, and adds memory units for information transmission. Therefore, although the number of parameters of the MCGDNN model of the present invention using three-layer GRU is only slightly smaller than that of the existing MCLDNN model using two-layer LSTM, the present invention achieves a 2.26-fold improvement in training speed compared to the MCLDNN model.

[0075] The number of training parameters of the MCGDNN model of the present invention is greater than that of the GRU model and the LSTM model, but less than that of other models. Combined with the information in Table 1, it can be seen that the training time of the MCGDNN model is second only to that of the GRU model, and it has the best weight for the fastest convergence. Therefore, it is believed that the training convergence speed of the present invention is the fastest. Although the GRU model has fewer parameters and usually requires only a small amount of training time, its simple mechanism and lack of memory units make it more difficult to learn the temporal structure features between consecutive time steps than the LSTM model. Therefore, a simple GRU model can usually only maintain poor recognition accuracy. In order to solve the above problems, the invented MCGDNN architecture first follows the method of reference

[10] in the background technology, and uses the CNN algorithm to model higher-level data in the early stage of the architecture to clarify the potential change factors in the input. In addition, the frequency domain data is fed into the model as input data, which enriches the information content and enhances the robustness. In the middle stage of the architecture, the CNN algorithm combines time, space and frequency information, which can more comprehensively understand the nature of the signal, improve the model's ability to understand complex signals, and thus improve the recognition accuracy.

[0076] The present invention combines the characteristics of frequency domain information, CNN algorithm and GRU algorithm, and makes use of their respective advantages to play a role while complementing each other's strengths and weaknesses. The three complement each other to fill their respective shortcomings, that is, relying on the advantages of GRU to achieve the goal of fewer overall architecture parameters and the fastest training convergence speed, while ensuring that the model has better recognition accuracy. The information in the table shows that the present invention has the highest average verification accuracy and the smallest minimum verification loss. Figure 3 and Figure 4 The information also proves that the model of the present invention has better recognition accuracy than other models.

[0077] Finally, to facilitate an intuitive understanding of the third advantage of the present invention, Table 2 shows the recognition accuracy of eleven modulation formats in a 0dB signal-to-noise ratio environment. MCGDNN significantly improves the recognition accuracy of 16-QAM, 64-QAM, and 8PSK modulation types, outperforming other models. It effectively resolves the signal confusion problem between 16-QAM and 64-QAM.

[0078] Table 2: Recognition accuracy based on the RML2016.10a dataset at 0dB signal-to-noise ratio

[0079] MCGDNN MCLDNN GRU LSTM ResNet18 ResNet50 CLDNN2 16-QAM 96.0% 92.7% 73.4% 85.3% 75.8% 83.7% 92.7% 64-QAM 96.6% 95.1% 83.9% 90.1% 80.0% 88.3% 85.9% 8PSK 96.1% 90.0% 86.7% 92.2% 81.1% 92.2% 92.2% WBFM 32.1% 37.7% 46.7% 36.8% 36.3% 25.9% 42.0% BPSK 97.5% 97.9% 97.0% 97.5% 96.5% 97.5% 97.5% CPFSK 99.5% 99.1% 99.1% 100% 99.5% 99.5% 100% AM-DSB 98.5% 96.4% 84.2% 93.4% 97.4% 96.4% 93.9% GFSK 98.9% 95.8% 97.4% 95.8% 94.2% 96.8% 94.2% PAM4 98.2% 98.8% 98.8% 98.8% 98.3% 98.3% 98.3% QPSK 97.0% 96.5% 93.6% 85.1% 89.1% 96.5% 98.5% AM-SSB 94.1% 92.6% 94.6% 94.6% 90.1% 93.1% 94.1%

[0080] To further understand the third advantage of the present invention: "The present invention has a better solution to the confusion problem of signals with high-dimensional modulation schemes such as 16-QAM, 64-QAM, etc., and achieves higher recognition accuracy." Figure 5 and Figure 6 As shown in Figure 2, by using t-SNE technology to perform 2D visualization of the test data input to the MCGDNN model, we can gain a deeper understanding of the recognition process of the MCGDNN model. Figure 5 In the

[15] , discrete data are spatially clustered by the CNN algorithm, and the frequency domain features (Conv4 to Conv5) are added to aggregate the discrete data of the same modulation type, making the data of different modulation types appear hierarchical and centralized. Subsequently, the GRU layer separates different types of data. Figure 5 The independent contributions of space, time and frequency are demonstrated, and the key roles of spatial feature extraction, frequency domain feature extraction and time modeling are emphasized. Figure 6 In order to have a more intuitive understanding of the role of frequency domain information in the present invention, a model named MCGDNN-A is designed, which removes the frequency domain data in MCGDNN and replaces it with time domain I / Q (in-phase / quadrature) data to feed it into Conv1. Figure 6The information shows that compared with time-domain features, MCGDNN is better at resolving signal confusion between 16-QAM and 64-QAM due to the introduction of frequency-domain features, thereby improving the recognition accuracy of high-dimensional modulated signals. Therefore, the present invention can extract features from received data in three dimensions: space, time, and frequency, and realize the recognition of the modulation type of unknown electromagnetic spectrum signals.

[0081] The above-described embodiments merely represent specific implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of protection of the present application. It should be noted that those skilled in the art will be able to make relevant modifications and improvements without departing from the technical concept of the present application, and these modifications and improvements fall within the scope of protection of the present application.

Claims

1. The automatic modulation recognition method based on three channels is characterized by: When a single-input single-output communication system receives an unknown signal, it first performs time-domain discrete sampling on the in-phase / orthogonal time-domain signals to form a time-domain discrete signal, then combines the time-domain discrete signal into one-dimensional complex data, performs Fourier time-frequency domain transform to obtain frequency domain data, combines the frequency domain data with the time-domain discrete signal to obtain time-frequency domain discrete signal data, separates the in-phase data and the orthogonal data of the time domain data in the time-frequency domain discrete signal data, then feeds the frequency data in the time-frequency domain discrete signal data into the two-dimensional convolution layer of the time-space-frequency fusion network model, and feeds the in-phase data and the orthogonal data into the corresponding one-dimensional convolution layer respectively. The time-space-frequency fusion network model infers and determines the type of modulation format of the received unknown signal based on the unique characteristics of the various modulation formats that have been learned; The space-time frequency fusion network model includes a convolutional neural network module, a recurrent neural network module and a fully connected deep neural network module, and the input consists of the frequency domain data, the in-phase data and the orthogonal data; In the convolutional neural module, the first two-dimensional convolutional layer is used to extract the frequency characteristics of the time-frequency domain discrete signal data, and the first one-dimensional convolutional layer and the second one-dimensional convolutional layer are used to extract the spatial characteristics of the time-frequency domain discrete signal data through convolution operations between consecutive frames, as well as the sliding window mechanism and edge detection capability of the convolution kernel. That is, the first two-dimensional convolutional layer inputs the frequency domain data in the time-frequency domain discrete signal data, the first one-dimensional convolutional layer inputs the in-phase data in the time-frequency domain discrete signal data, and the second one-dimensional convolutional layer inputs the orthogonal data in the time-frequency domain discrete signal data, forming a three-channel data input; After the convolutional neural module extracts spatial and frequency dimensional features from the input signal data, the data is fed into the gated recurrent unit layer of the recurrent neural network module. The gated recurrent unit layer has three layers and is used to learn the long-term dependencies and temporal correlations of the data sequence and extract the time domain features of the unknown signal. After the first one-dimensional convolution layer and the second one-dimensional convolution layer, there is also a second two-dimensional convolution layer, which fuses the in-phase output data and the orthogonal output data after spatial feature extraction and feeds them into the second two-dimensional convolution layer to learn higher-level abstract features in the signal, and then combines the higher-level abstract features with the frequency domain output data of the first two-dimensional convolution layer based on the channel dimension to be perpendicular to the spatial correlation; The output data of the second two-dimensional convolutional layer is fed into the third two-dimensional convolutional layer. The data channel dimension is compressed by setting the convolution kernel to fuse the spatiotemporal and frequency information, and the collaboration and complementarity between the data are used to aggregate useful information to complete the feature capture in the data.

2. The three-channel automatic modulation recognition method according to claim 1, wherein: The Fourier time-frequency domain transform adopts discrete Fourier time-frequency domain transform.

3. The three-channel automatic modulation recognition method according to claim 2, wherein: The Fourier time-frequency domain transform adopts fast Fourier time-frequency domain transform.

4. The three-channel automatic modulation recognition method according to claim 1, wherein: The output data of the recurrent neural network module is fed into the fully connected deep neural network module to map the features in the data to a space that is easier to separate. The fully connected layer of the fully connected deep neural network module includes a nonlinear activation function and a dropout layer. The final model output layer is a fully connected layer with n neurons, and the number n is determined by the type of modulation format to be identified input into the space-time frequency fusion network model during pre-training.

Citation Information

Patent Citations

  • Modulation signal identification method based on multi-domain mixed attention

    CN116471154A

Cited By

  • Modulation recognition method and system for missing signal based on time-frequency fusion collaborative recovery

    CN122554272A