Multi-mode radio signal modulation identification method based on deep learning

By constructing a three-stream fusion signal recognition model and combining deep learning technology to extract and fuse multimodal features of radio signals, the problem of low recognition accuracy in the existing technology is solved, and more efficient and accurate modulated signal recognition is achieved.

CN120075005AInactive Publication Date: 2025-05-30HANGZHOU DIANZI UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510220729.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-05-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art is difficult to effectively utilize the multimodal features of the signal in radio signal modulation recognition, resulting in low recognition accuracy, and the features designed by traditional methods are difficult to be widely applicable to various modulation methods.

Method used

Using a multimodal radio signal modulation recognition method based on deep learning, a three-stream fusion signal recognition model is constructed, combined with a point-by-point convolution network, attention module and a bidirectional gated cyclic unit network, a multimodal feature of the signal is extracted and fused, and finally classified through a full connection layer.

Benefits of technology

It improves the recognition accuracy of modulated signal patterns, enhances the diversity of features and recognition performance, reduces model parameters and calculation complexity, and achieves faster and more accurate signal recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120075005A_ABST
    Figure CN120075005A_ABST
Patent Text Reader

Abstract

The invention relates to the field of radio signal modulation identification, and particularly discloses a multi-mode radio signal modulation identification method based on deep learning, which comprises the following steps of: 1, preparing a data set which comprises 48 digital modulation signals of BPSK (Binary Phase Shift Keying), QPSK (Quadrature Phase Shift Keying), 8PSK (8PSK), 16QAM (Quadrature Amplitude Modulation), 64QAM, BFSK (Binary Frequency Shift Keying), CPFSK (Carrier Phase Frequency Shift Keying) and PAM (Pulse Amplitude Modulation) and 3 analog modulation signals of WB-FM (Wave Binary Frequency Modulation), AM-SSB (Amplitude Modulation-Frequency Shift Keying) and AM- Step 2, data preprocessing: converting I / Q signals into real part and imaginary part representations of A / P and 2D-FFT to form three data sets; step 3, constructing a three-stream fusion signal identification network model; 4, training the model, and inputting the training set and the verification set into the three-stream fusion signal identification network model for training; and 5, inputting the test set into the trained three-stream fusion signal identification model, automatically identifying the type of the modulation signal, and outputting the signal identification accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of radio signal modulation recognition, and more specifically to a multi-modal radio signal modulation recognition method, providing a technical option for the best modulation signal recognition solution in the communication industry. Background Art

[0002] With the rapid development of communication technology, the diversification of modulation methods, and the further improvement of the complexity of communication channels, the requirements for automatic modulation recognition are also getting higher and higher. Automatic Modulation Recognition (AMR) technology plays a key role in communication reconnaissance and cognitive electronic warfare. Efficient and accurate signal analysis and recognition are crucial for backend processing. Therefore, automatic modulation recognition technology remains a key research technology in the field of communication signal processing.

[0003] The purpose of the radio modulation classification problem is to identify the modulation type to which the signal belongs. This is actually a multi-class classification problem, and the number of classes is the number of modulation types in the signal dataset. Traditional modulation classification methods mainly rely on feature design, and the quality of the designed features directly determines the performance of recognition. However, these designed features are often related to specific modulations, and it is difficult to find features that are widely applicable to various modulations. As an end-to-end learning method, deep learning unifies the feature extraction and recognition tasks, avoiding the process of manually designing features, thereby greatly enhancing its general applicability.

[0004] In recent years, AMR can be mainly divided into two categories. The first category is to convert radio signals into images and then classify the images with reference to the network structure used in image classification to achieve the classification of radio signals. Currently, there are mainly two methods to convert in-phase quadrature (IQ) signals into images, namely constellation diagrams and time-frequency images. However, they may affect the modulation classification performance due to the loss of time-related information between signal sampling points. At the same time, the constellation diagram features of the signal are classified. The above research has made certain progress in the field of signal modulation classification, but most of them only focus on certain aspects of the signal features, ignoring other aspects of the features and not considering the multi-modal features of the signal. Summary of the Invention

[0005] In view of the above deficiencies in the prior art, the present invention provides a multi-modal three-stream fusion radio signal modulation mode recognition method (CAG), which increases the diversity of features, reduces the resource consumption rate, uses a pointwise convolutional network (PWConv) to reduce model parameters and computational complexity; performs feature fusion through an attention module (Co-Attention); and uses a bidirectional gated recurrent unit (BiGRU) network to extract high-level features of the signal for classification. The method of the present invention can identify the modulation signal mode more quickly and accurately, with a higher recognition accuracy. The specific technical solutions are as follows:

[0006] A multi-modal radio signal modulation recognition method based on deep learning, comprising the following steps:

[0007] Step 1: Prepare the data set;

[0008] The data set RML2016.10a used in the present invention is a data set for radio signal modulation recognition, which contains signal samples of various digital and analog modulation methods. These samples are generated by software-defined radio (SDR) technology, including training data and test data, and their types include 8 digital modulations such as BPSK, QPSK, 8PSK, 16QAM, 64QAM, BFSK, PAM4K, and CPFS, and 3 analog modulation signals such as WB-FM, AM-SSB, and AM-DSB, with a total of 220,000 samples, divided into two channels of I / Q data, and the length of each sample is 128. The signal-to-noise ratio is -20 dB to 18 dB, with an interval of 2 dB, that is, 20 different signal-to-noise ratio environments.

[0009] Step 2: Data preprocessing, converting the signal into time in-phase / quadrature (I / Q), amplitude / phase (A / P), and the real and imaginary parts of the Fourier transform (2D-FFT) representations to form three data sets;

[0010] (2.1) Construct the I / Q vector data set Obtained from the original complex signal, consisting of two parts, the in-phase component X I and the quadrature component X Q , and converted into an I / Q dual-channel vector:

[0011]

[0012] where X I , X Q ∈R N ,

[0013] In-phase component:

[0014]

[0015] Orthogonal component:

[0016]

[0017] Let \(j\) represent the number of samples in the data set, and \(N\) represent the number of sampling points.

[0018] (2.2) Construct the A / P data set It is composed of the amplitude component \(X\) A and the phase component \(X\) P and is converted into an A / P two-channel vector:

[0019]

[0020] The signal amplitude and phase vectors are calculated as follows:

[0021]

[0022] where \(X\) A , \(X\) P \(\in R\) N ,

[0023] Amplitude component:

[0024] \(j\)

[0025] represents the number of samples in the data set, represents the maximum amplitude in the data set;

[0026] Phase vector:

[0027] Let \(j\) represent the number of samples in the data set, and \(N\) represent the number of sampling points.

[0028] (2.3) The 2D-FFT vector is a mapping from the time domain to the frequency domain. The data set \(X\) FFT is composed of the real component \(X\) FFT-R of the complex number of the real-valued data vector and the imaginary component \(X\) FFT-I of the complex number:

[0029]

[0030] where \(R\{FFT[r(k)]\}\), \(J\{FFT[r(k)]\}\in R\) N , \(X\) FFT \(\in R\) 2×N

[0031] \(r\) n is the complex signal representation of the IQ channel:

[0032]

[0033] For the complex sequence r n Perform a discrete Fourier transform, which can be denoted as:

[0034]

[0035] Since the neural network model does not support complex number operations, for the obtained spectral signal X FFT (k), divide it into two channels of real part and imaginary part as the input of the neural network.

[0036]

[0037] Construct a 2D-FFT dataset, which consists of real part components and imaginary part components:

[0038] Where

[0039] j represents the number of dataset samples, and N represents the number of sampling points.

[0040] Step 3: Build the network structure of the three-stream fusion signal recognition model and set the basic network parameters;

[0041] (3.1) Based on the above signal processing, select the original IQ signal and its amplitude, phase, and frequency to form 3 modalities as the input of the network model, allowing the model to learn the feature information of the original modulation signal from different angles. Since signal modulation is essentially a process of transforming the amplitude, phase, and frequency of a communication signal according to certain rules, it is theoretically applicable to various modulation methods. In addition, the features learned from multiple modalities interact with each other, increasing the diversity of features and being more conducive to feature classification.

[0042] (3.2) The specific three-stream fusion signal recognition model includes multiple convolutional modules, attention fusion modules, BiGRU modules, and classification modules. Different feature extraction modules are used in the three streams to extract features from three different datasets. Then, Co-Attention is used to fuse the features of the three streams. The fused sequence extracts the temporal features of the signal through the BiGRU network, and finally, a fully connected layer is used for classification. The convolutional network structure of each stream is used to extract the spatial characteristics of the data. In the attention fusion module, the signal features extracted from the I / Q stream are regarded as key-value pair sequences, which are represented by K = {k 1 , k 2 ,..., k n} and V = {v 1 , v 2 ,..., v n} respectively. The other two streams are regarded as query sequences and represented by Q = {q 1 , q 2 ,..., q n}(denoted as). According to the attention mechanism structure, the co-attention between the two modalities is calculated as follows:

[0043]

[0044] The bidirectional gated recurrent unit network extracts the temporal characteristics of the fused features, and the output of the hidden layer is H = {h 1 , h 2 ,..., h T}}. The dropout method is used with a dropout rate of p = 0.5 to avoid overfitting; next, the output of the hidden layer is fed into the Dense layer to achieve dimensionality reduction and generate an 11-dimensional matrix, and finally, the maximum value of the predicted probability of the modulation signal is obtained through the argmax operation.

[0045] (3.3) The fusion of the three-stream network uses a convolutional neural network (CNN) to extract the spatial features of the signal and uses a joint attention mechanism to achieve feature fusion to enhance performance. The fused features use BiGRU to extract temporal features, which can extract more details and enrich the diversity of attributes compared to ordinary weighted fusion, thus further enhancing the performance.

[0046] Step 4: Train the model. Input the training set and the validation set into the three-stream fusion signal recognition model network for training. GeLu is used as the activation function for the convolution in the network structure, ReLU is used as the activation function for each layer of BiGRU, and the categorical cross-entropy error is used as the loss function:

[0047]

[0048] where yi represents the predicted value and N represents the number of training samples; the network model uses the Adam stochastic gradient descent algorithm as the optimizer, where the learning rate μ is set to 0.001, the loss function is minimized, and the network parameters are iteratively updated;

[0049] Step 5: Input the test set into the trained three-stream fusion signal recognition model to automatically identify the type of the modulation signal and output the signal recognition accuracy rate.

[0050] The advantages of the present invention compared with the existing inventions are:

[0051] 1. The three-stream fusion signal recognition model adopted by the present invention utilizes three different features of the signal to automatically extract the temporal and spatial characteristics of the modulation signal, uses the Co-Attention module to fuse the features of different streams, and uses PWConv to reduce the computational amount, accelerating the calculation process and maximizing the retention of more feature information;

[0052] 2. The present invention adopts a three-stream fusion signal recognition model network structure. By inputting multiple features of the signal, it enriches the data representation form of each modulation method, realizes the complementarity between different types of data features, and preferably solves the problem of signal modulation recognition. At the same time, the idea of multi-feature extraction can better extract and fuse the internal features of the signal, thereby improving the recognition accuracy of the modulation signal. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the embodiments will be briefly introduced below. The following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts. The proportional relationships of the components in the accompanying drawings of this specification do not represent the proportional relationships in actual material selection and design, where:

[0054] Figure 1 is the network structure diagram of the three-stream fusion signal recognition model;

[0055] Figure 2 is the network structure diagram of the bidirectional gated recurrent unit;

[0056] Figure 3 is the structure diagram of the gated recurrent unit;

[0057] Figure 4 is the recognition accuracy curve of the comparison model;

[0058] Figure 5 is the recognition accuracy curve graph of each type of modulation signal;

[0059] Figure 6 is the confusion matrix graph of the recognition result when the signal-to-noise ratio of the signal is 0 dB;

[0060] Figure 7 is the confusion matrix graph of the recognition result when the signal-to-noise ratio of the signal is 6 dB. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0061] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings: These embodiments are implemented on the premise of the technical solutions of the present invention, and detailed implementation methods and specific operation processes are given, so that the advantages and features of the present invention can be more simply and quickly understood by those skilled in the art. However, the protection scope of the present invention is not limited to the following embodiments.

[0062] The multi-modal radio signal modulation recognition method based on deep learning of the present invention includes the following steps:

[0063] Step 1: Prepare the data set;

[0064] The dataset used in this invention is the open-source modulation dataset RadioML2016.10a generated by GNU Radio, which contains 8 types of digital modulations including BPSK, QPSK, 8PSK, 16QAM, 64QAM, BFSK, CPFSK, and PAM4, and 3 types of analog modulation signals including WB-FM, AM-SSB, and AM-DSB. The training set and test set are divided in the ratio of 8:2.

[0065] Step 2: Data preprocessing. Convert the in-phase / quadrature (I / Q) signals into amplitude / phase (A / P) and the real and imaginary parts (R / J) of the 2D-FFT to form three datasets.

[0066] The specific steps are as follows:

[0067] Step 2.1: The modulation signal is an I / Q signal, which consists of two groups of sequences, the I-channel and the Q-channel. First, convert it into an I / Q vector:

[0068]

[0069] where

[0070] j represents the number of dataset samples, and N represents the number of sampling points.

[0071] Step 2.2: Calculate the amplitude and phase vectors:

[0072]

[0073] Construct the A / P dataset. The amplitude component and the phase component are combined and converted into an A / P two-channel vector:

[0074]

[0075] To reduce the range difference of the values in the data and improve the numerical stability, normalize the amplitude and phase respectively:

[0076]

[0077] represents the maximum amplitude value in the dataset;

[0078] j represents the number of dataset samples, and N represents the number of sampling points;

[0079] Step 2.3: The 2D-FFT vector is a mapping from the time domain to the frequency domain, which consists of two groups of real-valued data vectors to form the real component and the imaginary component of the complex number:

[0080]

[0081] r n Complex signal representation for the IQ channel:

[0082]

[0083] For the complex sequence r n Perform a discrete Fourier transform, which can be denoted as:

[0084]

[0085] Since the neural network model does not support complex number operations, for the obtained spectral signal X FFT (k), we divide it into two channels, the real part and the imaginary part, as the input of the neural network.

[0086]

[0087] Construct a 2D-FFT data set, which consists of real and imaginary components and is converted into an R / J vector:

[0088]

[0089] j represents the number of data set samples, and N represents the number of sampling points;

[0090] Step 3: Construct a three-stream fusion signal recognition (CAG) network model

[0091] As Figure 1 shown, the constructed lightweight convolutional attention recurrent neural network mainly includes four modules, namely the convolutional module, the attention fusion module, the double-layer GRU network module, and the classifier. Among them, the convolutional module is used to extract the spatial domain features of the input data; the attention fusion module is used to fuse the extracted three-stream signal features together; the double-layer GRU network module is used to extract the one-dimensional time characteristics of the signal, and finally classification is performed through the classifier.

[0092] Construct the convolutional module:

[0093] Specifically, each stream of the CAG model uses convolutional modules of different scales. The I / Q channel and the 2D-FFT channel use a combination of partial convolution (PConv) and pointwise convolution (PWConv) to extract the spatial features of the signal, and the A / P channel uses two-dimensional convolution to extract the spatial characteristics of the signal.

[0094] The I / Q and 2D-FFT channel convolution module contains two combined blocks of PConv and PWConv; the I / Q signal and 2D-FFT signal are processed into a size of 1*2*128 and input into the convolutional layer. Among them, some convolutional layers only perform convolution on the first quarter of the channels, with a convolutional kernel size of 3*3. After convolution, the convolved channels are connected to the non-convolved channels to form the input of the next layer; the pointwise convolutional kernel size is 1*1. This convolutional layer effectively extracts information from all channels and significantly reduces the number of convolutional parameters; to avoid the situation of gradient explosion and gradient disappearance during training, in the present invention, after each pointwise convolution, a batch normalization layer (Batch Normalization, BN) is used to accelerate the training and convergence speed of the network; the activation function GeLu is selected to enhance the nonlinearity of the network, prevent gradient disappearance, reduce overfitting and improve the training speed of the network; after passing through the above convolutional layer, feature information of 32*32*2 is output.

[0095] The A / P channel convolution module contains two two-dimensional convolutional layers; the A / P signal is processed into a size of 1*2*128 and input into the two-dimensional convolutional layer. The first convolutional layer has a convolutional kernel size of 1*3 and a filter size of 16, with one unit padded on each side in the width direction and no padding in the height direction. The second convolutional layer has a size of 1*3 and a filter size of 32, with one unit padded on each side in the width direction and no padding in the height direction. To further eliminate the interference of useless information and extract key information, max pooling is performed on the data after convolution, with a stride of 2 in the width direction and a convolutional kernel size of 1*2; at the same time, the activation function Relu is selected to enhance the nonlinearity of the network, prevent gradient disappearance, reduce overfitting and improve the training speed of the network; after passing through the above convolutional layer, feature information of 32*2*32 is output.

[0096] Construct an attention fusion module:

[0097] First, a multi-head attention unit is added behind each channel to learn the importance of each channel, and the attention mechanism is used to further enhance the features useful for modulation recognition and suppress the features useless for modulation recognition. After the multi-head attention features are output from each stream, Co-Attention is used for feature fusion. In the feature fusion module, the signal features extracted from the I / Q stream are regarded as key-value pair sequences, K = {k 1 ,k 2 ,...,k n} and V = {v 1 ,v 2 ,...,v n} are represented, and the other two streams are regarded as query sequences, represented by Q = {q 1 ,q 2 ,...,q n}\ It is represented as follows. According to the attention mechanism structure, the co-attention among the three modalities is calculated as follows:

[0098]

[0099] Finally, the I / Q stream is connected with the output of the multi-head attention and the output of the Co-Attention to obtain the input of the gated array unit.

[0100] Construct a two-layer GRU network:

[0101] The BiGRU model is a type of recurrent neural network as Figure 2 shown, which consists of two independent GRU units. One processes data forward in time series, and the other processes data backward in time series. Through this bidirectional structure, the BiGRU model can capture both the forward and backward information of the sequence data, thus better understanding and predicting the patterns in the sequence. For the two-layer BiGRU network we used, the first-layer BiGRU units are responsible for capturing the basic sequence features, and the role of the BiGRU in the second layer is to capture more complex dependencies and higher-level abstract features, which is beneficial for the model to better understand the long-term dependencies and more complex semantic relationships in the sequence data. As the size of the BiGRU hidden units increases, the classification accuracy increases; when the size of the hidden units doubles, the recognition accuracy will improve. However, as the size of the hidden units increases, the training time also increases. Therefore, 32 hidden units are selected. At this time, 64 eigenvalue outputs are generated at each time step. After flattening all the eigenvalues, they can be directly sent to the classification layer for classification.

[0102] There are two key components in the GRU network: the update gate and the reset gate. The role of the update gate is to determine how much of the previous moment's state information will be retained and passed to the current moment, while the reset gate controls how the new input information is fused with the memories of the previous time steps. When the value of the reset gate is set to 1 and the value of the update gate is set to 0, such a configuration actually degenerates the GRU model into a traditional RNN model. By adjusting these two gating vectors, GRU can effectively filter out the information finally used for output. Its uniqueness lies in that this gating mechanism can effectively retain the information in the long time series, preventing this information from being forgotten over time or discarded because it is considered irrelevant to the prediction. The GRU gating structure is as Figure 3 shown, and the calculation process is as follows:

[0103] Reset gate:

[0104] r t = σ(X t * W xr + h t-1 * Whr +b r ) (10)

[0105] Among them, X t represents the current input, h t-1 represents the hidden state of the previous time step, W xr , W hr and b r are learnable weight parameters, and σ is the sigmoid function. r t represents the output of the reset gate.

[0106] Update gate:

[0107] In the GRU architecture, an update gate is calculated at each time step. This process is implemented through a sigmoid activation function to ensure that the output value is between 0 and 1. When the output value of the update gate is close to 1, it means that the network tends to fully retain the information of the previous state; conversely, if the output value is close to 0, it indicates that the network hardly considers the past state information and mainly processes based on the current input data. In short, the update gate enables the GRU to flexibly determine the degree of retention of the previous state information, thus achieving selective memory or forgetting of historical information.

[0108] z t = σ(X t *W xz + h t-1 *W hz + b z ) (11)

[0109] Among them, X t represents the current input, h t-1 represents the hidden state of the previous time step, W xz , W hz and b z are learnable weight parameters, and σ is the sigmoid function. z t represents the output of the update gate.

[0110] New candidate state:

[0111]

[0112] Among them, X t represents the current input, h t-1 represents the hidden state of the previous time step, W h and b h are learnable weight parameters, r t represents the output of the reset gate. represents the new candidate state.

[0113] Update the hidden state:

[0114]

[0115] Among them, h t-1 represents the hidden state of the previous time step, z t represents the output of the update gate, represents the new candidate state.

[0116] h t is the new hidden state

[0117] Construct a classifier

[0118] Flatten the feature output of each time step and input it into the Dense layer to achieve dimensionality reduction and generate an 11-dimensional matrix. Finally, obtain the maximum value of the predicted probability of the modulation signal through the argmax operation.

[0119] Step 4: Train the model. Input the training set and the validation set into the three-stream fusion signal recognition model network for training. GeLu is used as the activation function for the convolution in the network structure, ReLU is used as the activation function for each layer of BiGRU, and the categorical cross-entropy error is used as the loss function:

[0120]

[0121] Among them, yi represents the predicted value, and N represents the number of training samples; the network model uses the Adam algorithm of stochastic gradient descent as the optimizer, where the learning rate μ is set to 0.001, the loss function is minimized, and the network parameters are iteratively updated;

[0122] Step 5: Input the test set into the trained three-stream fusion signal recognition model to automatically identify the type of the modulation signal and output the signal recognition accuracy rate.

[0123] Result analysis:

[0124] Figure 4 Shows the experimental results and recognition curves of six models on RadioML2016.10a. At each signal-to-noise ratio, the average recognition accuracy rate of the CAG model proposed by the present invention is 63.1%. Compared with other algorithms, the classification accuracy has been greatly improved. The average detection accuracy rate in the high signal-to-noise ratio case exceeds 90%. Compared with CLDNN in this field, the signal detection accuracy rate is 7.6% higher, and compared with GRU, it is 4.7% higher.

[0125] The recognition accuracy rates of the CAG model for various signals in RadioML2016.10a are as Figure 5 shown. As can be seen from Figure 5It can be seen that the CAG model has the best recognition rate for PAM4, and the WBFM has the worst recognition rate. Among them, the recognition rates of QPSK, BPSK, CPFSK, GFSK, QAM16, QAM64, AM-DSB, and 8PSK modulation signals are all above 90% in the SNR environment above 0 dB; the recognition rate of the AM-SSB modulation signal is about 85% in the SNR environment above 0 dB;

[0126] To further analyze the recognition performance of each modulation signal, a confusion matrix is drawn based on the recognition results of the CAG model. When the SNR is -6 dB, from Figure 6 it can be seen that there are still large classification misjudgments for other modulation methods except PAM4, AM-SSB, QAM16, GFSK, and QAM64; when the SNR reaches 0 dB, from Figure 7 it can be seen that most modulation methods can be effectively recognized, but the WBFM modulation method is still easily misjudged as the AM-DSB modulation method.

[0127] The above are only the embodiments of the present invention, and do not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be included in the patent protection scope of the present invention by the same token.

Claims

1. A multimodal radio signal modulation recognition method based on deep learning, characterized in that The following steps are involved: Step 1. Prepare the data set. Use the open source modulation data set RadioML2016.10a generated by GNU Radio, which contains 48 digital modulations including BPSK, QPSK, 8PSK, 16QAM, 64QAM, BFSK, CPFSK and PAM, and 3 analog modulation signals including WB-FM, AM-SSB and AM-DSB. Divide the training set and test set into 8:

2. Step 2. Data preprocessing: converting the I / Q signals into real and imaginary representations of A / P and 2D-FFT to form three data sets; Step 3. Construct a three-stream fusion signal recognition network model; Step 4. Training the model, inputting the training set and the validation set into the three-stream fusion signal recognition network model for training; Step 5: Input the test set into the trained three-stream fusion signal recognition model, automatically identify the type of modulation signal, and output the signal recognition accuracy.

2. The multimodal radio signal modulation recognition method based on deep learning according to claim 1, characterized in that: Step 2 The specific steps are as follows: Step 2.1: The I / Q signal consists of two sets of sequences, I and Q, which are first converted into I / Q vectors: in j represents the number of samples in the data set, and N represents the number of sampling points; Step 2.2, calculate the amplitude and phase vector: Construct an A / P data set, which consists of amplitude and phase components and is converted into an A / P dual-channel vector: Normalize the amplitude and phase separately: Indicates the maximum amplitude in the data set; represents the number of samples in the data set, and N represents the number of sampling points; Step 2.3, 2D-FFT vector consists of two sets of real-valued data vectors, the real component of the complex number and the imaginary component of the complex number: r n The complex signal representation of the IQ channel is: For complex sequence r n Performing discrete Fourier transform, it can be written as: The obtained spectrum signal X FFT (k) It is divided into two channels, real part and imaginary part, as the input of the neural network: Construct a 2D-FFT data set consisting of real and imaginary components and convert them into R / J vectors: represents the number of samples in the data set, and N represents the number of sampling points.

3. The multimodal radio signal modulation recognition method based on deep learning according to claim 1, characterized in that: The three-stream fusion signal recognition CAG network model constructed in step 3 includes a convolution module, an attention fusion module, a double-layer GRU network module and a classifier. The convolution neural network module is used to extract the spatial domain features of the input data; the attention fusion module is used to fuse the extracted three-stream signal features; The two-layer GRU network module is used to extract the one-dimensional temporal characteristics of the signal, and the classifier is used for classification.

4. The multimodal radio signal modulation recognition method based on deep learning as claimed in claim 3, characterized in that: The convolution module is constructed as follows: The I / Q channel and 2D-FFT channel convolution modules use a combination of partial convolution PConv and point-by-point convolution PWConv to extract the spatial features of the signal, and the A / P channel convolution module uses two-dimensional convolution to extract the spatial characteristics of the signal; The I / Q and 2D-FFT channel convolution modules include two PConv and PWConv combination blocks; the I / Q signal and 2D-FFT signal are processed into a size of 1*2*128 and input into the convolution layer, where some convolution layers only perform convolution on the first quarter of the channel, and the convolution kernel size is 3*3. After the convolution, the convolved channel and the unconvolved channel are connected to form the input of the next layer; the convolution kernel size of the point-by-point convolution is 1*1; after each point-by-point convolution, a batch standard layer BN is used to speed up the training and convergence of the network; GeLu is selected as the activation function; after passing through the above convolution layer, 32*32*2 feature information is output; The A / P channel convolution module includes two two-dimensional convolution layers; the A / P signal is processed into a size of 1*2*128 and input into the two-dimensional convolution layer. The convolution kernel size of the first convolution layer is 1*3, the filter size is 16, one unit is padded on each side in the width direction, and no padding is performed in the height direction. The size of the second convolution layer is 1*3, the filter size is 32, one unit is padded on each side in the width direction, and no padding is performed in the height direction; after convolution, the data is subjected to maximum pooling processing, with a step size of 2 in the width direction and a convolution kernel size of 1*2; at the same time, Relu is added as an activation function; after passing through the above convolution layers, 32*2*32 feature information is output.

5. The multimodal radio signal modulation recognition method based on deep learning as claimed in claim 3, characterized in that: The attention fusion module is constructed as follows: First, a multi-head attention unit is added after each channel to learn the importance of each channel. After each stream outputs the multi-head attention feature, Co-Attention is used for feature fusion. In the feature fusion module, the signal features extracted from the I / Q stream are regarded as a key-value pair sequence, K = {k1, k2, ..., k n } and V = {v1,v2,...,v n }, and the other two streams are regarded as query sequences and denoted by Q = {q1, q2, ..., q n } represents; according to the attention mechanism structure, the common attention between the three modalities is calculated as follows: Finally, the I / Q flow is connected through the output of multi-head attention and the output of Co-Attention to obtain the input of the gated array unit.

6. The multimodal radio signal modulation recognition method based on deep learning as claimed in claim 3, characterized in that: The two-layer GRU network is constructed as follows: The two-layer GRU network consists of two independent GRU units, one of which processes data forward in time series, and the other processes data in reverse time series; the first-layer BiGRU unit captures basic sequence features, and the second-layer BiGRU unit captures more complex dependencies and more advanced abstract features.

7. The multimodal radio signal modulation recognition method based on deep learning according to claim 6, characterized in that: The GRU network has two key components: the update gate and the reset gate; Reset the gate as follows: r t =σ(X t *W xr +h t-1 *W hr +b r ) (10) Among them, X t Indicates the current input, h t-1 represents the hidden state of the previous time step, W xr , W hr and b r is a learnable weight parameter, and σ is a sigmoid function. t represents the output of the reset gate; Update the gate as follows: z t =σ(X t *W xz +h t-1 *W hz +b z ) (11) Among them, X t Indicates the current input, h t-1 represents the hidden state of the previous time step, W xz , W hz and b z is a learnable weight parameter, and σ is a sigmoid function. t represents the output of the update gate; New candidate status: Among them, X t Indicates the current input, h t-1 represents the hidden state of the previous time step, W h and b h is a learnable weight parameter, r t Represents the output of the reset gate. Indicates a new candidate state; Update hidden state: Among them, h t-1 represents the hidden state of the previous time step, z t represents the output of the update gate, Indicates the new candidate state; h t is the new hidden state.

8. The multimodal radio signal modulation recognition method based on deep learning as claimed in claim 3, characterized in that: The classifier is constructed as follows: The feature output of each time step is flattened and passed into the Dense layer to achieve dimensionality reduction and generate an 11-dimensional matrix. Finally, the maximum value of the predicted probability of the modulated signal is obtained through the argmax operation.

9. The multimodal radio signal modulation recognition method based on deep learning according to claim 1, characterized in that: In step 4, GeLu is used as the activation function, each layer of BiGRU uses ReLU as the activation function, and the classification cross entropy error is used as the loss function: Where yi represents the predicted value, N represents the number of training samples; the network model uses the stochastic gradient descent algorithm Adam as the optimizer, where the learning rate μ is set to 0.001, the loss function is set to be minimized, and the network parameters are iteratively updated. Step 5: Input the test set into the trained three-stream fusion signal recognition model, automatically identify the type of modulation signal, and output the signal recognition accuracy.