Radio modulation signal identification method and system based on hybrid neural network

By using a hybrid neural network in radio modulated signal recognition, combined with FIM, interactive gated Transformer and Bi-LSTM, the problem of difficulty in global analysis and low recognition accuracy in noise environments in the prior art is solved, and high-precision modulated signal recognition and classification are achieved.

CN120050146APending Publication Date: 2025-05-27JIANGSU UNIV OF SCI & TECH
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510268499.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing automatic modulation recognition method based on deep learning is difficult to perform global analysis, and the recognition accuracy is not high in complex noise environments.

Method used

A hybrid neural network-based method is adopted to combine the frequency interaction module (FIM) for signal denoising, and the global and time-series characteristics of the signal are extracted using interactive gated Transformer and Bi-LSTM.

Benefits of technology

The recognition accuracy of radio modulated signals is improved and the robustness of the model's impact on noise is enhanced, so that accurate modulated signals can be recognized and classified in complex noise environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120050146A_ABST
    Figure CN120050146A_ABST
Patent Text Reader

Abstract

The invention discloses a radio modulation signal identification method and system based on a hybrid neural network, and the method comprises the steps: obtaining a radio modulation signal data set, and dividing the radio modulation signal data set into a training set, a test set and a verification set according to a proportion; the method comprises the following steps of: constructing a hybrid neural network model on the basis of a frequency interaction module (FIM), an interaction gate Transform and a Bi-LSTM (Bidirectional Long Short Term Memory); and training and verifying the hybrid neural network model through the training set and the verification set to obtain an identification model, and inputting the test set into the identification model to complete identification of the radio modulation signal in the test set. Denoising is carried out through FIE, the global feature extraction capability of Transform and the time sequence feature extraction capability of Bi-LSTM are effectively combined, the recognition precision is effectively improved through reasonable design, meanwhile, the robustness of noise influence is enhanced, and the model can complete accurate recognition and classification of modulation signals in a complex noise environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of wireless communication, relates to radio modulation signal recognition, and particularly relates to a method and system for radio modulation signal recognition based on a hybrid neural network. Background Art

[0002] With the rapid development of modern communication technologies, the electromagnetic environment where radio signals are located has become complex and variable. Since people have adopted more and more modulation schemes to achieve efficient data transmission, automatic modulation recognition (AMR), as a key step between signal detection and demodulation, has become increasingly important. AMR can quickly identify the modulation type of the received signal without prior knowledge, and shows broad application prospects in both civilian and military technologies, such as radio interference monitoring, spectrum optimization, and electronic countermeasures.

[0003] In recent years, with the rapid development of deep learning (DL) technologies, some researchers have attempted to apply DL technologies to AMR. In 2016, O'Shea et al. first proposed a modulation recognition method based on convolutional neural networks (CNN) for processing raw in-phase and quadrature (I / Q) signals, and opened the dataset RML2016.10a, which attracted a large number of researchers to participate and promoted the development of this field. In 2018, Rajendran et al. proposed a method for automatic modulation recognition using long short-term memory networks (LSTM), proving that LSTM can effectively extract signal temporal features. Xu et al. applied the classic network models ResNet and DenseNet in CNN to modulation recognition, and designed a convolutional long short-term deep neural network (CLDNN) model composed of a series connection of CNN and LSTM, which has better modulation recognition performance compared with the previous models that used CNN or LSTM alone. Based on the CLDNN structure, Njoku et al. proposed a network structure that combines a gated recurrent unit (GRU) improved from LSTM with CNN, achieving higher recognition accuracy with fewer parameters. Sun et al. used multiple layers of CNN to extract signal spatial features and used GRU to extract the temporal features of signals, enhancing the recognition performance by introducing residual connection structures, attention mechanisms, etc.

[0004] However, in the above DL-based AMR methods, few people have tried to conduct global analysis on signals. Modulation signals usually transmit bit sequences without semantic information, and the features of each modulation signal are evenly distributed in the signal waveform. Therefore, feature extraction should not be limited to the local area, but should extract the global features of the signal. In recent years, the Transformer structure has been widely used in the fields of natural language, speech, and image processing, demonstrating powerful feature extraction capabilities. By using multi-head self-attention (MHSA) on the input data to model the context information, the Transformer can obtain longer context information than RNNs, which conforms to the time series characteristics of the signal and can effectively extract the global features of the signal. Moreover, since the self-attention operation can be executed in parallel, GPU parallel computing can be utilized to accelerate the training speed. However, directly inputting the original signal into the Transformer will lead to a decline in recognition performance due to the lack of consideration of the intrinsic features of the modulation signal. In addition, during the propagation of radio signals, noise will inevitably be superimposed on the signal. If no processing is done, the signal and noise features will be confused, thus affecting the recognition accuracy. Summary of the Invention

[0005] Object of the Invention: In order to overcome the deficiencies in the prior art, a radio modulation signal recognition method and system based on a hybrid neural network are provided. Denoising is performed through a frequency interaction module (FIM), and the global feature extraction ability of the Transformer and the time series feature extraction ability of the Bi-LSTM are effectively combined. By reasonable design, while effectively improving the recognition accuracy, the robustness to the influence of noise is enhanced, enabling the model to accurately identify and classify modulation signals in a complex noise environment.

[0006] Technical Solution: To achieve the above object, the present invention provides a radio modulation signal recognition method based on a hybrid neural network, including the following steps:

[0007] S1: Obtain a radio modulation signal data set and divide it into a training set, a test set, and a validation set according to a ratio;

[0008] S2: Based on the frequency interaction module FIM, the interactive gated Transformer, and the Bi-LSTM, construct a hybrid neural network model;

[0009] S3: Train and validate the hybrid neural network model through the training set and the validation set respectively to obtain a recognition model, and input the test set into the recognition model to complete the recognition of radio modulation signals in the test set.

[0010] Further, the construction and operation of the hybrid neural network model in step S2 specifically include:

[0011] A1: Design the FIM. The FIM uses the Fourier transform to extract the frequency features of the signal, and uses a multi-layer perceptron (MLP) to enable the model to adaptively learn the weights and complete the reconstruction of the frequency-domain features. The reconstructed frequency-domain features are converted into time-domain features through the inverse Fourier transform, and added to the input signal to complete signal denoising;

[0012] A2: Segment the input signal, map the signal slices to feature vectors, use the attention mechanism to complete the weight assignment of the feature vectors, and then perform positional embedding to mark the position information for facilitating the subsequent extraction of global features;

[0013] A3: Construct an interactive gated Transformer module. By replacing the MLP in the Transformer with an interactive gated linear unit (IGLU), transform and enhance the context information output by the multi-head self-attention (MHSA), and filter the noise through the gated mechanism to strengthen the local key features, thereby extracting the dependency relationships between different feature vectors and outputting an encoded vector sequence containing global correlations;

[0014] A4: Input the vector sequence into the Bi-LSTM to extract the time-series features, and output a one-dimensional vector containing the modulation mode category through a linear layer, thereby realizing the recognition and classification of the modulation mode.

[0015] Further, step A1 specifically includes:

[0016] A1-1: Perform a Fourier transform (DFT) on the input signal X with length L n , where n is the number of input signals, and its formula is as follows:

[0017]

[0018] The above formula can be expressed as where X Re and X Im respectively represent the real part and the imaginary part of the signal;

[0019] A1-2: After linearly transforming and activating the real part and the imaginary part of the signal through two MLPs respectively with Relu, perform another linear transformation to obtain the corresponding feature weights W Re and W Im , and multiply the feature weights after being activated by Swish with the imaginary part XIm and the real part X Re are interactively fused, and its calculation process is expressed as:

[0020] W Re (X Re ) = Linear(Relu(Linear(X Re ))

[0021] W Im (X Im ) = Linear(Relu(Linear(X Im ))

[0022]

[0023] In the formula, Linear represents a linear layer, which is used to generate learnable feature weights. Relu and Swish are activation functions to increase the non-linear expression ability of features and help the network for training. represents element-wise multiplication. and respectively represent the feature vectors after interactive fusion;

[0024] A1-3: Concatenate the feature vectors after interaction to obtain the features after enhancing the frequency domain features:

[0025]

[0026] A1-4: Transform the enhanced frequency domain features back to the time domain through the inverse Fourier transform (IDFT). This process preserves their features in the time domain to adapt to subsequent feature extraction and transformation:

[0027]

[0028] A1-5: Add the signal after the inverse transformation to the original signal to obtain the denoised signal:

[0029]

[0030] Furthermore, the step A2 specifically includes:

[0031] A2-1: Segment the input signal by designing a sliding window; use a rectangular window with a length of L / N to intercept the signal with a length of L, and divide it into N segments. N is the input dimension of the transformer encoder and is a hyperparameter; if there is not enough data in the window for segmentation, zero values are filled at both ends of the input data to solve it.

[0032] A2-2: Use a linear layer to perform feature encoding on the segmented segments and map them to corresponding feature vectors x(i), where i = 1, …, N:

[0033]

[0034] A2-3: Use an attention mechanism to process the feature vectors to enhance the model's attention to different channels and complete the reallocation of feature weights;

[0035] A2-4: Perform position embedding, that is, generate a one-dimensional position information sequence, embed the position information into the feature sequence to complete the marking of position information; further, perform class embedding. The class embedding x classtoken is a learnable embedding, which is inserted at the zero position of the feature sequence for generating classification predictions; the specific calculation formula is as follows:

[0036]

[0037] where, W position represents the one-dimensional position information of the feature vector.

[0038] Further, the step A2-3 specifically includes:

[0039] A2-3-1: Perform global average pooling on the input feature vectors to embed global information. Assume that the shape of the input x feature vector is C×N, and the feature vector x in the c-th channel of channel C is c expressed as:

[0040]

[0041] A2-3-2: Use the obtained channel attention weights to stimulate the feature channels to obtain the output feature vectors

[0042]

[0043] where, δ represents the Relu activation function, σ represents the Sigmoid activation function, L 1 and L 2 are two linear transformations for learning the importance among channels.

[0044] Further, the step A3 specifically includes:

[0045] Pass the input feature sequence x k through K layers of the same-structured Transformer, calculate the correlation between feature vectors at different positions, and output an encoded vector sequence containing global correlation:

[0046] x′ k= LN(MHSA(x k-1 )) + x k-1 , k = 1, ..., K,

[0047] x k = LN(IGLU(x k ′)) + x k ′, k = 1, ..., K.

[0048] Among them, LN is layer normalization processing to avoid the phenomena of gradient vanishing and gradient explosion during training, making the model more stable; MHSA represents the multi-head self-attention module, and IGLU represents the interactive gating linear unit.

[0049] Furthermore, in step A3, the Transformer encoder uses the MHSA module to capture the correlation between feature vectors at different positions, obtaining an encoded vector containing position correlation; MHSA is expressed as:

[0050] MSHA = Concat(GN(head 1 , ..., head k ))

[0051] Among them, Concat means splicing different heads in the multi-head attention mechanism according to the feature dimension, and GN means batch normalization; the i-th head is expressed as:

[0052] head i = Attention(W Q x k , W K x k , W V x k )

[0053] Among them, W Q , W K , W V are the weight matrices of query, key, and value; the specific calculation process of Attention is expressed as:

[0054]

[0055] Among them, Q, K, and V are the feature matrices obtained by multiplying the input data with the corresponding weight matrices, Calculated through MHSA.

[0056] Further, in step A3, IGLU is used to replace the MLP structure of the traditional Transformer to filter the effective information processed by MHSA. Compared with the single-branch structure of MLP, IGLU transmits information through two branches and interacts with each other, enabling the Transformer to obtain better feature extraction capabilities. IGLU includes linear layers of two branches and the Relu activation function, and its formula is:

[0057]

[0058] Among them, W 1 ,W 2 ,W 3 are the weight vectors generated by the linear layer of the Attention output feature vector respectively. IGLU first performs weighted calculation on it and the input feature vector x′ k to filter the effective information in MHSA; after being transformed by the non-linear activation function, the feature information contained in the vector is separated to further extract the effective information; then, through element-wise multiplication, it completes the interactive fusion with the semantic information of the other branch to enhance the feature extraction ability; finally, the feature information of the two branches is added to complete the fusion of the two branches.

[0059] Further, step A4 specifically includes:

[0060] A4-1: Input the feature sequence x k into the Bi-LSTM to further enhance the ability of the model to extract temporal features. The Bi-LSTM consists of two LSTM sequences, including several LSTM units, which extract the hidden state vectors of the input feature sequence from the forward and backward directions respectively. The hidden state vectors obtained from the forward and backward directions at each time step are further fused, and finally, a feature sequence containing temporal features is output; assuming there are t LSTM units in total, the output X L of the Bi-LSTM is:

[0061]

[0062] Among them, H 1 ,H 2 ,…,H t represent the feature vectors output by each LSTM unit;

[0063] A4-2: Input the final output feature sequence X L of the Bi-LSTM into the classifier to obtain the modulation recognition result.

[0064] Based on the above content, the present invention provides a radio modulation signal recognition system based on a hybrid neural network, including:

[0065] A data acquisition and division module, which is used to obtain a radio modulation signal dataset and divide it into a training set, a test set, and a validation set according to a ratio;

[0066] A model construction module, which is used to construct a hybrid neural network model;

[0067] A model training module, which is used to train and validate the hybrid neural network model to obtain an identification model;

[0068] An output module, which is used to output the identification result of the radio modulation signal.

[0069] Beneficial effects: Compared with the prior art, the present invention has the following advantages:

[0070] 1. The present invention constructs a FIM to denoise the input IQ signal sequence. The FIM extracts the frequency and phase information of the input signal sequence through Fourier transform, uses an MLP to enable the model to adaptively learn weights and complete feature reconstruction, utilizes the internal correlation between the real part and the imaginary part of the signal for interactive fusion, and the fused feature information is converted into time-domain features through IDFT and fused with the original input signal to complete signal denoising. The FIM improves the identification ability of the model in a complex noise environment.

[0071] 2. By adding an attention mechanism to the positional embedding, the model fully learns the context feature information of different slices, thereby improving the feature expression ability of the positional embedding, enhancing the model's positional encoding ability, and enabling the subsequent transformer module to fully extract the global features of the model.

[0072] 3. The present invention replaces the existing MLP structure in the transformer structure with an IGLU to construct an interactive gating Transformer module. The IGLU uses an activation function to implement a gating mechanism, filters the information in the MHSA, passes the information through two branches and interacts with each other, enabling the Transformer to obtain better feature extraction ability.

[0073] 4. The present invention uses a Bi-LSTM to further increase the feature extraction ability of the model. Although positional encoding is used to represent relative position sequence information, the transformer still has limitations in capturing temporal information. The Bi-LSTM fully extracts the temporal features of the signal through two LSTMs with different orders, enhancing the model's feature extraction ability. Description of the drawings

[0074] Figure 1 is a flowchart of the method of the present invention;

[0075] Figure 2 is a structural diagram of the identification model in the present invention;

[0076] Figure 3 Schematic diagram of the attention embedding layer and the attention position encoding structure in the present invention;

[0077] Figure 4 Schematic diagram of the interactive gating Transformer encoder in the present invention;

[0078] Figure 5 Schematic diagram of the confusion matrix for modulating and identifying samples with a signal-to-noise ratio of 0 dB in the present invention;

[0079] Figure 6 Schematic diagram of the confusion matrix for modulating and identifying samples with a signal-to-noise ratio of 10 dB in the present invention;

[0080] Figure 7 Comparison chart of the accuracy rates of the method of the present invention and different methods on the test set. Detailed implementation manners

[0081] The following further clarifies the present invention in conjunction with the accompanying drawings and specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, various equivalent forms of modification of the present invention by those skilled in the art fall within the scope defined by the appended claims of this application.

[0082] Embodiment 1:

[0083] As Figure 1 shown, this embodiment provides a method for identifying radio modulation signals based on a hybrid neural network, including the following steps:

[0084] S1: Obtain a radio modulation signal data set, perform label encoding on data samples, and divide the training set, test set, and validation set in a ratio of 6:2:2;

[0085] S2: Based on the frequency interaction module FIM, the interactive gating Transformer, and the Bi-LSTM, construct a hybrid neural network model;

[0086] Referring to Figure 2 the model structure shown, the construction and operation of the hybrid neural network model specifically include:

[0087] A1: Design the FIM. The FIM uses the Fourier transform to extract the frequency features of the signal, and uses a multi-layer perceptron MLP to enable the model to adaptively learn the weights and complete the reconstruction of the frequency domain features. The reconstructed frequency domain features are converted into time domain features through the inverse Fourier transform, and added to the input signal to complete signal denoising;

[0088] Step A1 specifically includes:

[0089] A1-1: Perform a Fourier transform (DFT) on the signal X with an input length of L = 128. Here, n is the number of input signals, and the formula is as follows: n The above formula can be expressed as

[0090]

[0091] where X and X Re and X Im represent the real and imaginary parts of the signal respectively;

[0092] A1-2: After performing linear transformation and Relu activation on the real and imaginary parts of the signal through two MLPs respectively, perform another linear transformation to obtain the corresponding feature weights W Re and W Im . Then, interact and fuse the feature weights after Swish activation with the imaginary part X Im and the real part X Re components. The calculation process is expressed as:

[0093] W Re (X Re ) = Linear(Relu(Linear(X Re )))

[0094] W Im (X Im ) = Linear(Relu(Linear(X Im )))

[0095]

[0096] In the formula, Linear represents the linear layer, which is used to generate learnable feature weights. Relu and Swish are activation functions, which are used to increase the non-linear expression ability of features and help the network for training. represents element-wise multiplication, and represent the feature vectors after interaction and fusion respectively;

[0097] A1-3: Concatenate the feature vectors after interaction to obtain the features after enhancing the frequency domain features:

[0098]

[0099] A1-4: Transform the enhanced frequency domain features back to the time domain through the inverse Fourier transform (IDFT). The length of the transformed signal is still 128. This process preserves their features in the time domain to adapt to subsequent feature extraction and transformation:

[0100]

[0101] A1-5: Add the signal after inverse transformation to the original signal to obtain the denoised signal:

[0102]

[0103] A2: Segment the input signal, map the signal slices to feature vectors, use the attention mechanism to complete the weight assignment of the feature vectors, and then perform position embedding to mark the position information for subsequent extraction of global features; the structure of the attention embedding layer is as Figure 3 shown;

[0104] Step A2 specifically includes:

[0105] A2-1: Segment the input signal by designing a sliding window; use a rectangular window with a length of 4 to intercept a signal with a length of 128, and divide it into N = 32 segments, where N is the input dimension of the transformer encoder; if there is not enough data in the window for segmentation, pad zeros at both ends of the input data to solve it;

[0106] A2-2: Use a linear layer to perform feature encoding on the segmented segments and map them to corresponding feature vectors x(i), i = 1,..., N:

[0107]

[0108] A2-3: Use the attention mechanism to process the feature vectors to enhance the model's attention to different channels and complete the reallocation of feature weights;

[0109] Step A2-3 specifically includes:

[0110] A2-3-1: Perform global average pooling on the input feature vectors for global information embedding. Assume that the shape of the input x feature vector is C×N, and the feature vector x in the c-th channel in channel C is c expressed as:

[0111]

[0112] A2-3-2: Use the obtained channel attention weights to stimulate the feature channels to obtain the output feature vectors

[0113]

[0114] where, δ represents the Relu activation function, σ represents the Sigmoid activation function, L 1 and L 2 are two linear transformations for learning the importance among channels.

[0115] A2-4: Perform positional embedding, that is, generate a one-dimensional sequence of positional information, embed the positional information into the feature sequence, and complete the marking of the positional information; further, perform class embedding, and the class embedding x classtoken is a learnable embedding, which is inserted at the zero position of the feature sequence for generating classification predictions; the specific calculation formula is as follows:

[0116]

[0117] where, W position represents the one-dimensional positional information of the feature vector.

[0118] A3: Construct an interactive gated Transformer module as shown in Figure 4 . By replacing the MLP in the Transformer with an interactive gated linear unit (IGLU), transform and enhance the context information output by the multi-head self-attention (MHSA), and filter out noise through the gating mechanism to strengthen local key features, so as to extract the dependencies between different feature vectors and output a sequence of encoded vectors containing global correlations;

[0119] Step A3 specifically includes:

[0120] Pass the input feature sequence x k through K = 8 layers of Transformer with the same structure, calculate the correlations between feature vectors at different positions, and output a sequence of encoded vectors containing global correlations:

[0121] x′ k = LN(MHSA(x k-1 )) + x k-1 , k = 1,..., K,

[0122] x k = LN(IGLU(x k ′)) + x k ′, k = 1,..., K.

[0123] where, LN is layer normalization processing to avoid the phenomena of gradient disappearance and gradient explosion during the training process and make the model more stable; MHSA represents the multi-head self-attention module, and IGLU represents the interactive gated linear unit.

[0124] The Transformer encoder uses the MHSA module to capture the correlations between feature vectors at different positions and obtains encoded vectors containing positional correlations; MHSA is expressed as:

[0125] MSHA = Concat(GN(head 1 ,..., head k ))

[0126] Among them, Concat means to splice different heads in the multi-head attention mechanism according to the feature dimension, and GN means batch normalization; the i-th head is expressed as:

[0127] head i = Attention(W Q x k , W K x k , W V x k )

[0128] Among them, W Q , W K , W V are the weight matrices of query, key, and value; the specific calculation process of Attention is expressed as:

[0129]

[0130] Among them, Q, K, and V are the feature matrices obtained by multiplying the input data with the corresponding weight matrices, and are obtained through MHSA calculation.

[0131] Use IGLU to replace the MLP structure of the traditional Transformer to filter the effective information processed by MHSA. Compared with the MLP single-branch structure, IGLU transmits information through two branches and interacts with each other, enabling the Transformer to obtain better feature extraction capabilities. IGLU includes the linear layers of two branches and the Relu activation function, and its formula is:

[0132]

[0133] Among them, W 1 , W 2 , W 3 are the weight vectors generated by the linear layer of the Attention output feature vector respectively. IGLU first performs weighted calculation with the input feature vector x' k to filter the effective information in MHSA; after passing through the transformation of the non-linear activation function, the feature information contained in the vector is separated to further extract the effective information; then through element-wise multiplication, it completes the interaction and fusion with the semantic information of the other branch to enhance the feature extraction ability; finally, the feature information of the two branches is added to complete the fusion of the two branches.

[0134] A4: Input the vector sequence into the Bi-LSTM to extract the time series features, and output a one-dimensional vector containing the modulation mode category through a linear layer, thereby realizing the recognition and classification of the modulation mode;

[0135] Step A4 specifically includes:

[0136] A4-1: Input the feature sequence x k into the Bi-LSTM to further enhance the ability of the model to extract temporal features. The Bi-LSTM consists of two LSTM sequences, including several LSTM units, which respectively extract the hidden state vectors of the input feature sequence from the forward and backward directions. The hidden state vectors obtained from the forward and backward directions at each time step are further fused, and finally a feature sequence containing temporal features is output. Suppose there are a total of t = 16 LSTM units, then the output X L of the Bi-LSTM is:

[0137]

[0138] where, H 1 , H 2 ,…, H t represent the feature vectors output by each LSTM unit;

[0139] A4-2: Input the final output feature sequence X L of the Bi-LSTM into the classifier to obtain the modulation recognition result.

[0140] S3: Select the cross-entropy as the loss function and use the stochastic gradient descent method for iterative training. During the training process, when the loss function remains unchanged for consecutive rounds, it is considered that the network converges and the training stops to obtain the trained model. Otherwise, update the parameters obtained during the training process and retrain. Input the validation set in the sample dataset into the trained model for validation to obtain the final recognition model. Finally, input the test set into the final recognition model to complete the recognition of the radio modulation signals in the test set.

[0141] Embodiment 2:

[0142] To implement the method of Embodiment 1, this embodiment provides a radio modulation signal recognition system based on a hybrid neural network, including:

[0143] A data acquisition and division module, used to obtain the radio modulation signal dataset and divide it into a training set, a test set, and a validation set according to a ratio;

[0144] A model construction module, used to construct a hybrid neural network model;

[0145] A model training module for training and validating a hybrid neural network model to obtain a recognition model;

[0146] An output module for outputting the recognition result of the radio modulation signal.

[0147] Example 3:

[0148] To verify the feasibility of the recognition model provided by the present invention, based on the solutions of Examples 1 and 2, this example uses the RML2016.10a wireless communication dataset for simulation experiments. The dataset considers the time-varying random channel effects common in most wireless systems, including center frequency offset, sampling rate offset, additive Gaussian white noise, multipath, and fading, and can simulate the modulation recognition ability of the model in complex environments. Subsequently, the classification results are analyzed and compared with existing methods. In the simulation experiment, randomly divide each category of different signal-to-noise ratios into training set, test set, and validation set according to the ratio of 6:2:2, set the number of training rounds to 50, and set the learning rate to 10 -3 , when the validation loss does not decrease within 10 iterations, reduce the learning rate to 10 -4 , adopt the Adam optimizer, and the batch size is 512.

[0149] Figure 5 and Figure 6 are the confusion matrices at signal-to-noise ratios of 0 dB and 10 dB. It can be found that the model has a good recognition effect on most modulation methods. From Figure 7 it can be found that as the SNR increases, the accuracy of the modulation recognition method proposed by the present invention continuously increases. When SNR = 10 dB, the recognition rate of most modulation patterns is close to 100%. The model has a poor recognition effect on the QAM modulation method, mainly because QAM is a phase modulation method, and the two QAM modulation methods are easily confused. In addition, the two modulation patterns of WBFM and AM-DSB are easily confused, mainly because the dataset samples the analog audio signal to generate these two modulation data, there are some silent areas in the signal, and the dataset does not perform frequent sampling, resulting in the model being difficult to distinguish the modulation signal characteristics between them.

[0150] To demonstrate the advantages of the method proposed by the present invention, in the simulation experiment, it is compared with the CNN, LSTM, CLDNN, MCformer, and FEA-T models. The first three methods all implement feature extraction and classification based on the CNN and LSTM models, and the FEA-T model directly inputs the signal data into the Transformer to implement feature extraction and classification. As Figure 7 shown, the method of the present invention achieves the highest recognition rate in the test set when SNR is greater than -4 dB, and the highest recognition rate reaches 91.36%, thus proving that the model has a high recognition accuracy in complex environments.

[0151] In summary, considering the disadvantage of low recognition accuracy in the existing methods for modulation recognition under complex noise conditions, the present invention constructs a FIM to denoise the input signal, increases the anti-noise ability of the model, and improves the recognition performance; in the position encoding, the attention mechanism is used to complete the weight allocation of the feature vectors, enhancing the model's ability to extract context feature information; the interactive gating Transformer model is used, and the IGLU is used to effectively extract the dependencies between feature vectors at different positions, so as to output a feature sequence containing global correlations; the time series features are extracted by using Bi-LSTM, further improving the recognition accuracy of the model; finally, the obtained time series characteristics are input into the fully connected layer to obtain the recognition result of the signal modulation type. The present invention improves the Transformer model, and by introducing FIM, attention embedding and Bi-LSTM, while effectively improving the recognition accuracy, the robustness of the model to the influence of noise is enhanced.

Claims

1. A radio modulation signal recognition method based on a hybrid neural network, characterized in that: The steps include: S1: Obtain a radio modulation signal dataset and divide it into a training set, a test set, and a validation set according to the ratio; S2: Construct a hybrid neural network model based on frequency interaction module FIM, interactive gated Transformer and Bi-LSTM; S3: The hybrid neural network model is trained and verified through the training set and the verification set respectively to obtain a recognition model, and the test set is input into the recognition model to complete the recognition of the radio modulation signal in the test set.

2. The radio modulation signal recognition method based on hybrid neural network according to claim 1 is characterized in that: The construction and operation of the hybrid neural network model in step S2 specifically includes: A1: Design FIM. FIM uses Fourier transform to extract the frequency features of the signal, and uses multi-layer perceptron MLP to make the model adaptively learn weights and complete the reconstruction of frequency domain features. The reconstructed frequency domain features are converted into time domain features through inverse Fourier transform and added to the input signal to complete signal denoising. A2: Slice the input signal and map the signal slices into feature vectors. Use the attention mechanism to assign weights to the feature vectors, and then perform position embedding to mark the position information. A3: Construct an interactive gated Transformer module. By replacing the MLP in the Transformer with an interactive gated linear unit, the context information output by the multi-head self-attention is transformed and enhanced, and the noise is filtered through the gating mechanism to strengthen the local key features, thereby extracting the dependency between different feature vectors and outputting a sequence of encoding vectors containing global correlations. A4: Input the vector sequence into Bi-LSTM to extract the time series features, and output a one-dimensional vector containing the modulation mode category through the linear layer to realize the recognition and classification of the modulation mode.

3. The radio modulation signal recognition method based on hybrid neural network according to claim 2 is characterized in that: The step A1 specifically includes: A1-1: For a signal X with an input length of L n Perform Fourier transform, n is the number of input signals, and the formula is as follows: The above formula is expressed as Where X Re and X Im Represent the real and imaginary parts of the signal respectively; A1-2: After two MLPs perform linear transformation on the real and imaginary parts of the signal and ReLU activation, perform another linear transformation to obtain the corresponding feature weight W Re and W Im , the feature weight after Swish activation and the imaginary part X Im and real part X Re The components are interactively fused, and the calculation process is expressed as: W Re (X Re )=Linear(Relu(Linear(X Re )) W Im (X Im )=Linear(Relu(Linear(X Im )) In the formula, Linear represents the linear layer, which is used to generate learnable feature weights. Relu and Swish are activation functions to increase the nonlinear expression ability of features and help the network to train. represents element-wise multiplication, and They represent the feature vectors after interactive fusion respectively; A1-3: Concatenate the feature vectors after the interaction to obtain the features after enhancing the frequency domain features: A1-4: Transform the enhanced frequency domain features back to the time domain through inverse Fourier transform: A1-5: Add the inverse transformed signal to the original signal to obtain the denoised signal:

4. The radio modulation signal recognition method based on hybrid neural network according to claim 2 is characterized in that: The step A2 specifically includes: A2-1: Split the input signal by designing a sliding window; use a rectangular window of length L / N to intercept the signal of length L and divide it into N segments N is the transformer encoder input dimension, which is a hyperparameter; A2-2: Use the linear layer to encode the features of the segmented fragments and map them into corresponding feature vectors x(i), i=1,…,N: A2-3: Use the attention mechanism to process the feature vector to enhance the model's attention to different channels and redistribute the feature weights; A2-4: Perform position embedding, that is, generate a one-dimensional position information sequence, embed the position information into the feature sequence, and complete the labeling of the position information; perform class embedding, class embedding x classtoken It is a learnable embedding that is inserted into the zero position of the feature sequence to generate classification predictions; the specific calculation formula is as follows: Among them, W position Represents the one-dimensional position information of the feature vector.

5. The radio modulation signal recognition method based on hybrid neural network according to claim 4 is characterized in that: The step A2-3 specifically includes: A2-3-1: Perform global average pooling on the input feature vector to embed global information. Assume that the shape of the input x feature vector is C×N, and the feature vector x of the cth channel in channel C is c It is expressed as: A2-3-2: Use the obtained channel attention weight to stimulate the feature channel and obtain the output feature vector Among them, δ represents the Relu activation function, σ represents the Sigmoid activation function, and L1 and L2 are two linear transformations used to learn the importance of each channel.

6. The radio modulation signal recognition method based on hybrid neural network according to claim 2 is characterized in that: The step A3 specifically includes: The input feature sequence x k After passing through K layers of Transformer with the same structure, the correlation between feature vectors at different positions is calculated, and the encoding vector sequence containing global correlation is output: x′ k =LN(MHSA(x k-1 ))+x k-1 ,k=1,...,K, x k =LN(IGLU(x k ′))+x k ′,k=1,...,K. Among them, LN is layer normalization processing; MHSA represents the multi-head self-attention module, and IGLU represents the interactive gated linear unit.

7. The radio modulation signal recognition method based on hybrid neural network according to claim 6 is characterized in that: In step A3, the Transformer encoder uses the MHSA module to capture the correlation between feature vectors at different positions to obtain a coding vector containing position correlation; MHSA is expressed as: MSHA=Concat(GN(head1,...,head k )) Among them, Concat means concatenating different heads in the multi-head attention mechanism according to the feature dimension, GN means batch normalization; the i-th head is expressed as: head i =Attention(W Q x k ,W K x k ,W V x k ) Among them, W Q ,W K ,W V is the weight matrix of query, key, and value; the specific calculation process of Attention is expressed as: Among them, Q, K, and V are the feature matrices obtained by multiplying the input data with the corresponding weight matrix. Calculated by MHSA.

8. The radio modulation signal recognition method based on hybrid neural network according to claim 6 is characterized in that: In step A3, the IGLU includes two branched linear layers and a Relu activation function, and its formula is: Among them, W1, W2, and W3 are the weight vectors generated by the Attention output feature vector through the linear layer. IGLU first compares it with the input feature vector x′ k A weighted calculation is performed to filter out valid information in MHSA. After that, the feature information contained in the vector is separated after being transformed by a nonlinear activation function to further extract valid information. Then, the vector is interactively fused with the semantic information of another branch by element-by-element multiplication to enhance the feature extraction capability. Finally, the feature information of the two branches is added to complete the fusion of the two branches.

9. The radio modulation signal recognition method based on hybrid neural network according to claim 2 is characterized in that: The step A4 specifically includes: A4-1: Change the feature sequence x k The input to Bi-LSTM further increases the model's ability to extract time series features. Bi-LSTM consists of two LSTM sequences, which include several LSTM units. The hidden state vectors of the input feature sequence are extracted from the forward and reverse directions respectively. The hidden state vectors obtained from the forward and reverse directions are further fused at each time step, and finally the feature sequence containing the time series features is output. Assuming there are t LSTM units in total, the output X of Bi-LSTM is L for: Among them, H1, H2, …, H t The feature vector representing the output of each LSTM unit; A4-2: Bi-LSTM finally outputs the feature sequence X L Input into the classifier to obtain the modulation recognition result.

10. A radio modulation signal recognition system based on a hybrid neural network, characterized in that: include: A data acquisition and division module is used to obtain a radio modulation signal data set and divide it into a training set, a test set, and a validation set according to a ratio; Model building module, used to build hybrid neural network models; A model training module is used to train and verify the hybrid neural network model to obtain a recognition model; The output module is used to output the recognition result of the radio modulation signal.

Citation Information

Cited By

  • Complex electromagnetic signal demodulation method based on multi-task learning

    CN120896824A

  • Signal identification system of neural network based on biological superficial brain architecture

    CN120910783A

  • Electric energy quality disturbance denoising method based on artificial intelligence

    CN121092861A