A Joint Modulation Recognition Method Based on Attention Mechanism and Residual Structure

By using BoTAMCNet and Transformer networks in the high and low signal-to-noise ratio intervals, combining deep separation convolution and global deep convolution feature reconstruction, the existing modulation recognition methods have solved the problem of low recognition accuracy and high complexity under low signal-to-noise ratio, and achieved more efficient modulation recognition effect.

CN115514597BActive Publication Date: 2025-07-08ZHENGZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211125460.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-15
Publication Date
2025-07-08
Estimated Expiration
2042-09-15

AI Technical Summary

Technical Problem

The existing modulation recognition method based on deep learning is too low in recognition accuracy and complexity under low signal-to-noise ratio conditions, and the number of network training parameters is huge, which has the problem of wasted computing resources.

Method used

A joint modulation recognition method based on attention mechanism and residual structure is adopted, and signal classification is used using the signal-to-noise ratio blind estimation calculation method. The high signal-to-noise ratio interval is used by the BoTAMCNet network, and the low signal-to-noise ratio interval is used by the Transformer network. Combined with deep separable convolution and global deep convolution feature reconstruction, feature extraction and recognition are performed through the multi-head self-attention mechanism.

Benefits of technology

The accuracy of recognition in the low signal-to-noise ratio interval is significantly improved, and the recognition rate of the high signal-to-noise ratio interval is further improved. At the same time, the network complexity and computing burden are reduced, and more efficient modulation recognition performance is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115514597B_ABST
    Figure CN115514597B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of wireless communication technologies, and particularly to a joint modulation recognition method based on an attention mechanism and a residual structure. Aiming at the problems in the prior art such as too low recognition accuracy under low signal-to-noise ratio conditions, too high complexity, too many multiplication and addition operations under the same conditions, and a large number of network trainable parameters, the following solutions are proposed, including the following steps: Step A: For the received signal type to be recognized, the signal subspace dimension is estimated by eigenvalue decomposition to estimate the signal-to-noise ratio of the received signal; The purpose of the present invention is to verify the recognition effectiveness of the joint structure through a large number of data experiments on the DeepSig open-source dataset RadioML2018.01A. The recognition accuracy in the low signal-to-noise ratio interval is significantly improved, and the recognition rate in the high signal-to-noise ratio interval is also improved. Through the simulation of the number of network trainable parameters and the inference time, it is verified that BoTAMCNet has relatively low complexity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of wireless communication, and particularly relates to a joint modulation recognition method based on an attention mechanism and a residual structure. Background Technique

[0002] For the classical blind SNR estimation algorithm, the eigenvalues of the matrix are obtained by constructing the autocovariance matrix of the received signal and performing eigenvalue decomposition, and the minimum description length criterion, originally called Minimum Description Length in English, abbreviated as MDL, is used to estimate the dimension of the signal subspace so as to estimate the SNR.

[0003] The multi-head self-attention mechanism, originally called Multi-Head Self Attention in English, abbreviated as MHSA, its main idea is: without any convolutional layer or recurrent layer, it is implemented through an encoder-decoder model. The attention layer determines the attention weights of different parts of the input sequence at each position during the encoding process, and then calculates the hidden vector representation of this part in the form of a weighted sum. The advantage of the self-attention mechanism is that it is easier to capture the features of medium- and long-distance interdependencies in the sentence, and it also directly helps to increase the parallelism of the calculation.

[0004] The residual structure, originally called Residual Network in English, abbreviated as ResNet, its main idea is: to solve the degradation problem in deep networks, a shortcut connection method is used to add the feature matrices across layers, weakening the strong connection between each layer. Through preprocessing of the data, a deeper network structure, the residual module, is used, and BatchNormalization is used to accelerate the training. It has been proven that as the network is continuously deepened using the residual structure, not only the degradation problem of the network is solved, but also the overall effect of the network is improved.

[0005] The depthwise separable convolution, originally called Deep Separable Convolution in English, abbreviated as DSC, its main idea is: first, the depthwise convolutional layer is used to perform a separate convolution operation on each channel of the input, and then the pointwise convolutional layer is used to fuse all channels. The depthwise convolutional kernel is single-channel and the number of convolutional kernels is equal to the number of input channels. Only one depthwise convolutional kernel is used for filtering each channel. Completing single-channel filtering and multi-channel fusion in sequence can obtain the same output dimension. Its advantage is that compared with the commonly used standard convolution, under the same condition settings, the number of multiply-add operations and the number of network trainable parameters of the DSC operation are significantly reduced compared with the standard convolution, the complexity is significantly reduced, and it will accelerate the forward inference of the model.

[0006] The global depthwise convolution feature reconstruction algorithm, originally known as Global Depthwise Convolution and abbreviated as GDWConv, has the main idea that by learning the importance of features at different positions, the classification ability of the reconstructed features is enhanced through the ReLU activation function and the bias term. Its essence is a depth convolution layer, and the dimension of the convolution kernel is the same as the size of the input feature map. Its advantage is that it can reflect the importance difference of features at different positions within a single channel of the feature map, thereby enhancing the overall classification ability of the network. This difference is caused by the different characteristics presented by signal samples within different time ranges due to the randomness and variability of the transmitted signal.

[0007] Existing modulation recognition methods based on deep learning have achieved good results, but there are problems such as too high complexity and the recognition rate can be further improved, especially the problem of too low recognition accuracy under low signal-to-noise ratio conditions. Therefore, it is very important to propose a joint modulation recognition method based on the attention mechanism and the residual structure. Summary of the Invention

[0008] Aiming at the problems in the prior art such as too low recognition accuracy under low signal-to-noise ratio conditions, too high complexity, a large number of multiply-accumulate operations under the same conditions, and a large number of network trainable parameters, a joint modulation recognition method based on the attention mechanism and the residual structure is proposed.

[0009] To achieve the above objectives, the present invention adopts the following technical solutions:

[0010] A joint modulation recognition method based on the attention mechanism and the residual structure, comprising the following steps:

[0011] Step A: For the received signal type to be recognized, perform signal-to-noise ratio blind estimation detection, and estimate the signal-to-noise ratio of the received signal by estimating the dimension of the signal subspace using eigenvalue decomposition.

[0012] Step B: In the high signal-to-noise ratio interval, design a superimposed and optimized convolutional neural network BoTAMCNet. Use the basic structure of ResNet18, replace the standard convolution with depthwise separable convolution, and replace the original convolution of the two residual units in the residual block with a multi-head self-attention (abbreviated as MHSA) module. In this way, the residual structure is superimposed for feature extraction, and then the global depth convolution operation is used for feature reconstruction to increase the feature dimension.

[0013] Step C: In the low signal-to-noise ratio range, design a Transformer architecture, optimize it by adding different fully connected layers and activation functions on the basis of the Transformer Block, use the multi-head self-attention mechanism for recognition, and let the attention layer judge the importance of different regions in the input sequence;

[0014] Step D: Binarize and classify the received signals blindly estimated by signal-to-noise ratio in Step A into two parts: high signal-to-noise ratio (0 dB to 20 dB) signals and low signal-to-noise ratio signals (-20 dB to 0 dB). The high signal-to-noise ratio signals are input into BoTAMCNet for recognition, and the low signal-to-noise ratio signals are input into the Transformer for recognition.

[0015] Preferably, the specific steps in Step A are as follows:

[0016] A1: Construct the autocovariance matrix of the received signal and perform eigenvalue decomposition to extract the eigenvalues;

[0017] A2; Estimate the dimension of the signal subspace through the minimum description length criterion;

[0018] A3: Estimate the signal power and noise power through the dimension of the signal subspace, so as to estimate the signal-to-noise ratio.

[0019] Preferably, the specific steps in A1 are as follows:

[0020] Let the received signal be y(n), that is: y(n) = x(n) + σ(n), construct the M-order autocovariance matrix R of the received signal and decompose it into:

[0021]

[0022] where I is the M-order identity matrix, is the (M - k)-fold eigenvalue of R, R ss is the autocorrelation matrix of the modulation signal, assume its rank is d (d < M), and perform eigenvalue decomposition on it to get: R ss = U∑U H , where U is the orthogonal matrix composed of the eigenvectors of R ss , its diagonal elements are the eigenvalues of R ss . So the diagonal elements of are the eigenvalues of R:

[0023]

[0024] Preferably, the specific steps in A2 are as follows:

[0025] Estimate the dimension of the signal subspace through the Minimum Description Length (MDL) criterion. First, define the test function:

[0026]

[0027] The MDL criterion combines the log-likelihood function with a penalty function and uses the penalty function for bias correction to make the MDL criterion consistent:

[0028]

[0029] Therefore, the dimension of the signal subspace can be estimated as:

[0030]

[0031] Preferably, the specific steps in A3 are as follows:

[0032] From the estimated value of the dimension of the signal subspace, the estimated noise power can be obtained as:

[0033]

[0034] Estimate the signal power as:

[0035]

[0036] The ratio of the above-estimated signal power to the estimated noise power is the estimated value of the signal-to-noise ratio:

[0037]

[0038] The above is the principle of the signal-to-noise ratio blind estimation algorithm. In the joint structure proposed by the present invention, this algorithm is used to divide 24 kinds of signal data with a signal-to-noise ratio of -20 dB to 0 dB in the data set into low signal-to-noise ratio signals through setting a threshold and input them into the Transformer for recognition; 24 kinds of signal data with a signal-to-noise ratio of 0 dB to 20 dB in the data set are divided into high signal-to-noise ratio signals and input into the BoTAMCNet for recognition.

[0039] Preferably, the specific steps in B are as follows:

[0040] B1: Deploy a standard convolutional layer with 64 convolutional kernels at the beginning of the network, aiming to extract sufficient initial information from the input signal;

[0041] B2: Use the depthwise separable convolution to introduce the skip connection method to design the residual structure, and multiple DSC residual blocks are used in series;

[0042] B3: In the residual structure designed in step B2, use the multi-head self-attention module to replace the original convolution between two residual units in the last residual block to extract more effective feature information;

[0043] B4: After the DSC operations in steps B2 and B3, use global depth convolution operation for feature reconstruction, learn the feature importance at different positions, and then enhance the classification ability of the reconstructed features through the ReLU activation function and bias term.

[0044] Preferably, the specific content of B2 is as follows:

[0045] In the feature extraction part, multiple DSC residual blocks are used in series. At the beginning of the DSC residual block, a linear 1×1 convolution is deployed to perform channel (feature) fusion on the output of the maximum pooling of the previous residual block. At the end of the residual block, a maximum pooling operation is deployed to reduce the feature dimension.

[0046] The DSC operation process is completed in two steps: The first step is that the depth convolution layer performs convolution operations on each input channel separately; the second step is that the point convolution layer performs fusion on all channels. In the depth convolution layer, the depth convolution kernel is single-channel and the number of convolution kernels is equal to the number of input channels. The height and width still use 5×5. The point convolution layer is a standard convolution layer with a height and width of 1 for the convolution kernel. The size of a single point convolution kernel is 1×1×3. The input and output of a single point convolution kernel operation are the same in the height and width dimensions, but the output becomes single-channel, in the form of 8×8×1.

[0047] Preferably, the specific content of B3 is as follows:

[0048] 4 DSC residual blocks are deployed in the feature extraction part. In the fourth DSC residual block, the MHSA module is used to replace the original convolution. After the convolution of two residual units, they are added after passing through the MHSA module, and finally pass through the maximum pooling. At the same time, a layer of standard convolution is deployed in front of and behind the 4 DSC residual blocks respectively. To reduce the number of parameters and the computational burden, the number of convolution kernels in the first two DSC residual blocks is designed to be 32. And to increase the feature dimension and extract sufficient information from the current input for feature extraction, the number of convolution kernels in the first standard convolution layer (First Conv) and the last two DSC residual blocks is designed to be 64, that is, the number of features is increased from 32 categories to 64 categories.

[0049] Preferably, the specific content of B4 is as follows:

[0050] After the DSC operation, in the feature reconstruction part, BoTAMCNet proposes to use the GDWConv operation to learn the feature importance at different positions, and then enhance the classification ability of the reconstructed features through the ReLU activation function and bias term. The global depth convolutional layer is a depth convolutional layer whose convolutional kernel dimension is equal to the size of the input feature map. The output of the global depth convolutional layer can be expressed as:

[0051]

[0052] In the formula is the feature map (full English name: Last Feature Map, abbreviated as LFM), with a size of W×H×M; is the global depth convolutional kernel, with a size of W×H×M; is the bias term; is the size of the output classification feature vector, which is 1×1×M; in addition, (i, j) represents the spatial position index; m represents the channel index. Further, the size of the global depth convolutional kernel is equal to that of the input LFM The feature reconstruction operation occurs between the corresponding channels of and For a specific channel m, first multiply the values at the same position (i, j) of and , then sum all the dot product results and add the bias term Finally, the output after passing through the ReLU activation function is a classification feature;

[0053] After feature reconstruction, an FC layer activated by Softmax (normalized exponential function) is added to convert the prediction result into a probability output. During the training process, Lazy Adam is used as the model optimizer instead of the commonly used SGD optimizer because in this structure, the SGD optimizer is prone to getting stuck in local minima and has a slow descent speed. Lazy Adam can handle sparse updates more effectively.

[0054] Preferably, the specific steps in step C include the following steps:

[0055] C1: The encoder in the Transformer maps the input sequence (x1,..., x n ) represented by symbols to a sequence z = (z1,..., z n ) of continuous representations;

[0056] C2: The decoder generates a symbol output sequence (y1,..., y m ) one element at a time. In each step, the model is autoregressive, and when generating the next step, the previously generated symbol is used as an additional input;

[0057] C3: The Transformer follows the overall architectures of C2 and C3, and uses stacked self-attention and point-wise fully connected layers in the encoder and decoder therein;

[0058] C4: Global average pooling, batch normalization, Alpha-dropout, etc. are added on the basis of the Transformer Block to ensure various characteristics of the structure.

[0059] Preferably, C1 specifically includes the following:

[0060] The encoder consists of a stack of N = 6 identical layers. Each layer has two sub-layers. The first layer is the multi-head self-attention mechanism, and the second layer is a simple fully connected feed-forward network. Residual connections are used around each of the two sub-layers and layer normalization is performed. The output of each sub-layer is LayerNorm(x + Sulayer(x)), where Sulayer(x) is a function implemented by the sub-layer itself. To optimize these residual connections, all sub-layers in the model as well as the embedding layer generate outputs with a dimension of d model = 512.

[0061] Preferably, C2 specifically includes the following:

[0062] The decoder also consists of a stack of N = 6 identical layers. In addition to the two sub-layers in each encoder layer, the decoder also inserts a third sub-layer that performs multi-head attention on the output of the encoder. Similar to the encoder, residual connections are used around each sub-layer and layer normalization is performed. At the same time, a masking layer that guarantees sequence information is added to ensure that the prediction at position i can only depend on the known outputs at positions less than i.

[0063] Preferably, C3 specifically includes the following:

[0064] The input sequence passes through the attention mechanism, the weights are calculated by the compatibility function, and the output is the weighted sum, which can be expressed by the following formula:

[0065]

[0066] Among them, K, Q, and V in the attention function respectively represent Key, Query, and Value, and d k is the number of columns of the Q and K matrices, that is, the vector dimension. The attention function maps the Query and a set of key-value pairs to the output, where Query, Key, Value, and Output are all vectors. The output is the weighted sum of the Values, and the weight assigned to each value is calculated by the correlation function between the Query and the current Key.

[0067] Except for the attention sub-layers, each layer of the encoder and decoder contains a fully connected feed-forward network, which consists of two linear transformations connected by the ReLU activation function.

[0068] FFN(x) = max(0, xW1 + b1)W2 + b2,

[0069] The linear transformation at each position is the same, but the parameters between different layers are different. The input and output dimensions of the network are both d model = 512, but the dimension of the intermediate layer is d ff = 2048.

[0070] Preferably, the C4 specifically includes the following:

[0071] The Transformer architecture used in the present invention first adds global average pooling (English full name: Global Average Pooling, abbreviation: GAP) on the basis of the Transformer Block to enhance the structural stability. After the traditional max pooling layer, one or n fully connected layers are required before the softmax classification can be used, which is prone to overfitting. Here, GAP is used instead of max pooling, and GAP is used to replace the fully connected layer. The purpose is to achieve dimensionality reduction, greatly reduce the parameters in the network, reduce the computational complexity, and prevent overfitting at the same time. After the GAP layer, batch normalization (English full name: Batch Normalization, abbreviation: BN) is used to accelerate the convergence speed and improve the model accuracy. The BN layer estimates the sample mean and variance of the entire training data set through moving average, learns the appropriate offsets and scalings, and uses them during prediction to obtain a definite output. After the BN layer, Alpha-dropout is used to ensure that the self-normalizing characteristics of the previous BN layer remain unchanged. At the same time, Alpha-dropout is used together with the SeLU activation function to ensure that the mean and variance of the input are maintained at their original values and prevent overfitting during the training process. During the training process, Alpha-Dropout will zero some elements with a probability sampled from the Bernoulli distribution, but does not delete the weights on the path. During each forward call, the remaining elements are random and will be scaled and shifted with the mean and variance unchanged. At the same time, using the SeLU activation function can ensure the self-normalizing characteristics of the attenuation layer. Finally, the Softmax activation function (normalized exponential function) is added to convert the prediction result into a probability output. During the training process, Lazy Adam is used as the model optimizer. On the basis that Adam designs independent adaptive learning rates for different parameters by calculating the first-order moment estimate and second-order moment estimate of the gradient, it can further update the first-order momentum and second-order momentum according to the current batch at each iteration, and only update the moving average accumulator of the sparse variable index that appears in the current batch.

[0072] The beneficial effects of the present invention are as follows: For wireless communication signals, a joint modulation recognition method based on the attention mechanism and the residual structure is proposed. The modulation type of the received signal is binary classified by using a signal-to-noise ratio blind estimation algorithm and then input into two networks for automatic recognition. First, in the high signal-to-noise ratio range, a stacked optimized convolutional neural network BoTAMCNet is proposed. The depthwise separable convolution is used to introduce the skip connection method to stack the residual structure, and at the same time, the multi-head self-attention mechanism is added to replace part of the convolution. Secondly, in the low signal-to-noise ratio range, the self-attention mechanism of Transformer is used for recognition, and the attention layer judges the importance of different regions of the input sequence. Through a large number of data experiments on the DeepSig open-source dataset RadioML2018.01A, the recognition effectiveness of the joint structure is verified. The recognition accuracy in the low signal-to-noise ratio range is significantly improved, and the recognition rate in the high signal-to-noise ratio range is further enhanced. At the same time, through the simulation of the number of trainable network parameters and the inference time, it is verified that BoTAMCNet has a relatively low complexity. Description of the Drawings

[0073] Figure 1 is the flowchart of the present invention;

[0074] Figure 2 is the schematic diagram of the modulation type joint structure recognition classifier;

[0075] Figure 3 is the schematic diagram of the DSC residual unit and the DSC residual block;

[0076] Figure 4 is the schematic diagram of the DSC depth convolution operation;

[0077] Figure 5 is the schematic diagram of the DSC point convolution operation;

[0078] Figure 6 is the block diagram of the Transformer structure;

[0079] Figure 7 is the schematic diagram of the multi-head self-attention network architecture;

[0080] Figure 8 is the schematic diagram of the recognition performance comparison of the two networks and other networks on the dataset;

[0081] Figure 9 is the schematic diagram of the recognition performance comparison of the joint structure and other networks on the dataset. Detailed Embodiment

[0082] Next, the technical solutions in the embodiments of the present invention will be described clearly and completely. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments.

[0083] Example 1

[0084] Reference Figures 1-9 , a joint modulation recognition method based on the attention mechanism and the residual structure, includes the following steps:

[0085] Step A: For the received signal type to be recognized, perform blind SNR estimation detection, and estimate the SNR of the received signal by using eigenvalue decomposition to estimate the dimension of the signal subspace;

[0086] Step B: In the high SNR interval, use depthwise separable convolution to replace the standard convolution, and use the multi-head self-attention (abbreviated as MHSA, full English name Multi-Head Self Attention) module to replace the original convolution of the two residual units in the residual block, so as to stack the residual structure for feature extraction, and then use global depth convolution operation for feature reconstruction to increase the feature dimension;

[0087] Step C: In the low SNR interval, design a Transformer architecture, add different fully connected layers and activation functions for optimization based on the Transformer Block, and judge the importance of different regions in the input sequence by the attention layer;

[0088] Step D: Binarize and classify the received signal blindly estimated by SNR in Step A, input the high SNR signal into BoTAMCNet for recognition, and input the low SNR signal into Transformer for recognition.

[0089] In this embodiment, the specific steps in Step A include the following:

[0090] A1: Construct the autocovariance matrix of the received signal and perform eigenvalue decomposition to extract eigenvalues;

[0091] A2; Estimate the dimension of the signal subspace through the minimum description length criterion;

[0092] A3: Estimate the signal power and noise power through the dimension of the signal subspace, so as to estimate the SNR.

[0093] In this embodiment, the specific content in A1 includes the following:

[0094] Let the received signal be y(n), that is: y(n) = x(n) + σ(n), construct the M-order autocovariance matrix R of the received signal and decompose it into:

[0095]

[0096] where I is the M-order identity matrix, is the (M-k)-fold eigenvalue of R, Rss Let \(R\) be the autocorrelation matrix of the modulation signal, and assume its rank is \(d\) (\(d \lt M\)). Perform eigenvalue decomposition on it to obtain: \(R\) ss \( = U\sum U^H\) H where \(U\) is an orthogonal matrix composed of the eigenvectors of \(R\) ss , and its diagonal elements are the eigenvalues of \(R\). So ss the diagonal elements of \(\sum\) are the eigenvalues of \(R\): \(\lambda_1, \lambda_2, \cdots, \lambda_d\)

[0097]

[0098] In this embodiment, the specific steps in \(A2\) are as follows:

[0099] Estimate the dimension of the signal subspace through the Minimum Description Length (MDL) criterion. First, define the test function:

[0100]

[0101] The MDL criterion combines the log-likelihood function with a penalty function, and uses the penalty function to correct the bias so that the MDL criterion has consistency:

[0102]

[0103] So the dimension of the signal subspace can be estimated as:

[0104]

[0105] In this embodiment, the specific steps in \(A3\) are as follows:

[0106] From the estimated value of the dimension of the signal subspace, the estimated noise power can be obtained as:

[0107]

[0108] The estimated signal power is:

[0109]

[0110] The ratio of the above estimated signal power to the estimated noise power is the estimated value of the signal-to-noise ratio:

[0111]

[0112] The above is the principle of the blind estimation algorithm for the signal-to-noise ratio. Input it into the Transformer for recognition; 24 kinds of signal data from 0 dB to 20 dB in the dataset are classified as high signal-to-noise ratio signals and input into the BoTAMCNet for recognition.

[0113] In this embodiment, the specific steps in step B are as follows:

[0114] B1: Deploy a standard convolutional layer with 64 convolutional kernels at the beginning of the network to extract sufficient initial information from the input signal;

[0115] B2: Use the depthwise separable convolution to introduce a skip connection method to design a residual structure, and multiple DSC residual blocks are used in series;

[0116] B3: In the residual structure designed in B2, use the multi-head self-attention module to replace the original convolution between two residual units in the last residual block to extract more effective feature information;

[0117] B4: After the DSC operations in B2 and B3, use the global depth convolution operation for feature reconstruction to learn the importance of features at different positions, and then enhance the classification ability of the reconstructed features through the ReLU activation function and the bias term.

[0118] In this embodiment, the specific steps in B2 are as follows:

[0119] In the feature extraction part, multiple DSC residual blocks are used in series. At the beginning of the DSC residual block, a linear 1×1 convolution is deployed to perform channel (feature) fusion on the output of the max pooling of the previous residual block, and a max pooling operation is deployed at the end of the residual block to reduce the feature dimension.

[0120] The DSC operation process is completed in two steps: the first step is that the depth convolution layer performs convolution operations on each input channel separately; the second step is that the point convolution layer fuses all channels. In the depth convolution layer, the depth convolution kernel is single-channel and the number of convolution kernels is equal to the number of input channels, and the height and width still use 5×5. The point convolution layer is a standard convolutional layer with a height and width of 1 for the convolution kernel. The size of a single point convolution kernel is 1×1×3. The input and output of a single point convolution kernel operation are the same in the height and width dimensions, but the output becomes a single-channel 8×8×1 form.

[0121] In this embodiment, the specific steps in B3 are as follows:

[0122] Four DSC residual blocks are deployed in the feature extraction part. In the fourth DSC residual block, the MHSA module is used to replace the original convolution. After two residual units are convolved and passed through the MHSA module, they are added together, and finally max pooling is performed. At the same time, a layer of standard convolution is deployed in front of and behind the four DSC residual blocks respectively. To reduce the number of parameters and computational burden, the number of convolution kernels in the first two DSC residual blocks is designed to be 32. And to increase the feature dimension and extract sufficient information from the current input for feature extraction, the number of convolution kernels in the first standard convolution layer (First Conv) and the last two DSC residual blocks is designed to be 64, that is, the number of features is increased from 32 classes to 64 classes.

[0123] In this embodiment, the specific content of B4 is as follows:

[0124] After DSC operation, the classification ability of the reconstructed features is enhanced through the ReLU activation function and the bias term. The global depth convolution layer is a depth convolution layer, and the dimension of its convolution kernel is the same as the size of the input feature map. The output of the global depth convolution layer can be expressed as:

[0125]

[0126] In the formula is the feature map (English full name: Last Feature Map, abbreviation: LFM), and its size is W×H×M; is the global depth convolution kernel, and its size is W×H×M; is the bias term; is the output classification feature vector with a size of 1×1×M; in addition, (i, j) represents the spatial position index; m represents the channel index. For a specific channel m, first multiply the values at the same position (i, j) of and Then sum all the dot product results and add the bias term Finally, after passing through the ReLU activation function, a classification feature is output.

[0127] After feature reconstruction, an FC layer activated by Softmax (normalized exponential function) is added to convert the prediction result into a probability output. During the training process, Lazy Adam is used as the model optimizer instead of the commonly used SGD optimizer because in this structure, the SGD optimizer is prone to falling into local minima and has a slow descent speed. Lazy Adam can handle sparse updates more effectively.

[0128] In this embodiment, the specific steps in step C are as follows:

[0129] C1: In the Transformer, the encoder encodes the input sequence (x1,..., xn ) The sequence z=(z1,...,z n ) mapped to a continuous representation;

[0130] C2: The decoder generates a symbolic output sequence (y1,...,y m ) one element at a time. At each step, the model is autoregressive and uses the previously generated symbols as additional inputs when generating the next step;

[0131] C3: The Transformer follows the overall architecture of C2 and C3, using stacked self-attention and point-wise fully connected layers for both the encoder and the decoder;

[0132] C4: Global average pooling, batch normalization, Alpha-dropout, etc. are added on top of the Transformer Block to ensure various characteristics of the structure.

[0133] In this embodiment, the specific content of C1 is as follows:

[0134] The encoder consists of a stack of N = 6 identical layers. Each layer has two sub-layers. The first layer is the multi-head self-attention mechanism, and the second layer is a simple fully connected feed-forward network. Residual connections are used around each of the two sub-layers and layer normalization is performed. The output of each sub-layer is LayerNorm(x + Sulayer(x)), where Sulayer(x) is the function implemented by the sub-layer itself. All sub-layers in the model as well as the embedding layer produce outputs with a dimension of d model = 512.

[0135] In this embodiment, the specific content of C2 is as follows:

[0136] The decoder also consists of a stack of N = 6 identical layers. In addition to the two sub-layers in each encoder layer, a third sub-layer is inserted into the decoder. Similar to the encoder, residual connections are used around each sub-layer and layer normalization is performed. At the same time, a masking layer that ensures sequence information is added to ensure that the prediction at position i can only depend on the known outputs at positions less than i.

[0137] In this embodiment, the specific content of C3 is as follows:

[0138] The input sequence passes through the attention mechanism, and the weights are calculated through the compatibility function. The output is a weighted sum, which can be expressed by the following formula:

[0139]

[0140] Among them, K, Q, and V in the attention function represent Key, Query, and Value respectively, and dk is the number of columns of the Q and K matrices, which is also the vector dimension. The number of attention heads maps a Query and a set of key-value pairs to an output, where Query, Key, Value, and Output are all vectors. The output is a weighted sum of the Values, and the weight assigned to each value is calculated by a correlation function between the Query and the current Key.

[0141] Except for the attention sublayer, each layer of the encoder and decoder contains a fully connected feed-forward network, which consists of two linear transformations connected by the ReLU activation function.

[0142] FFN(x) = max(0, xW1 + b1)W2 + b2,

[0143] The linear transformation at each position is the same, but the parameters between different layers are different. The input and output dimensions of the network are both d model = 512, but the dimension of the middle layer is d ff = 2048.

[0144] In this embodiment, the specific content of C4 is as follows:

[0145] The Transformer architecture used in the present invention first adds global average pooling (English full name: Global Average Pooling, abbreviation: GAP) on the basis of the Transformer Block to enhance the structural stability. Here, GAP is used instead of max pooling, and GAP is used to replace the fully connected layer. The purpose is to achieve dimensionality reduction, greatly reduce the parameters in the network, reduce the computational complexity, and prevent overfitting at the same time.

[0146] After the GAP layer, batch normalization (English full name: Batch Normalization, abbreviation: BN) is used to accelerate the convergence speed and improve the model accuracy. The BN layer estimates the sample mean and variance of the entire training dataset through moving average, learns the appropriate offset and scale, and uses them during prediction to obtain a definite output.

[0147] After the BN layer, Alpha-dropout is used to ensure that the self-normalizing property of the previous BN layer remains unchanged. At the same time, Alpha-dropout is used together with the SeLU activation function to ensure that the mean and variance of the input are maintained at their original values, preventing overfitting during training. During each forward call, the retained elements are random and will be scaled and shifted with the mean and variance remaining unchanged. At the same time, using the SeLU activation function can ensure the self-normalizing property of the attenuation layer. Finally, the Softmax activation function (normalized exponential function) is added to convert the prediction result into a probability output. During training, Lazy Adam is used as the model optimizer. Based on Adam's calculation of the first-order moment estimate and second-order moment estimate of the gradient to design independent adaptive learning rates for different parameters, it can further update the first-order momentum and second-order momentum according to the current batch at each iteration, and only update the moving average accumulator of the sparse variable index that appears in the current batch.

[0148] Example 2

[0149] Refer to Figures 1-9 , a joint modulation recognition method based on the attention mechanism and residual structure, includes the following steps:

[0150] Step A: For the received signal type to be recognized, perform blind estimation detection of the signal-to-noise ratio, and estimate the signal subspace dimension using eigenvalue decomposition to estimate the signal-to-noise ratio of the received signal;

[0151] Step B: In the high signal-to-noise ratio range, design a superimposed and optimized convolutional neural network BoTAMCNet. Using the basic structure of ResNet18, replace the standard convolution with depthwise separable convolution, and replace the original convolution of the two residual units in the residual block with a multi-head self-attention (English full name Multi-Head Self Attention, abbreviation MHSA) module, so as to superimpose the residual structure for feature extraction, and then use global depth convolution operation for feature reconstruction to increase the feature dimension;

[0152] Step C: In the low signal-to-noise ratio range, design a Transformer architecture, add different fully connected layers and activation functions for optimization on the basis of the Transformer Block, and use the multi-head self-attention mechanism for recognition, and the attention layer judges the importance of different regions in the input sequence;

[0153] Step D: Binary classify the received signals blindly estimated by SNR in Step A into two parts: high SNR (0 dB to 20 dB) signals and low SNR signals (-20 dB to 0 dB). The high SNR signals are input into BoTAMCNet for recognition, and the low SNR signals are input into Transformer for recognition.

[0154] In this embodiment, the specific steps in Step A are as follows:

[0155] A1: Construct the autocovariance matrix of the received signal and perform eigenvalue decomposition to extract eigenvalues;

[0156] A2: Estimate the dimension of the signal subspace through the Minimum Description Length (MDL) criterion;

[0157] A3: Estimate the signal power and noise power through the dimension of the signal subspace, thereby estimating the SNR.

[0158] In this embodiment, the specific steps in A1 are as follows:

[0159] Let the received signal be y(n), that is: y(n) = x(n) + σ(n). Construct the M-order autocovariance matrix R of the received signal and decompose it as:

[0160]

[0161] where I is the M-order identity matrix, is the (M-k)-fold eigenvalue of R, R ss is the autocorrelation matrix of the modulation signal. Let its rank be d (d < M), and perform eigenvalue decomposition on it to obtain: R ss = U∑U H , where U is the orthogonal matrix composed of the eigenvectors of R ss , its diagonal elements are the eigenvalues of R ss . So the diagonal elements of are the eigenvalues of R:

[0162]

[0163] In this embodiment, the specific steps in A2 are as follows:

[0164] Estimate the dimension of the signal subspace through the Minimum Description Length (MDL) criterion. First, define the test function:

[0165]

[0166] The MDL criterion combines the log-likelihood function with a penalty function and uses the penalty function for bias correction to make the MDL criterion consistent:

[0167]

[0168] Therefore, the dimension of the signal subspace can be estimated as:

[0169]

[0170] In this embodiment, the specific content of A3 is as follows:

[0171] From the estimated value of the dimension of the signal subspace, the estimated noise power can be obtained as:

[0172]

[0173] The estimated signal power is:

[0174]

[0175] The ratio of the above estimated signal power to the estimated noise power is the estimated value of the signal-to-noise ratio:

[0176]

[0177] The above is the principle of the signal-to-noise ratio blind estimation algorithm. In the joint structure proposed by the present invention, this algorithm is used to divide 24 kinds of signal data with a signal-to-noise ratio of -20 dB to 0 dB in the dataset into low signal-to-noise ratio signals by setting a threshold and input them into the Transformer for recognition; 24 kinds of signal data with a signal-to-noise ratio of 0 dB to 20 dB in the dataset are divided into high signal-to-noise ratio signals and input into the BoTAMCNet for recognition.

[0178] In this embodiment, the specific steps in step B are as follows:

[0179] B1: Deploy a standard convolutional layer with 64 convolutional kernels at the beginning of the network to extract sufficient initial information from the input signal;

[0180] B2: Use the depthwise separable convolution to introduce the skip connection method to design a residual structure, and multiple DSC residual blocks are used in series;

[0181] B3: In the residual structure designed in step B2, use the multi-head self-attention module to replace the original convolution between two residual units in the last residual block to extract more effective feature information;

[0182] B4: After the DSC operations in steps B2 and B3, global depth convolution operation is used for feature reconstruction to learn the importance of features at different positions, and then through the ReLU activation function and bias term to enhance the classification ability of the reconstructed features.

[0183] In this embodiment, the specific steps in B2 are as follows:

[0184] In the feature extraction part, multiple DSC residual blocks are used in series. At the beginning of the DSC residual block, a linear 1×1 convolution is deployed to perform channel (feature) fusion on the output of the maximum pooling of the previous residual block, and a maximum pooling operation is deployed at the end of the residual block to reduce the feature dimension.

[0185] The DSC operation process is completed in two steps: In the first step, the depth convolution layer performs convolution operations on each input channel separately; in the second step, the point convolution layer fuses all channels. The point convolution layer is a standard convolution layer with a kernel height and width of 1. The size of a single convolution kernel is 1×1×3. The input and output of a single point convolution kernel operation are consistent in the height and width dimensions, but the output becomes a single-channel 8×8×1 form.

[0186] In this embodiment, the specific steps in B3 are as follows:

[0187] In the feature extraction part, 4 DSC residual blocks are deployed. In the fourth DSC residual block, the MHSA module is used to replace the original convolution. After the convolution of two residual units, they are added after passing through the MHSA module, and finally through maximum pooling. At the same time, a layer of standard convolution is deployed in front of and behind the 4 DSC residual blocks respectively. The number of convolution kernels in the first two DSC residual blocks is designed to be 32. In order to increase the feature dimension and extract sufficient information from the current input for feature extraction, the number of convolution kernels in the first standard convolution layer (First Conv) and the last two DSC residual blocks are both designed to be 64, that is, the number of features is increased from 32 classes to 64 classes.

[0188] In this embodiment, the specific steps in B4 are as follows:

[0189] After the DSC operation, in the feature reconstruction part, BoTAMCNet proposes to use the GDWConv operation to learn the importance of features at different positions. The global depth convolution layer is a depth convolution layer, and its convolution kernel dimension is equal to the size of the input feature map. The output of the global depth convolution layer can be expressed as:

[0190]

[0191] In the formula is the feature map (full English name Last Feature Map, abbreviation LFM), and its size is W×H×M; is the global depth convolution kernel with a size of W×H×M; is the bias term; is the output classification feature vector with a size of 1×1×M; in addition, (i, j) represents the spatial position index; m represents the channel index. Further, the feature reconstruction operation occurs between and the corresponding channels. For a specific channel m, first multiply the values at the same position (i, j) of and , then sum all the multiplication results and add the bias term Finally, the output is obtained through the ReLU activation function to get a classification feature.

[0192] After feature reconstruction, an FC layer activated by Softmax (normalized exponential function) is added to convert the prediction result into a probability output. During the training process, Lazy Adam is used as the model optimizer instead of the commonly used SGD optimizer because in this structure, the SGD optimizer is prone to falling into local minima and has a slow descent speed. Lazy Adam can handle sparse updates more effectively.

[0193] In this embodiment, the specific steps in step C include the following steps:

[0194] C1: In the Transformer, the encoder maps the input sequence (x1,..., x n ) represented by symbols to a sequence z = (z1,..., z n ) of continuous representations;

[0195] C2: The decoder generates a symbol output sequence (y1,..., y m ) one element at a time. In each step, the model is autoregressive, and when generating the next step, the previously generated symbol is used as an additional input;

[0196] C3: The Transformer follows the overall architecture of C2 and C3, and uses stacked self-attention and point-wise fully connected layers in the encoder and decoder;

[0197] C4: Global average pooling, batch normalization, Alpha-dropout, etc. are added on the basis of the Transformer Block to ensure various characteristics of the structure.

[0198] In this embodiment, the specific content in C1 includes the following:

[0199] The encoder consists of a stack of N = 6 identical layers. Residual connections are used around each of the two sub-layers and layer normalization is performed. The output of each sub-layer is LayerNorm(x + Sulayer(x)), where Sulayer(x) is a function implemented by the sub-layer itself. To optimize these residual connections, all sub-layers in the model as well as the embedding layer produce outputs of dimension d model = 512.

[0200] In this embodiment, the specific content of C2 includes the following:

[0201] The decoder also consists of a stack of N = 6 identical layers. In addition to the two sub-layers in each encoder layer, the decoder inserts a third sub-layer that performs multi-head attention on the output of the encoder. Similar to the encoder, residual connections are used around each sub-layer and layer normalization is performed. At the same time, a masking layer that guarantees sequence information is added to ensure that the prediction at position i can only depend on the known outputs at positions less than i.

[0202] In this embodiment, the specific content of C3 includes the following:

[0203] The input sequence passes through the attention mechanism, the weights are calculated by the compatibility function, and the output is a weighted sum, which can be expressed by the following formula:

[0204]

[0205] Among them, K, Q, and V in the attention function represent Key, Query, and Value respectively, and d k is the number of columns of the Q and K matrices, that is, the vector dimension. The attention function maps the Query and a set of key-value pairs to the output, where Query, Key, Value, and Output are all vectors. The output is a weighted sum of the Values, and the weight assigned to each value is calculated by the correlation function between the Query and the current Key.

[0206] In addition to the attention sub-layer, each layer of the encoder and decoder contains a fully connected feed-forward network, which consists of two linear transformations connected by the ReLU activation function.

[0207] FFN(x) = max(0, xW1 + b1)W2 + b2,

[0208] The linear transformation at each position is the same, but the parameters between different layers are different. The input and output dimensions of the network are both d model = 512, but the dimension of the intermediate layer is d ff = 2048.

[0209] In this embodiment, the specific content of C4 is as follows:

[0210] In the Transformer architecture used in the present invention, based on the Transformer Block, global average pooling (English full name: Global Average Pooling, abbreviation: GAP) is first added. After the traditional max pooling layer, one or n fully connected layers are required before the softmax classification can be utilized. Here, GAP is used instead of max pooling, and GAP is used to replace the fully connected layer, greatly reducing the parameters in the network and lowering the computational complexity.

[0211] After the GAP layer, batch normalization (English full name: Batch Normalization, abbreviation: BN) is used. The BN layer estimates the sample mean and variance of the entire training dataset through moving average, learns suitable offsets and scalings, and uses them during prediction to obtain a definite output.

[0212] After the BN layer, Alpha-dropout is used to ensure that the self-normalizing property of the previous BN layer remains unchanged. At the same time, Alpha-dropout is used together with the SeLU activation function to ensure that the mean and variance of the input are maintained at their original values, preventing overfitting during the training process. During the training process, Alpha-Dropout will set some elements to zero with a probability sampled from the Bernoulli distribution, but does not delete the weights on the path. During each forward call, the remaining elements are random and will be scaled and shifted while the mean and variance remain unchanged. At the same time, using the SeLU activation function can ensure the self-normalizing property of the attenuation layer. Finally, the Softmax activation function (normalized exponential function) is added to convert the prediction result into a probability output. During the training process, Lazy Adam is used as the model optimizer. Based on the fact that Adam designs independent adaptive learning rates for different parameters by calculating the first-order moment estimate and second-order moment estimate of the gradient, it is further able to update the first-order momentum and second-order momentum according to the current batch at each iteration, and only update the moving average accumulator of the sparse variable index that appears in the current batch.

[0213] As shown in Table 1, it is the parameter configuration of the BoTAMCNet proposed by the present invention.

[0214] Table 1

[0215]

[0216] As shown in Table 2, it is the parameter configuration of the network after the Transformer proposed by the present invention includes the above C1, C2, C3, and C4.

[0217] Table 2

[0218]

[0219] As shown in Table 3, these are the dataset parameters used in the algorithm performance simulation of the present invention.

[0220] Table 3

[0221]

[0222] The algorithm performance simulation is as Figure 8 、 Figure 9 and shown in Table 4 and Table 5.

[0223] Table 4

[0224]

[0225] Table 5

[0226]

[0227] As shown in Table 4, the blind SNR estimation algorithm is used to perform blind SNR estimation on each signal data in the sampled dataset. Through threshold setting for binary classification, it is divided into low SNR signals (-20dB to 0dB) and high SNR signals (0dB to 20dB), and they are respectively input into Transformer and BoTAMCNet for signal classification to complete the recognition of various signal modulation types by the joint structure. The binary classification accuracy is shown in Table 4.

[0228] From the data in Table 4, it can be seen that the estimation performance of this algorithm for signals near -2dB and 0dB SNR is poor, and there is an aliasing with the SNR estimation value of signals at 0 - 1dB. However, the recognition rates of the two networks in the joint structure at -2dB and 0dB are not very different, and the impact on the overall recognition effect of the joint structure is relatively small.

[0229] As Figure 8As shown, it is the performance comparison of BoTAMCNet, Transformer and other networks on the RadioML2018.01A dataset. The recognition of the dataset by BoTAMCNet achieved better recognition results in the high signal-to-noise ratio range. Under the condition of a signal-to-noise ratio of 20 dB, the recognition accuracy was as high as 97.21%, which was 1.14% - 7.41% higher than other networks under the same conditions. Transformer had better recognition results for the RadioML2018.01A dataset in the low signal-to-noise ratio range. The recognition accuracy reached over 60% in the range from -20 dB to 0 dB, with the highest reaching 65.49%, showing a significant improvement in recognition accuracy. However, in the high signal-to-noise ratio range, Transformer had poor classification effects for various signals of the dataset. Therefore, the present invention adopts a combined structure, using Transformer for recognition in the low signal-to-noise ratio range and BoTAMCNet for recognition in the high signal-to-noise ratio range, which can achieve a relatively high recognition rate in the overall signal-to-noise ratio range from -20 dB to 20 dB.

[0230] As shown in Table 5, it is the comparison of the model complexity between the BoTAMCNet network and other networks, which is reflected by the number of trainable parameters of the network and the inference time. By using depthwise separable convolution to replace standard convolution, controlling the number of convolutional kernels and using MHSA to replace convolution, the model complexity of BoTAMCNet has been significantly reduced.

[0231] Under the same conditions, the recognition accuracy of BoTAMCNet reached the highest, and the number of trainable parameters of the network was reduced by about 18% - 30% compared with other solutions except MRNN. Compared with the ResNet network, which achieved the highest recognition accuracy among other networks, BoTAMCNet saved about 10.2% of the inference time.

[0232] As Figure 9As shown, it is the performance comparison of the joint structure proposed by the present invention and other different network structures on the RadioML2018.01A dataset. It includes the classic CNN structure (CNN / VGG), the residual neural network (ResNet), the modulation classification convolutional neural network (MCNet) specially designed for the modulation recognition task, and the multi-hop residual neural network (MRNN). It can be seen from this that: (1) BoTAMCNet in the joint structure has the highest accuracy at high signal-to-noise ratios, with an accuracy improvement of 1.14%-7.41% compared with other networks at 20 dB; (2) The classification accuracy of Transformer in the joint structure has been significantly improved compared with other algorithms at low signal-to-noise ratios, and the classification accuracy reaches over 60% in the signal-to-noise ratio range from -20 dB to 0 dB; (3) The overall classification accuracy of the VGG network is the worst because its model design is relatively simple and the number of convolutional layers is small, making it difficult to extract the best features, and the computational cost of feature extraction using standard convolutional operations is too large; (4) The performance of the MRNN model is generally better than that of VGG and MCNet, but the classification performance is still inferior to that of ResNet and the joint structure.

[0233] As can be seen from Figure 9 it that the joint structure for modulation type recognition proposed by the present invention has achieved good recognition effects in the overall low signal-to-noise ratio to high signal-to-noise ratio range. The recognition accuracy has been greatly improved in the low signal-to-noise ratio range, reaching up to 65.49% at most. In the high signal-to-noise ratio range, while reducing the computational complexity, the accuracy has also been improved. The recognition accuracy is as high as 97.21% under the condition of 20 dB. The overall recognition rate of the joint structure is higher than that of other existing algorithms, achieving higher recognition performance.

[0234] For the specific calculation process of the joint modulation recognition method system based on the attention mechanism and the residual structure, reference can be made to the above embodiments, and the embodiments of the present invention will not be elaborated here.

[0235] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features, but these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A joint modulation recognition method based on the attention mechanism and the residual structure, characterized in that It includes the following steps: Step A: For the received signal type to be recognized, perform blind SNR estimation detection, and use eigenvalue decomposition to estimate the signal subspace dimension, thereby estimating the SNR of the received signal; Step B: In the high SNR interval, design a superimposed optimized convolutional neural network BoTAMCNet. Use the basic structure of ResNet18, replace the standard convolution with depthwise separable convolution, and replace the original convolutions of two residual units in the residual block with a multi-head self-attention module. In this way, superimpose the residual structure for feature extraction, and then use global depth convolution operation for feature reconstruction to increase the feature dimension; Step C: In the low SNR interval, design a Transformer architecture. Optimize it by adding different fully connected layers and activation functions on the basis of the Transformer Block, and use the multi-head self-attention mechanism for recognition. The attention layer judges the importance of different regions in the input sequence; Step D: Binarize and classify the received signal blindly estimated by SNR in Step A into two parts: high SNR signal and low SNR signal. The high SNR signal is input into BoTAMCNet for recognition, and the low SNR signal is input into Transformer for recognition.

2. The joint modulation recognition method based on the attention mechanism and the residual structure according to claim 1, wherein The specific steps of Step A include the following steps: A1: Construct the autocovariance matrix of the received signal and perform eigenvalue decomposition to extract eigenvalues; A2: Estimate the dimension of the signal subspace through the MDL criterion; A3: Estimate the signal power and noise power through the dimension of the signal subspace, thereby estimating the SNR.

3. The joint modulation recognition method based on the attention mechanism and the residual structure according to claim 2, wherein The specific content of A1 includes the following: Let the received signal be y(n), that is: y(n) = x(n) + σ(n), construct the M-order autocovariance matrix R of the received signal and decompose it as: where \(I\) is an \(M\times M\) identity matrix, is an eigenvalue of multiplicity \((M - k)\) of \(R\), and \(R\) ss is the autocorrelation matrix of the modulation signal. Let its rank be \(d(d < M)\). Perform eigenvalue decomposition on it to obtain: \(R\) ss \(= U\sum U\) H , where \(U\) is an orthogonal matrix composed of the eigenvectors of \(R\) ss . Its diagonal elements are the eigenvalues of \(R\) ss . Therefore, the diagonal elements of are the eigenvalues of \(R\):

4. The joint modulation recognition method based on the attention mechanism and the residual structure according to claim 3, characterized in that, The specific content of A2 includes the following: Estimate the dimension of the signal subspace through the MDL criterion. First, define the test function: The MDL criterion combines the log-likelihood function and the penalty function, and uses the penalty function for bias correction to make the MDL criterion consistent: The dimension of the signal subspace can be estimated as:

5. A joint modulation recognition method based on the attention mechanism and the residual structure according to claim 4, characterized in that The specific content of A3 includes the following: The estimated noise power can be obtained from the estimated value of the dimension of the signal subspace as: The estimated signal power is: The ratio of the above estimated signal power to the estimated noise power is the SNR estimated value: The above is the principle of the SNR blind estimation algorithm. In the joint structure, use this algorithm to divide 24 kinds of signal data with SNR from -20dB to 0dB in the dataset into low SNR signals by setting a threshold and input them into Transformer for recognition; divide 24 kinds of signal data with SNR from 0dB to 20dB in the dataset into high SNR signals and input them into BoTAMCNet for recognition.

6. The joint modulation recognition method based on the attention mechanism and the residual structure according to claim 5, wherein, The specific steps of Step B include the following steps: B1: Deploy a standard convolutional layer with 64 convolutional kernels at the beginning of the network to extract sufficient initial information from the input signal; B2: Use the depthwise separable convolution to introduce the skip connection method to design the residual structure, and multiple DSC residual blocks are used in series; B3: In the residual structure designed by B2, the multi-head self-attention module is used to replace the original convolution between two residual units in the last residual block to extract more effective feature information; B4: After the DSC operations of B2 and B3, global depth convolution operation is used for feature reconstruction to learn the importance of features at different positions, and then through the ReLU activation function and bias term to enhance the classification ability of the reconstructed features.

7. A joint modulation recognition method based on attention mechanism and residual structure according to claim 6, characterized in that, Specifically, B2 includes the following: In the feature extraction part, multiple DSC residual blocks are used in series. At the beginning of the DSC residual block, a linear 1×1 convolution is deployed to perform channel fusion on the output of the max pooling of the previous residual block. At the end of the residual block, a max pooling operation is deployed to reduce the feature dimension. The DSC operation process is completed in two steps: the first step is that the depth convolution layer performs convolution operations on each channel of the input separately; the second step is that the point convolution layer fuses all channels. In the depth convolution layer, the depth convolution kernel is single-channel and the number of convolution kernels is equal to the number of input channels, and the height and width still use 5×5; the point convolution layer is a standard convolution layer with a convolution kernel height and width of 1, and the size of a single convolution kernel is 1×1×3. The input and output of a single point convolution kernel operation are consistent in the height and width dimensions, but the output becomes a single-channel 8×8×1 form.

8. A joint modulation recognition method based on the attention mechanism and the residual structure according to claim 6, characterized in that Specifically, B3 includes the following: In the feature extraction part, 4 DSC residual blocks are deployed. In the fourth DSC residual block, the MHSA module is used to replace the original convolution. After the convolution of two residual units, the results are added after passing through the MHSA module, and finally max pooling is performed; at the same time, a layer of standard convolution is deployed in front of and behind the 4 DSC residual blocks. The number of convolution kernels in the first two DSC residual blocks is designed to be 32 to extract sufficient information from the current input for feature extraction. The number of convolution kernels in the first standard convolution layer and the last two DSC residual blocks is designed to be 64, that is, the features are upgraded from 32 classes to 64 classes.

9. The joint modulation recognition method based on the attention mechanism and the residual structure according to claim 6, wherein Specifically, B4 includes the following: After the DSC operation, in the feature reconstruction part, BoTAMCNet proposes to use the GDWConv operation to learn the importance of features at different positions, and then through the ReLU activation function and bias term to enhance the classification ability of the reconstructed features. The global depth convolution layer is a depth convolution layer, and its convolution kernel dimension is the same as the size of the input feature map. The output of the global depth convolution layer can be expressed as: where is the feature map with a size of W×H×M; is the global depth convolution kernel with a size of W×H×M; is the bias term; is the output classification feature vector with a size of 1×1×M; in addition, (i, j) represents the spatial position index; m represents the channel index; the global depth convolution kernel has the same size as the of the input LFM, and the feature reconstruction operation occurs between the and corresponding channels; for a specific channel m, first multiply the values at the same position (i, j) of and , then sum all the dot product results and add the bias term Finally, after passing through the ReLU activation function, a classification feature is output. After feature reconstruction, an FC layer activated by Softmax is added to convert the prediction result into a probability output. During the training process, LazyAdam is used as the model optimizer.

10. A joint modulation recognition method based on the attention mechanism and the residual structure according to claim 9, characterized in that The specific steps of step C are as follows: C1: In the Transformer, the encoder maps the input sequence (x1,..., x n ) of symbolic representations to a sequence z = (z1,..., z n ); C2: The decoder generates a sequence of symbol outputs (y1,..., y m ) one element at a time. At each step, the model is autoregressive and uses the previously generated symbols as additional input when generating the next step; C3: Transformer follows the overall architecture of C2 and C3, and uses stacked self-attention and point-wise fully connected layers of the encoder and decoder in it; C4: Global average pooling, batch normalization, and Alpha-dropout are added on the basis of the Transformer Block to ensure various characteristics of the structure; Specifically, C1 includes the following: The encoder consists of a stack of N = 6 identical layers, each layer having two sub-layers. The first layer is the multi-head self-attention mechanism, and the second layer is a simple fully-connected feed-forward network. Residual connections are used around each of the two sub-layers and layer normalization is performed. The output of each sub-layer is LayerNorm(x + Sulayer(x)), where Sulayer(x) is the function implemented by the sub-layer itself. All sub-layers in the model, as well as the embedding layer, produce outputs of dimension d model = 512; Specifically, C2 includes the following: The decoder is also composed of a stack of N = 6 identical layers. In addition to the two sub-layers in each encoder layer, a third sub-layer is inserted into the decoder to perform multi-head attention on the output of the encoder. Similar to the encoder, residual connections are used around each sub-layer and layer normalization is performed. At the same time, a masking layer that guarantees sequence information is added to ensure that the prediction at position i can only depend on the known outputs at positions less than i; Specifically, C3 includes the following: The input sequence passes through the attention mechanism, and the weights are calculated through the compatibility function. The output is a weighted sum, which can be expressed by the following formula: Among them, K, Q, and V in the attention function represent key, query, and value respectively, and d k is the number of columns of the Q and K matrices, that is, the vector dimension. The attention function maps a Query and a set of key-value pairs to an output. Among them, Q, K, V, and Output are all vectors, and the output is a weighted sum of V. The weight assigned to each value is calculated by the correlation function of Q and the current K; Except for the attention sub-layer, each layer of the encoder and decoder contains a fully connected feed-forward network, which consists of two linear transformations connected by the ReLU activation function: FFN(x) = max(0, xW1 + b1)W2 + b2, The linear transformation at each position is the same, but the parameters between different layers are different; the input and output dimensions of the network are both d model = 512, but the dimension of the middle layer is d ff = 2048; Specifically, C4 includes the following: On the basis of the Transformer Block, the Transformer architecture first adds global average pooling to enhance the structural stability. After the traditional max pooling layer, 1 or n fully connected layers are required before the softmax classification can be used, which is likely to cause overfitting. Using global average pooling instead of the fully connected layer greatly reduces the parameters in the network, reduces the computational complexity, and prevents overfitting; Batch normalization is used after the global average pooling layer to accelerate the convergence speed and improve the model accuracy. The batch normalization layer estimates the sample mean and variance of the entire training dataset through the moving average, learns the appropriate offsets and scales, and uses them during prediction to obtain a definite output; After the batch normalization layer, Alpha-dropout is used to ensure that the self-normalization characteristics of the previous batch normalization layer remain unchanged. At the same time, Alpha-dropout is used together with the SeLU activation function to ensure that the mean and variance of the input are maintained at their original values; during the training process, Alpha-Dropout will set some elements to zero with a probability sampled from the Bernoulli distribution, but does not delete the weights on the path; during each forward call, the remaining elements are random and will be scaled and shifted with the mean and variance unchanged. At the same time, using the SeLU activation function can ensure the self-normalization characteristics of the attenuation layer. Finally, the Softmax activation function is added to convert the prediction result into a probability output. During the training process, Lazy Adam is used as the model optimizer. On the basis that Adam designs independent adaptive learning rates for different parameters by calculating the first-order moment estimate and second-order moment estimate of the gradient, it can further update the first-order momentum and second-order momentum according to the current batch at each iteration, and only update the moving average accumulator of the sparse variable indexes that appear in the current batch.

Citation Information

Patent Citations

  • Neural network design method for signal modulation type recognition

    CN113657491A

  • Automatic identification method for radar signal modulation type

    CN114564982A