Modulation identification method and gradient-based screening smooth score class activation map interpretation method
Patent Information
- Application Number
- CN202610984254.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-03
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2046-07-03
AI Technical Summary
Bai等人提出结合传统FB方法与代理模型的调制识别可解释性方法,通过训练局部决策树模型将复杂决策转化为If-Then判决规则,但该方法高度依赖人工特征的完备性,当调制类型增加时受到约束,且决策树作为代理模型在拟合非线性边界时存在保真度下降的风险
[0036](1)设计了一种改进的ResNet-18调制识别网络,通过重构输入层、剪裁冗余深层模块以及嵌入自适应空间注意力机制,增强了对SPWVD时频图像中细微特征的提取能力。
Smart Images

Figure CN122533895B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of electrical digital signal processing technology, and relates to the identification and model interpretation of modulated signals, specifically to a modulation identification method and a gradient-filtered smooth fractional activation mapping interpretation method. Background Technology
[0002] Modulation signal recognition (MSR) is a key technology in wireless communication signal processing, with important applications in military reconnaissance, spectrum monitoring, and electromagnetic countermeasures. With the rapid development of deep learning technology, MSR methods based on convolutional neural networks have achieved recognition accuracy far exceeding that of traditional algorithms. However, deep neural networks are often considered "black box" models, with opaque decision-making logic and unreliable results, severely restricting their deployment and application in communication systems with extremely high reliability requirements. Therefore, this paper designs an interpretation method for MSR networks. After revealing the key time-frequency features upon which the network's classification decisions depend, these features can be used as prior criteria for evaluating the importance of network channels. This guides the structured pruning and model lightweighting of complex neural networks, accurately eliminating redundant parameters that do not contribute to the decision or cause interference. Simultaneously, by locating the interference time-frequency regions marked on the heatmap when the model misclassifies based on the interpretation results, further guidance can be provided for the formulation of data augmentation or feature enhancement strategies.
[0003] Currently, research on the interpretability of deep neural networks is mainly divided into two categories: self-explainable models and ex post-interpretive models. Self-explainable models embed specific algorithms within the network, enabling it to generate explanations for the decision-making process while making decisions. Typical self-explainable models include linear regression and decision trees; these algorithms are simple in form and are often used to explain linear relationships or white-box models. However, the fitting accuracy of simple self-explainable models is usually limited, making it difficult to adapt to the high-dimensional, nonlinear feature distributions in complex modulation recognition tasks. This is especially true given the trend towards miniaturization and intelligence in modern devices, where the problems of complex networks become more pronounced. Ex post-interpretive models provide posterior explanations for already trained black-box models, offering a wider range of applications, and are currently the focus of most interpretability research.
[0004] Because modulated signals possess unique properties such as time-series correlation, frequency domain characteristics, and channel sensitivity, and because practical application scenarios with low signal-to-noise ratios and multipath fading demand higher reliability and stability of interpretation results, although the aforementioned general interpretability methods have matured in fields such as computer vision and natural language processing, they face unique challenges when transferred to the field of wireless communication signal processing.
[0005] Currently, research on the interpretability of modulation recognition networks, both domestically and internationally, is still in its early exploratory stages. Class Activation Mapping (CAM) and its improved methods have attracted considerable attention due to their intuitive visual interpretation effects. Huang et al. visualized modulation features by introducing class activation vectors, revealing the differences in features of interest between CNNs and LSTMs using Grad-CAM and mask-based optimization methods, respectively. However, the interpretation methods are inconsistent and lack quantitative evaluation standards. Chen et al. used Grad-CAM to obtain feature regions that promote classification in the input image and mapped them onto two-dimensional amplitude and phase maps and constellation maps. However, the input is a two-dimensional raw sampling point, which may lack actual physical connections, and objective quantitative evaluation indicators are still lacking. Bai et al. proposed a modulation recognition interpretability method combining traditional FB methods and surrogate models. By training a local decision tree model, complex decisions are transformed into If-Then decision rules. However, this method is highly dependent on the completeness of manually generated features, which is constrained when the modulation type increases. Furthermore, the decision tree as a surrogate model risks a decrease in fidelity when fitting nonlinear boundaries. Summary of the Invention
[0006] To address the shortcomings of existing technologies, this invention proposes a modulation recognition method and a gradient-filtered smooth fractional class activation map interpretation method. Based on a gradient filtering mechanism, invalid feature channels are eliminated, and only these channels are subjected to smooth fractional calculations to generate high-resolution class activation heatmaps. This identifies key time-frequency features upon which the model's decisions depend, guiding the lightweight pruning of the modulation recognition network structure by providing parameter removal, data augmentation, or feature enhancement. This solves the problems of existing class activation map interpretability methods struggling to balance heatmap fidelity and feature localization accuracy, and exhibiting computational redundancy and key feature ambiguity in resource-constrained scenarios.
[0007] Firstly, the present invention proposes a modulation recognition method, the specific steps of which are as follows:
[0008] Step 1: Time-frequency feature extraction based on SPWVD
[0009] A smoothed pseudo-Wigner-Vell distribution (SPWVD) is used to represent the time-frequency characteristics of the modulated signal. By simultaneously smoothing the received signal x(t) in both time t and frequency direction f, a two-dimensional feature representation I0 is generated.
[0010]
[0011] in, For time delay variables, and These represent the window functions added in the time domain and frequency domain, respectively. This is the instantaneous autocorrelation function.
[0012] Step 2: Design a modulation recognition network
[0013] Construct a ResNet-18 network, replace the convolution kernel size with 3*3 in the input layer, adjust the stride to 1, adjust the number of padding layers to 1, and replace the max pooling layer in the original network with an identity mapping.
[0014] Remove the last feature extraction layer. Each feature extraction layer contains two residual blocks, and each residual block includes a 3x3 convolutional layer, a batch normalization layer, and an activation layer.
[0015] A lightweight spatial attention module is added after the feature extraction layer. The multi-channel features are compressed into a single-channel attention weight map through 1*1 convolution. The weights are then mapped to the [0,1] interval by the Sigmoid activation function. Finally, the weights are multiplied element-wise with the original feature map to enhance the features of key regions.
[0016] Step 3: Network Training and Modulation Recognition
[0017] The improved ResNet-18 from step 2 is used as the modulation recognition network, with the two-dimensional feature representation I0 of the modulation signal as input, to predict the modulation method. The cross-entropy loss function is used to measure the difference between the model's predicted distribution and the true label distribution.
[0018] Secondly, this invention proposes a gradient-filtered smooth fractional activation mapping interpretation method:
[0019] Calculate the prediction score F of the trained modulation recognition network F() for the target modulation category c. c (I0) Feature map A of the k-th channel relative to the output of the target convolutional layer k gradient The gradient weights are then aggregated using the global average pooling operation GAP() to obtain channel-level gradient weights. :
[0020] k=1,2,…n
[0021]
[0022] Where I0 represents the SPWVD time-frequency image input to the modulation recognition network, and n represents the total number of channels in the target convolutional layer.
[0023] Configure a gradient filtering mechanism to retain only... If a feature channel > 0 and possesses an activation response, a set of feature maps S with a positive impact on decision-making is obtained:
[0024]
[0025] In sequence, each feature map S in the feature map group Sk The input mask M is obtained by upsampling and normalizing the samples respectively. k Then in the mask M k Based on this, random noise sampled from a Gaussian distribution is superimposed. To generate N smooth masks with tiny perturbations. :
[0026] m=1,2,…N
[0027] After the final masked image is fed back into the modulation recognition network, its Channel-wise Increase of Confidence (CIC) score is calculated to measure feature importance. :
[0028]
[0029] In the formula, This is the final mask input. b It is an all-zero matrix with the same size as the I0 space, representing the reference state without any time-frequency characteristic input.
[0030] For a feature map S k N smooth masks The average of the N CIC scores is used as the final weight for the k-th feature channel. :
[0031]
[0032] Finally, the weights With the corresponding original feature map S k Perform linear weighted summation and retain the positive contribution using the ReLU function to generate an activation-like heatmap. :
[0033]
[0034] Based on the core feature region of the class activation heatmap when the model makes a correct decision, the contribution of each channel to modulation recognition is quantified. Redundant channels with low contribution in the modulation recognition network are pruned according to the target pruning rate. The performance of the pruning model is restored by fine-tuning, resulting in a pruning model that balances recognition accuracy and lightweight design. The SPWVD time-frequency image of the modulation signal is then input into the pruning model for modulation recognition.
[0035] The present invention has the following beneficial effects:
[0036] (1) An improved ResNet-18 modulation recognition network was designed. By reconstructing the input layer, pruning redundant deep modules, and embedding an adaptive spatial attention mechanism, the ability to extract subtle features from SPWVD time-frequency images was enhanced.
[0037] (2) A gradient-based smooth fractional activation mapping method, SS-CAM+, is proposed. This method utilizes gradient information to filter positively contributing feature channels, effectively reducing the computational cost of heatmap generation while retaining the high robustness of SS-CAM, and significantly improving the feature localization accuracy of the visualization results. The localization results can be used to identify time-frequency feature regions that the model's decision depends on, thereby supporting model reliability verification, structured pruning, and misjudgment tracing, enabling the interpretation results to directly serve the engineering optimization of the modulation recognition system.
[0038] (3) The quantitative evaluation index verified the comprehensive advantages of SS-CAM+ in terms of interpretation accuracy and running time. The generation time of a single heat map was shortened by about 58.9% compared with SS-CAM, providing a new technical path for the study of the interpretability of modulation recognition networks. Attached Figure Description
[0039] Figure 1 This is a deep learning-based modulation recognition framework;
[0040] Figure 2 The overall framework of the SS-CAM+ method;
[0041] Figure 3 The time-domain waveforms of eight digital modulation signals in the RadioML 2016.10a dataset at 0dB are shown.
[0042] Figure 4 SPWVD time-frequency plots of 8 digital modulation signals at 0dB;
[0043] Figure 5 For residual learning framework;
[0044] Figure 6 Flowchart of the SS-CAM+ method;
[0045] Figure 7 SS-CAM+ heatmaps for different modulation signals at 6dB;
[0046] Figure 8 Heatmaps showing four interpretation methods for different modulation signals;
[0047] Figure 9 A heatmap comparing the correct and incorrect modulation type decisions made by the model;
[0048] Figure 10A comparison of AD and IC indices for four interpretation methods under different sample sizes. Detailed Implementation
[0049] The present invention will be further explained below with reference to the accompanying drawings;
[0050] Deep learning-based modulation recognition of communication signals is essentially a pattern recognition problem for modulated signals, such as... Figure 1 As shown, this process mainly includes three stages: preprocessing of the signal to be identified, feature extraction through a neural network, and classification decision. Based on this modulation recognition framework, this invention proposes a gradient-filtered smoothed score class activation mapping (SS-CAM+) method to reveal the decision logic of the model in the classification process, such as... Figure 2 As shown, firstly, a smoothed pseudo-Wigner-Ville distribution (SPWVD) is used to transform the one-dimensional signal into a two-dimensional time-frequency image; secondly, an improved ResNet-18 is used to extract features and classify them; finally, through gradient filtering and smoothing mechanisms, a class activation heatmap is generated to reveal the key time-frequency regions for model decisions. The specific steps are as follows:
[0051] Step 1: Data Preparation
[0052] The modulation signal dataset used in the experiment was RadioML 2016.10a, which contains 11 types of modulation signals: 8 digital signals (QPSK, PAM4, BPSK, 8PSK, QAM16, QAM64, GFSK, and CPFSK) and 3 analog signals (AM-DSB, AM-SSB, and WBFM). The signal-to-noise ratio (SNR) range for each modulation type is [-20dB, 18dB], with a step size of 2dB. There are 1000 samples at each SNR, for a total of 220,000 modulation signal samples. Each signal sample has a dimension of 2*128, meaning it contains two channels: a co-directional component and a quadrature component, each with 128 sampling points. This experiment used the 8 digital modulation signals from the RadioML 2016.10a dataset for neural network training and subsequent feature visualization. 80% of the dataset was used for network training, and 20% was used for testing network performance. Figure 3 The time-domain waveforms of these modulation signals at a signal-to-noise ratio of 0dB are shown.
[0053] Step 2: Time-frequency feature extraction based on SPWVD
[0054] Modulated signals exhibit time-domain and frequency-domain coupling characteristics, and traditional single-domain analysis easily leads to the loss of modulation feature information. To fully explore the deep features of the signal, a smoothed pseudo-Wigner-Vell distribution (SPWVD) is used to represent the time-frequency features of the modulated signal. Based on the WVD distribution, smoothing window functions are added in both the time and frequency domains. By simultaneously performing smoothing filtering in the time and frequency directions, deep suppression of cross-term interference is achieved, providing rich two-dimensional feature representations for deep learning models.
[0055]
[0056] Where x(t) represents the received signal, Let t be the time delay variable, and f be the time variable and frequency variable, respectively. It is the instantaneous autocorrelation function. and These represent the window functions added in the time domain and frequency domain, respectively.
[0057] Specifically, in this embodiment, Hamming windows of lengths 13 and 33 are used for time-domain smoothing and frequency-domain smoothing, respectively, and a time-frequency image of size 128*128 is generated by autocorrelation calculation and FFT. Figure 4 The SPWVD time-frequency plots of eight digital modulation signals at 0dB are shown.
[0058] Analysis of the time-frequency plots reveals corresponding time-frequency texture characteristics for different modulation signals: GFSK and CPFSK are continuous phase modulations, exhibiting continuous time-frequency ridges with concentrated energy; GFSK has a smooth trajectory, while CPFSK shows a slight wavy pattern. BPSK, QPSK, and 8PSK produce energy discontinuities and bandwidth broadening due to phase abrupt changes; as the modulation order increases, the discontinuities become denser, and the ridge continuity strengthens. PAM4 is amplitude modulation, showing discrete patches of alternating light and dark on the energy lines. QAM signals modulate both amplitude and phase simultaneously, resulting in a broadened time-frequency energy distribution; QAM16 exhibits a grid texture, while QAM64, due to its dense constellation points, has finer texture particles and a more homogeneous energy distribution. Higher-order QAM signals show less difference in characteristics, making them more prone to identification errors when data length is limited.
[0059] Step 3: Design of a Modulation Recognition Network Based on Improved ResNet
[0060] To meet the requirements of deep feature extraction of modulated signals and subsequent interpretable visualization, improvements were made to the classic ResNet-18 network, including reconstructing the input layer structure, pruning redundant deep modules, and embedding an adaptive spatial attention mechanism. This ensures that detailed features can be extracted while adapting to the input SPWVD time-frequency image.
[0061] In the standard ResNet-18 network architecture, the first convolutional layer has a 7x7 kernel size, a stride of s of 2, and 3 layers padded with zeros around the edges. It then passes through a max-pooling layer with a size of 3. Since the SPWVD time-frequency map size is 128x128x3, the original ResNet-18 network's downsampling strategy is quite aggressive, easily leading to the loss of high-frequency details before they enter deeper layers. Therefore, the improved ResNet-18 network replaces the kernel size with a 3x3 kernel in the input layer, adjusts the stride s to 1, and reduces the number of padding layers to 1. The corresponding convolutional output dimension... for:
[0062]
[0063] Among them, H in Let K be the input size and K be the kernel size. This improvement ensures that the spatial size of the feature map remains unchanged after passing through the first convolutional layer, reducing the loss of time-frequency features of the original signal. Meanwhile, since max pooling discards local non-maximum information, it may lead to excessive compression of fine time-frequency details of the modulated signal. Removing the max pooling layer from the original network and replacing it with an identity mapping further avoids spatial resolution loss.
[0064] To avoid overly abstract features and to achieve network lightweighting, reducing computational complexity and the risk of overfitting, the last feature extraction layer of the original ResNet-18 was removed. Each feature extraction layer now contains two residual blocks, such as... Figure 5 As shown, each residual block includes a 3x3 convolutional layer, a batch normalization layer, and an activation layer. For the l-th residual block, its input is defined as x. l Its output x l+1 for:
[0065] x l+1 =h(x l )+F(x l W l )
[0066] Among them, W l The weights of the convolution kernel are represented by F(), the residual function is represented by h(), and the skip connection mapping is represented by h().
[0067] To enhance the model's ability to extract key time-frequency feature regions and suppress background noise and redundant feature interference, a lightweight spatial attention module is added after the feature extraction layer. This module compresses multi-channel features into a single-channel attention weight map through 1*1 convolution, then maps the weights to the [0,1] interval using the Sigmoid activation function, and finally multiplies it element-wise with the original feature map to enhance the features of key regions.
[0068] The improved ResNet-18 was used as the modulation recognition network. The specific parameters and tensors of each layer change through the corresponding dimensions before and after the changes are shown in Table 1.
[0069] Table 1
[0070]
[0071] To train the above network, the simulation platform environment and related hardware configuration are as follows: operating system is Windows 11, CPU is AMD Ryzen 9 7940H, GPU is NVIDIA GeForce RTX 4070, GPU memory is 8GB, programming language is Python 3.9, programming software is PyCharm, and deep learning framework is PyTorch 2.5.1.
[0072] The cross-entropy loss function is used to measure the difference between the model's predicted distribution and the true label distribution.
[0073] The optimizer uses Adam, with an initial learning rate of 0.001, a first moment estimate decay of 0.9, a second moment estimate decay of 0.999, a weight decay of 0.0001, a batch size of 128, and 30 training epochs.
[0074] Step 4: SS CAM+ Interpretation Method for Modulation Recognition Network
[0075] This study proposes a smooth fractional activation mapping method, SS-CAM+, which combines a gradient selection mechanism with the SS-CAM interpretation method. The core idea of this method is to introduce a gradient-based feature map selection mechanism, prioritizing the selection of positively effective feature maps that contribute most to the classification of modulated signals, while eliminating redundant feature channels that negatively impact the model's decision-making or are irrelevant. Subsequently, the smoothing score of random noise is calculated only for the selected key feature channels. SS-CAM+ retains the high robustness of the SS-CAM method while effectively reducing the computational cost of heatmap generation and significantly improving the clarity and relevance of the visualization results, ensuring that the model can accurately capture the key feature regions relied upon for classification decisions, such as... Figure 6 As shown:
[0076] For the input SPWVD time-frequency image I0, the prediction score of the trained modulation recognition network for the target modulation category c is F. c (I0). The feature map set A = {A1, A2, ..., A...} output by the target convolutional layer. n}, where n represents the total number of channels. Calculate the target category score F.c (I0) Feature map A relative to the k-th channel k gradient :
[0077] k=1,2,…n
[0078] Able to reflect feature map A k The features of each pixel in the prediction score F c The positive or negative effects of (I0).
[0079] Gradient information for each feature channel is obtained by using the global average pooling operation GAP(). Aggregate the gradients to obtain channel-level gradient weights. This allows for the quantification of the overall contribution of each feature channel to the classification of the target category.
[0080]
[0081] like >0 indicates that the activation of this feature channel generally promotes the model to identify the input as category c, that is, the feature map contains key features of the target category; if If the value is ≤0, it indicates that the activation of this feature channel has an inhibitory effect on the prediction of the target category c, or that the channel extracts irrelevant features related to background noise or interference. A gradient filtering mechanism is set to retain only... If a feature channel > 0 and possesses an activation response, a set of feature maps S with a positive impact on decision-making is obtained:
[0082]
[0083] In sequence, each feature map S in the feature map group S k Each image is upsampled using the UP() function to ensure that the size of the subsequent mask corresponds to the input image.
[0084]
[0085] In this embodiment, the upsampling operation UP() is bilinear interpolation.
[0086] Subsequently, to ensure that the mask value is within the range of [0,1], the upsampled feature map is... Normalization is performed to obtain a standardized mask input M. k Then in the mask M k Based on this, random noise sampled from a Gaussian distribution is superimposed to generate N smooth masks with small perturbations. :
[0087] m=1,2,…N
[0088] in, This indicates that the mean is 0 and the variance is 0. Gaussian noise.
[0089] Obtaining a smooth mask Then, the final masked image is fed back into the neural network, and its Channel-wise Increase of Confidence (CIC) score is calculated to measure feature importance. :
[0090]
[0091] In the formula, For the final mask input, I b This is the baseline input.
[0092] For a feature map S k N smooth masks The average of the N CIC scores is used as the final weight for the k-th feature channel. :
[0093]
[0094] By averaging the scores of the smoothing mask, prediction fluctuations caused by local noise can be effectively smoothed, making the calculated weights more stable.
[0095] Finally, the weights With the corresponding original feature map S k Perform linear weighted summation and retain the positive contribution using the ReLU function to generate an activation-like heatmap. :
[0096]
[0097] To verify the effectiveness of the proposed SS-CAM+ interpretability method, it was compared with Grad-CAM, Score-CAM, and other interpretable methods. The SS-CAM+ interpretability method used a sampling count of 20 and a noise standard deviation of 2. First, SS-CAM+ was independently validated at a 6dB signal-to-noise ratio by inputting the SPWVD time-frequency plots of eight digitally modulated signals into a trained improved ResNet-18 model. Figure 7The characteristics of these eight modulation signals under SS-CAM+ are shown, with the red highlighted areas representing the time-frequency regions of highest interest during model decision-making. Results show that the highlighted areas of GFSK and CPFSK are closely aligned with the central energy band; the heatmaps of PSK signals accurately correspond to the spectral breaks caused by phase abrupt changes; the heatmaps of PAM4, QAM16, and QAM64 highlight energy intensity and bandwidth broadening characteristics, with the QAM signal heatmap showing wider coverage, reflecting the model's ability to perceive higher-order modulation constellation mappings.
[0098] Then, a horizontal comparison was made with other interpretable methods. Samples of -6dB QAM16, 12dB PAM4, and -10dB GFSK signals were selected, and heatmaps were generated using the four methods respectively. Figure 8 As shown in the results, Grad-CAM heatmaps are accompanied by background noise and have low spatial resolution; Score-CAM exhibits false activation regions and coarse localization; SS-CAM, while suppressing noise, suffers from reduced contrast and poor focusing due to smoothing all channels; while SS-CAM+ effectively suppresses background noise and significantly improves focusing by eliminating invalid channels through a gradient filtering mechanism.
[0099] The high-resolution heatmaps generated by this SS-CAM+ interpretability method not only allow observation of regions of interest to the model but also enable inference of the model's thought process when misjudging. This mechanism correlates the misjudged regions revealed by the heatmap with channel impairments such as sudden interference and multipath fading, helping to pinpoint the underlying causes of model performance degradation. These results not only provide a basis for subsequent data augmentation strategy design and network structure optimization but also serve as a key evaluation indicator for prediction reliability during actual model deployment. Figure 9As shown, taking the CPFSK signal as an example, the model misclassified it as PAM4 at -18dB. The heatmap shows that the originally continuous and smooth energy lines became discontinuous and chaotic energy fragments. The heatmap marks a bright spot slightly to the left of the image center as a feature of the model's judgment of the PAM4 signal, while the heatmap for correct identification focuses on the continuous energy lines in the lower right. This indicates that core time-frequency features are missing at low signal-to-noise ratios, and the model focuses on secondary features, leading to misclassification. In another example, the QAM64 signal was misclassified as QAM16 at -4dB. The heatmap shows that the model focuses on texture and densely populated areas. Due to noise blurring and sample length limitations, the originally denser and higher-frequency fine-grained texture features of QAM64 were not fully represented, instead being confused with the coarse texture of QAM16, which is the main reason for the misclassification. Based on the core feature region of the class activation heatmap when the model makes a correct decision, the contribution of each channel to modulation recognition is quantified. Redundant channels with low contribution in the modulation recognition network are pruned according to the target pruning rate. The performance of the pruning model is restored by fine-tuning, resulting in a pruning model that balances recognition accuracy and lightweight design. The SPWVD time-frequency image of the modulation signal is then input into the pruning model for modulation recognition.
[0100] To more objectively and comprehensively evaluate the performance of the SS-CAM+ method, quantitative evaluation metrics are introduced. SS-CAM+ is systematically compared with Grad-CAM, Score-CAM, and SS-CAM from two dimensions: interpretability and computational efficiency. Interpretation accuracy is assessed using two metrics: Average Drop (AD) and Increase in Confidence (IC). AD measures the decrease in the model's predicted score relative to the original image's predicted score when only the highlighted areas of the heatmap are retained; a lower value indicates more complete key information contained in the heatmap's covered area. IC measures the proportion of samples where the model's predicted confidence is higher than that of the original image after input masking; a higher value indicates more accurate localization and stronger noise resistance of the interpretation method.
[0101] The experiment included two control groups with sample sizes of 50 and 200, covering different modulation types and signal-to-noise ratios. Figure 10As shown, (a) and (b) are the comparison results of AD and IC indices when the sample size is 50 and 200, respectively. In terms of AD, SS-CAM+ has a significant advantage, with AD indices of 49.6% and 55.8% when the sample size is 50 and 200, respectively, indicating that retaining only its labeled region can still maintain about half of the prediction confidence. Score-CAM has the highest AD, indicating that its heatmap contains too many invalid regions, resulting in the loss of a large number of key features after masking. SS-CAM reduces AD through smoothing, but it is still higher than SS-CAM+, indicating that while simple smoothing increases the accuracy of the heatmap in locating modulated features, it does not eliminate the influence of negative features. In terms of IC, SS-CAM+ also performs best. When the sample size is 50, the IC index of SS-CAM+ is 16%, significantly higher than Grad-CAM's 10%, Score-CAM's 6%, and SS-CAM's 14%. Even with a sample size of 200, SS-CAM+ still maintains its lead, verifying its stability in capturing core features. The experimental results show that SS-CAM is significantly better than Score-CAM, demonstrating the contribution of the smoothing mechanism to noise reduction. SS-CAM+ further surpasses SS-CAM in evaluation metrics by eliminating negative contribution channels through gradient filtering, resulting in a cleaner heatmap background.
[0102] Regarding computational efficiency, the average time for generating a single heatmap using SS-CAM and SS-CAM+ was statistically analyzed. Five sets of images were randomly selected from the SPWVD dataset for testing, and the results are shown in Table 2. SS-CAM took an average of 20.88 seconds, which is due to the computational burden caused by the need for multiple forward propagation sampling of all feature channels. In contrast, SS-CAM+ took an average of only 8.58 seconds, a reduction of approximately 12.3 seconds, representing an average improvement of 58.90%. The experiments show that SS-CAM+ significantly improves computational speed without sacrificing interpretation accuracy through a gradient filtering mechanism.
[0103] Table 2
[0104]
Claims
1. A gradient-filtered smooth fractional activation map interpretation method, characterized in that: A heatmap is generated to reveal the key time-frequency features upon which the classification decision of the modulation recognition method depends; the heatmap generation method is as follows: Calculate the prediction score F of the trained modulation recognition network F() for the target modulation category c. c (I0) Feature map A of the k-th channel relative to the output of the target convolutional layer k gradient The gradient weights are then aggregated using global average pooling to obtain channel-level gradient weights. ; Configure a gradient filtering mechanism to retain only... Feature maps S with a value greater than 0 and possessing an activation response are obtained, resulting in a set of feature maps S that positively influence decision-making; each feature map S in the feature map set S is then processed sequentially. k The input mask M is obtained by upsampling and normalizing the samples respectively. k Then in mask M k Based on this, random noise sampled from a Gaussian distribution is superimposed. To generate N smooth masks with tiny perturbations. ; After the final masked image is fed back into the modulation recognition network, its confidence increment score is calculated to measure feature importance. : In the formula, Input the final mask; I b It is an all-zero matrix with the exact same size as the I0 space; For a feature map S k N smooth masks The average of the N CIC scores is used as the final weight for the k-th feature channel. Finally, the weights With the corresponding original feature map S k Perform linear weighted summation and retain the positive contribution using the ReLU function to generate an activation-like heatmap. : 。 2. The gradient-filtered smooth fractional activation mapping interpretation method as described in claim 1, characterized in that: The modulation recognition method is as follows: a smooth pseudo-Wigner-Vell distribution is used to represent the time-frequency characteristics of the modulation signal, and an SPWVD time-frequency image I0 is generated; Construct a ResNet-18 network, adjust the parameters of the input layer according to the size of the SPWVD time-frequency image, and replace the max pooling layer in the original ResNet-18 network with an identity mapping; remove the last feature extraction layer, and add a spatial attention module before the fully connected layer; An improved ResNet-18 was used as the modulation recognition network, and the cross-entropy loss function was used to measure the difference between the predicted distribution and the true label distribution for network training. The SPWVD time-frequency image I0 of the modulation signal was input to predict the modulation method.
3. The gradient-filtered smooth fractional activation mapping interpretation method as described in claim 2, characterized in that: Replace the kernel size in the input layer of the ResNet-18 network with 3*3, adjust the stride to 1, and adjust the number of padding layers to 1.
4. The gradient-filtered smooth fractional activation mapping interpretation method as described in claim 2, characterized in that: Each feature extraction layer contains two residual blocks, and each residual block includes a 3*3 convolutional layer, a batch normalization layer, and an activation layer.
5. The gradient-filtered smooth fractional activation mapping interpretation method as described in claim 2, characterized in that: The spatial attention module compresses multi-channel features into a single-channel attention weight map through 1*1 convolution, then maps the weights to the [0,1] interval through the Sigmoid activation function, and finally multiplies them element-wise with the original feature map to enhance the features of key regions.
6. The gradient-filtered smooth fractional activation mapping interpretation method as described in claim 2, characterized in that: The contribution of each channel of the modulation recognition network to the modulation recognition result is quantified by using class activation heatmaps, and the network pruning strategy is determined.
7. The gradient-filtered smooth fractional activation mapping interpretation method as described in claim 6, characterized in that: Based on the core feature regions of the activation heatmap when the model makes a correct decision, redundant channels with low contribution in the modulation recognition network are pruned according to the target pruning rate. The SPWVD time-frequency image of the modulation signal is then input into the pruning model for modulation recognition.
8. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to perform the method of any one of claims 1 to 7.