Pooling and up-sampling fusion CAM method based on information entropy and attention mechanism

By combining information entropy and attention mechanism, the attention weight matrix is designed and pooled and upsampled of feature maps is generated to generate high-quality class activation maps, which solves the overfitting and underfitting problems of CNN models in complex image processing, and improves the feature representation and interpretability of the model.

CN120356049APending Publication Date: 2025-07-22CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510432174.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

Existing CNN models are prone to overfitting or underfitting when processing complex images, making it difficult to effectively extract and utilize key information, resulting in insufficient generalization and interpretability of the model.

Method used

Combining the information entropy and attention mechanism, we evaluate the complexity and importance of the feature map by calculating the information entropy of the feature map, design the attention weight matrix, perform weighting, pooling and upsampling of the feature map, and generate a class activation map based on residual connections.

Benefits of technology

It significantly improves the model's ability to process complex images, enhances the accuracy and interpretability of feature representations, and improves the performance of computer vision tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356049A_ABST
    Figure CN120356049A_ABST
Patent Text Reader

Abstract

The invention relates to a pooling and up-sampling fusion CAM method based on information entropy and an attention mechanism. The method comprises the steps of image preprocessing, feature extraction, information entropy calculation, attention weight matrix design, weighted feature graph and pooling, up-sampling and residual connection, category weight assignment, CAM generation and the like. By introducing the information entropy theory, the complexity and importance of information in the feature map can be accurately evaluated, attention weight design is guided, key information is focused, redundancy is suppressed, and the capability of processing complex images by the model is remarkably improved. And in combination with an attention mechanism, feature representation is more effective and accurate. Through pooling and up-sampling technology fusion, the size of a feature map is flexibly adjusted, key features are reserved, and subsequent tasks are facilitated. Through the application of global average pooling and a pre-training classifier, a class activation mapping graph is generated, the performance and accuracy of a computer vision task are improved, and the explanatory performance of a model on picture processing is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision, and particularly to a pooling and upsampling fusion CAM method based on information entropy and attention mechanism. Background Art

[0002] The statements in this part only provide background technical information related to the present disclosure, and these statements may constitute prior art. During the implementation of the present invention, the inventors found that there are at least the following problems in the prior art.

[0003] In the field of computer vision, deep learning techniques, especially convolutional neural networks (CNNs), have become the core methods for tasks such as image classification, object detection, and semantic segmentation. However, with the complexity of application scenarios and the explosion of data volume, how to effectively extract and utilize the key information in images and improve the generalization ability and interpretability of models has become a hot and difficult issue in current research.

[0004] Traditional CNN models extract the feature representations of images layer by layer through structures such as stacking convolutional layers, pooling layers, and fully connected layers. However, this way of extracting features layer by layer often ignores the internal correlations and importance differences between features, resulting in the model being prone to overfitting or underfitting when dealing with complex images. To solve this problem, researchers have proposed various improvement methods, among which information entropy and attention mechanism are two techniques that have received much attention.

[0005] Information entropy is a basic concept in information theory, used to measure the uncertainty or complexity of information. In image processing, the feature map can be regarded as a probability distribution matrix, where each pixel value or the feature vector of a local region represents an event. By calculating the information entropy of the feature map, the complexity and importance of the information in the image can be evaluated, thereby guiding subsequent feature extraction and classification tasks.

[0006] On the other hand, the attention mechanism is an important feature of the human visual system, which allows us to quickly locate and focus on key information when dealing with complex scenes. In deep learning, the attention mechanism is widely used in fields such as natural language processing and computer vision. By introducing an additional weight matrix to emphasize the feature regions that contribute significantly to the task, the performance of the model can be improved.

[0007] However, how to combine and apply information entropy and attention mechanism in the model to generate high-quality class activation maps (CAMs) when dealing with complex images, so as to better improve the performance and accuracy of computer vision tasks, increase the interpretability of the model for image processing, and enhance the accuracy of image classification and localization, is a difficult problem in this field at present. Summary of the Invention

[0008] In view of the above problems, the object of the present invention is to solve a part of the problems in the prior art, or at least alleviate these problems.

[0009] A pooling and upsampling fusion CAM method based on information entropy and attention mechanism, comprising the following steps:

[0010] Perform image preprocessing on the original input image to obtain the preprocessed input image;

[0011] Use a pre-trained deep convolutional neural network (CNN) to perform forward propagation on the preprocessed input image to extract feature maps at multiple levels;

[0012] For each feature map, regard it as a probability distribution matrix, where each pixel value or the feature vector of a local region represents an event; calculate the probability distribution of each pixel value or the feature vector of a local region, that is, P(x i )), where x i represents the pixel value or the feature vector of a local region in the feature map; apply the information entropy formula to calculate the information entropy of each feature map to obtain an information entropy feature map, which is used to evaluate the complexity and importance of the information in the feature map; the information entropy formula is:

[0013]

[0014] According to the calculated information entropy feature map, design an attention weight matrix; each element of the attention weight matrix represents the weight of the corresponding feature map position; the calculation of the attention weight matrix is based on the spatial distribution of the information entropy feature map and is normalized by the softmax function (normalized exponential function) to ensure that the sum of the weights is 1;

[0015] Multiply the attention weight matrix and the feature map element by element to obtain a weighted feature map, enhance the feature regions containing important information, and suppress redundant information; perform a pooling operation on the weighted feature map to obtain a pooled feature map to reduce the size of the feature map and extract key features;

[0016] Through upsampling technology, restore the pooled feature map to the original size or the target size to obtain a restored feature map;

[0017] For the restored feature map, use global average pooling (GAP) to convert the feature map of each channel into a single value, and the single value represents the channel contribution value of the channel to the class prediction; use a pre-trained classifier to assign class weights to the restored feature map, and according to the class weights, perform weighted summation on the channel contribution values to generate a class activation mapping graph.

[0018] After restoring the pooled feature map to the original size or the target size through the upsampling technique, a residual connection is further included; the residual connection is that after the pooled feature map is restored to the original size or the target size through the upsampling technique, in combination with the residual connection strategy, the weighted feature map (i.e., the feature map before pooling) and the upsampled feature map are added element-wise or concatenated channel-wise to improve the feature representation ability.

[0019] When designing the attention weight matrix, a weighting method based on the local correlation of the feature map is adopted. For each feature map, the correlation between its adjacent pixels or regions is calculated to construct the spatial context matrix R, and then R is combined with the information entropy feature map to jointly guide the generation of the attention weight matrix; the specific formula of the attention weight matrix is as follows:

[0020] W = α·H λ (F) + β·R

[0021] where α and β are weight coefficients; H λ (F) is the information entropy feature map.

[0022] When applying the information entropy formula to calculate the information entropy of each feature map, a sparsity regularization term is introduced to perform sparsity constraints on the local regions of the feature map to suppress redundant information and improve the representation ability of the feature map; when calculating the information entropy, the sparsity regularization term is combined with the probability distribution of the feature map to jointly guide the calculation of the information entropy.

[0023] Furthermore, in the calculation process of the information entropy feature map, a regulation factor λ is also introduced to adjust the calculation sensitivity of the information entropy; the adjusted information entropy feature map is:

[0024] H λ (F i ) = H(F i ) + λ·Sparse(F i )

[0025] where Sparse(F i ) represents the sparsity regularization term of the feature map F i ;

[0026] The value range of the regulation factor λ is dynamically adjusted according to the actual application scenario. Specifically: in the image classification task, when the image background is complex and the target object is small, the value of λ is increased to highlight the information entropy feature of the target object; in the image segmentation task, when fine division of the image region is required, the value of λ is decreased to balance the information entropy features of different regions.

[0027] Furthermore, when calculating the information entropy feature map, the concept of local contrast is also introduced. By calculating the contrast between each pixel or local region in the feature map and its surrounding pixels or regions, the local information expression ability of the feature map is enhanced, so as to more accurately evaluate the complexity and importance of the information in the feature map.

[0028] Preferably, when using a pre-trained deep convolutional neural network (CNN) to extract feature maps at multiple levels, a multi-scale feature extraction strategy is adopted, that is, feature maps of different scales are extracted simultaneously to capture multi-scale information in the picture.

[0029] Preferably, when performing a pooling operation on the weighted feature map, a hybrid pooling strategy is adopted, that is, by combining the advantages of max pooling and average pooling, the best pooling method is adaptively selected according to different regions and contents of the feature map to retain more key features and reduce information loss.

[0030] When restoring the pooled feature map to the original size or target size through upsampling technology, a generative adversarial network (GAN) is used as the upsampling module. By introducing an adversarial training strategy to optimize the upsampling process, higher-quality feature maps are generated, and artifacts and blurring phenomena that may occur during the upsampling process are reduced.

[0031] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the pooling and upsampling fusion CAM method based on information entropy and attention mechanism described in any one of the above are implemented.

[0032] The present invention has the following beneficial effects:

[0033] By introducing the information entropy theory, the present invention can accurately evaluate the complexity and importance of the information in the feature map, guide the design of attention weights, focus on key information, suppress redundancy, and significantly improve the ability of the model to process complex images. Combined with the attention mechanism, the feature representation is more effective and accurate. Through the fusion of pooling and upsampling technologies, the size of the feature map is flexibly adjusted, key features are retained, and it is convenient for subsequent tasks. The use of global average pooling and pre-trained classifiers generates class activation mapping diagrams, improves the performance and accuracy of computer vision tasks, and increases the interpretability of the model for image processing.

[0034] Other features and advantages of the present invention will be described in the subsequent specification, and part of them will become obvious from the specification or be understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the structures pointed out in the specification, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1This is the flowchart of the pooling and upsampling fusion CAM method based on information entropy and attention mechanism of the present invention. Detailed implementation manners

[0036] The following further describes the present invention with reference to the accompanying drawings. The embodiments of the present invention are only used to illustrate the present invention rather than limit the present invention. Without departing from the technical idea of the present invention, various substitutions and changes made according to common general knowledge and conventional means in the art shall be included within the scope of the present invention.

[0037] As Figure 1 shown, the present invention proposes a pooling and upsampling fusion CAM method based on information entropy and attention mechanism, aiming to propose an innovative image processing technology. By combining information entropy, attention mechanism, hybrid pooling strategy, generative adversarial network upsampling, and residual connection strategy, high-quality class activation maps (CAMs) are generated. CAMs can reveal the regions in the input image that play a key role in the prediction of a specific class, thereby enhancing the accuracy of image classification and localization. The main implementation steps are as follows.

[0038] One: Image preprocessing:

[0039] First, adjust the size of the original input image to meet the input requirements of the pre-trained deep convolutional neural network (CNN);

[0040] Normalize the resized image to convert the pixel values to the range of [0, 1] or [-1, 1] to improve the stability and efficiency of model training;

[0041] According to needs, perform data augmentation operations on the image, such as rotation, cropping, flipping, color jittering, etc., to increase the data diversity and improve the generalization ability of the model;

[0042] Obtain the preprocessed input image, which is ready for subsequent feature extraction.

[0043] Two: Feature extraction

[0044] Then, use a pre-trained deep convolutional neural network (CNN), such as VGG (Visual Geometry Group) 16 or ResNet (Residual Network) 50, to perform forward propagation on the preprocessed input image to extract feature maps at multiple levels. When the pre-trained CNN extracts feature maps, it adopts a multi-scale feature extraction strategy, that is, feature maps at different scales are extracted simultaneously to capture the multi-scale information in the image.

[0045] Let the feature map output by the i-th layer of the CNN be F i , then the multi-scale feature extraction can be expressed as:

[0046] F multi-scale={F1,F2,…,F n}

[0047] where Fi i represents the feature map output by the i-th layer of the convolutional neural network (CNN), and n represents the number of levels of the feature maps extracted. These feature maps contain information at different scales and levels in the image, providing rich feature representations for subsequent processing.

[0048] III. Information Entropy Calculation

[0049] Each feature map is regarded as a probability distribution matrix, where each pixel value or the feature vector of a local region represents an event.

[0050] Calculate the probability distribution of each pixel value or the feature vector of a local region. Suppose the feature vector of a certain pixel or local region in the feature map Fi i is x, then its probability distribution can be expressed as:

[0051]

[0052] where x j represents all the feature vectors of the pixels or local regions in the feature map Fi i .

[0053] Apply the information entropy formula to calculate the information entropy of each feature map, obtaining an information entropy feature map for evaluating the complexity and importance of the information in the feature map. Apply the information entropy formula:

[0054]

[0055] When calculating the information entropy, a sparsity regularization term and a tuning factor λ can also be introduced to obtain the adjusted information entropy:

[0056] H λ (Fi i ) = H(Fi i ) + λ · Sparse(Fi i )

[0057] where Sparse(Fi i ) represents the sparsity regularization term of the feature map Fi i , and the specific calculation method can be set according to the actual situation, such as the L1 norm, etc. The value range of the tuning factor λ is dynamically adjusted according to the actual application scenario. Specifically:

[0058] In the image classification task, when the image background is complex and the target object is small, increase the value of λ to highlight the information entropy feature of the target object;

[0059] In the image segmentation task, when fine-grained division of the image area is required, the value of λ is reduced to balance the information entropy features of different regions.

[0060] When calculating the information entropy feature map, the concept of local contrast is also introduced. By calculating the Euclidean distance between pixel values or feature vectors, etc., the contrast between each pixel or local region in the feature map and its surrounding pixels or regions is calculated to enhance the local information expression ability of the feature map, so as to more accurately evaluate the complexity and importance of the information in the feature map.

[0061] IV. Attention Weight Matrix Design

[0062] According to the calculated information entropy feature map, the attention weight matrix W is designed. Each element of the weight matrix represents the weight of the corresponding position in the feature map.

[0063] The calculation of the weight matrix is based on the spatial distribution of the information entropy feature map and is normalized by the softmax function to ensure that the sum of the weights is 1. The specific formula is as follows:

[0064]

[0065] where, W ij represents the element in the i-th row and j-th column of the weight matrix W; H λ (F) represents the information entropy feature map, H λ (F) ij is the entropy value at the position (i,j) in the information entropy feature map H λ (F), λ is a tuning parameter representing the base of entropy, H λ (F) ij The larger the value, the higher the feature complexity and uncertainty at that position; H λ (F) k represents the information entropy value at the k-th position in the information entropy feature map H λ (F), and k is the index for traversing all positions.

[0066] When designing the attention weight matrix, a weighting method based on the local correlation of the feature map is adopted. Specifically, for each feature map, the correlation between its adjacent pixels or regions is calculated to construct the spatial context matrix R. Then R is combined with the information entropy feature map to jointly guide the generation of the attention weight matrix. The specific formula is as follows:

[0067] W = α·H λ (F) + β·R

[0068] Among them, α and β are weight coefficients. The spatial context matrix R can be calculated by a method based on a graph attention network (GAT). Specifically, each pixel or local region in the feature map is regarded as a node in the graph, and the GAT is used to learn the correlation between nodes to obtain the spatial context matrix C, which is used to guide the generation of the attention weight matrix;

[0069] When calculating the attention weight matrix W, the weight coefficients α and β are determined in the following way:

[0070] Set the initial values of α and β according to empirical values;

[0071] During the training process, α and β are optimized through the backpropagation algorithm so that the generated attention weight matrix W can better guide the subsequent image processing tasks.

[0072] V. Weighted Feature Map and Pooling

[0073] Multiply the attention weight matrix W element-wise with the feature map (i.e., the feature map obtained by feature extraction in step 2) to obtain the weighted feature map, enhance the feature regions containing important information, and suppress redundant information.

[0074] F weighted = W ⊙ F

[0075] Among them, ⊙ represents the element-wise multiplication operation;

[0076] Perform a pooling operation on the weighted feature map, such as max pooling or average pooling, to reduce the size of the feature map and extract key features. In this embodiment, a hybrid pooling strategy is adopted, that is, combining the advantages of max pooling and average pooling, and adaptively selecting the best pooling method according to different regions and contents of the feature map. The specific formula is as follows:

[0077] F pooled = AdaptivePool(F weighted )

[0078] Among them, AdaptivePool represents the adaptive pooling operation.

[0079] VI. Upsampling and Residual Connection

[0080] Through upsampling techniques, such as bilinear interpolation or transposed convolution, restore the pooled feature map to the original size or the target size. In this embodiment, a generative adversarial network (GAN) is used as the upsampling module, and by introducing an adversarial training strategy to optimize the upsampling process, a higher-quality feature map is generated. It is expressed as:

[0081] F upsampled = GAN(F pooled )

[0082] Among them, GAN represents a Generative Adversarial Network.

[0083] After the feature map is restored to the original size or the target size, combined with the residual connection strategy, the feature map before pooling (i.e., the weighted feature map) is element-wise added or channel-wise concatenated with the upsampled feature map to obtain the fused feature map, so as to improve the ability of feature representation. The fused feature map is expressed as:

[0084] F fused = F weighted + F upsampled or F fused = Concat(F weighted , F upsampled )

[0085] Among them, F weighted represents the feature map before pooling, Concat represents the channel-wise concatenation operation. When performing element-wise addition, the sizes of F weighted and F upsampled are the same; when performing channel-wise concatenation, their number of channels can be different, and other dimensions (such as height and width) match each other.

[0086] The restored feature map in this application can refer to either the upsampled feature map restored only to the original size or the target size, or the fused feature map after upsampling and residual connection.

[0087] VII. Class Weight Assignment and CAM Generation

[0088] For the fused feature map F fused , the global average pooling (GAP) is used to convert the feature map of each channel into a single value, representing the channel contribution value of this channel to the class prediction. It is expressed as:

[0089]

[0090] Among them, S c represents the contribution value of the c-th channel, and H and W represent the height and width of the feature map respectively.

[0091] The pre-trained classifier is used to assign class weights to the fused feature map. The classifier receives the fused feature map F fused as input, and processes the fused feature map through its internal parameters (these parameters have been optimized through the backpropagation algorithm during the training process), and finally outputs a prediction score vector. Each element in the prediction score vector corresponds to the prediction score of a class;

[0092] In this embodiment, we assume that the last layer of the classifier is a fully connected layer, and its weight matrix is Wclassifier Then, for the c-th category, its weight vector w class (c) can be obtained by extracting the c-th column from W classifier . Here, the weight vector w class (c) is extracted from the trained classifier parameters.

[0093] Generate a Class Activation Map (CAM) based on the class weights and channel contribution values. The CAM reveals the regions in the input image that play a key role in the prediction of a specific category. Specifically, for the c-th category, its CAM can be expressed as:

[0094]

[0095] where w class (c, c′) represents the weight value corresponding to the c′-th channel in the weight vector of the c-th category (here c′ is the channel index and c is the category index), and F fused (c′, :, :) represents the feature map of the c′-th channel in the fused feature map (i.e., the two-dimensional feature map after removing the channel dimension).

[0096] Through the above formula, we can generate a CAM for each category, thus visually showing which regions of the input image the model mainly focuses on when making category predictions.

[0097] By introducing the information entropy theory, the present invention can accurately evaluate the complexity and importance of information in the feature map, guide the design of attention weights, focus on key information, suppress redundancy, and significantly improve the model's ability to process complex images. Combined with the attention mechanism, the feature representation is more effective and accurate. Through the fusion of pooling and upsampling techniques, the feature map size is flexibly adjusted, key features are retained, and it is convenient for subsequent tasks. The use of global average pooling and pre-trained classifiers generates class activation maps, improves the performance and accuracy of computer vision tasks, and increases the interpretability of the model's image processing, thereby enhancing the accuracy of image classification and localization.

[0098] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the pooling and upsampling fusion CAM method based on information entropy and attention mechanism are implemented.

Claims

1. A pooling and upsampling fusion CAM method based on information entropy and attention mechanism, characterized in that, It includes the following steps: Perform image preprocessing on the original input image to obtain the preprocessed input image; Use a pre-trained deep convolutional neural network (CNN) to perform forward propagation on the preprocessed input image to extract feature maps at multiple levels; For each feature map, consider it as a probability distribution matrix, where each pixel value or the feature vector of a local region represents an event; calculate the probability distribution of each pixel value or the feature vector of a local region, that is, P(x i ), where x i represents the pixel value or the feature vector of a local region in the feature map; apply the information entropy formula to calculate the information entropy of each feature map to obtain an information entropy feature map; the information entropy formula is as follows: Design an attention weight matrix based on the calculated information entropy feature map; Each element of the attention weight matrix represents the weight of the corresponding feature map position; The calculation of the attention weight matrix is based on the spatial distribution of the information entropy feature map and is normalized by the softmax function (normalized exponential function) to ensure that the sum of the weights is 1; Multiply the attention weight matrix element-wise with the feature map to obtain the weighted feature map; perform a pooling operation on the weighted feature map to obtain the pooled feature map; Restore the pooled feature map to the original size or the target size through upsampling technology to obtain the restored feature map; For the restored feature map, use global average pooling (GAP) to convert the feature map of each channel into a single value, and the single value represents the channel contribution value of the channel to the class prediction; use a pre-trained classifier to assign class weights to the restored feature map, and based on the class weights, perform weighted summation on the channel contribution values to generate a class activation mapping.

2. The CAM method for pooling and upsampling fusion based on information entropy and attention mechanism according to claim 1, characterized in that, After restoring the pooled feature map to the original size or the target size through upsampling technology, it also includes a residual connection; the residual connection is that after the pooled feature map is restored to the original size or the target size through upsampling technology, in combination with the residual connection strategy, the weighted feature map and the upsampled feature map are added element-wise or concatenated channel-wise.

3. The pooling and upsampling fusion CAM method based on information entropy and attention mechanism according to claim 1, wherein When designing the attention weight matrix, a weighting method based on the local correlation of the feature map is adopted. For each feature map, calculate the correlation between its adjacent pixels or regions to construct a spatial context matrix R, and then combine R with the information entropy feature map to jointly guide the generation of the attention weight matrix; the specific formula of the attention weight matrix is as follows: W = α·H λ (F)+β·R where α and β are weight coefficients; H λ (F) is the information entropy feature map.

4. The pooling and upsampling fusion CAM method based on information entropy and attention mechanism according to claim 1, wherein When applying the information entropy formula to calculate the information entropy of each feature map, introduce a sparsity regularization term to perform sparsity constraint on the local area of the feature map to suppress redundant information and improve the representation ability of the feature map; when calculating the information entropy, combine the sparsity regularization term with the probability distribution of the feature map to jointly guide the calculation of the information entropy.

5. The pooling and upsampling fusion CAM method based on information entropy and attention mechanism according to claim 4, wherein During the calculation process of the information entropy feature map, a regulation factor λ is also introduced to adjust the calculation sensitivity of the information entropy; the adjusted information entropy feature map is: H λ (F i ) = H(F i ) + λ·Sparse(F i ) Among them, Sparse(F i ) represents the sparsity regularization term of the feature map F i ; The value range of the regulation factor λ is dynamically adjusted according to the actual application scenario. Specifically: in an image classification task, when the image background is complex and the target object is small, increase the value of λ to highlight the information entropy feature of the target object; in an image segmentation task, when fine segmentation of the image area is required, decrease the value of λ to balance the information entropy features of different regions.

6. The pooling and upsampling fusion CAM method based on information entropy and attention mechanism according to claim 5, characterized in that When calculating the information entropy feature map, the concept of local contrast is also introduced to enhance the local information expression ability of the feature map by calculating the contrast between each pixel or local region in the feature map and its surrounding pixels or regions.

7. The pooling and upsampling fusion CAM method based on information entropy and attention mechanism according to claim 1, wherein When using a pre-trained deep convolutional neural network (CNN) to extract feature maps at multiple levels, a multi-scale feature extraction strategy is adopted, that is, feature maps at different scales are extracted simultaneously to capture multi-scale information in the picture.

8. The pooling and upsampling fusion CAM method based on information entropy and attention mechanism according to claim 1, characterized in that When performing pooling operations on the weighted feature maps, a hybrid pooling strategy is adopted, that is, by combining the advantages of max pooling and average pooling, the best pooling method is adaptively selected according to different regions and contents of the feature maps to retain more key features and reduce information loss.

9. The pooling and upsampling fusion CAM method based on information entropy and attention mechanism according to claim 1, characterized in that, When restoring the pooled feature maps to the original size or the target size through the upsampling technique, a generative adversarial network (GAN) is used as the upsampling module. By introducing an adversarial training strategy to optimize the upsampling process, higher-quality feature maps are generated, and artifacts and blurring phenomena that may occur during the upsampling process are reduced.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the pooling and upsampling fusion CAM method based on information entropy and attention mechanism according to any one of claims 1 to 9.

Citation Information

Cited By

  • Sample difficulty assessment and adaptive data enhancement method based on thermogram

    CN122416167A

  • Thermographic-based sample difficulty assessment and adaptive data augmentation method

    CN122416167B