Prostate cancer image classification method based on attention guidance
By introducing a multi-directional residual band pooling attention module and a multi-scale feature-guided self-attention module in the DenseNet121 model, the problem of insufficient feature capture in the prostate cancer image classification is solved, and higher diagnostic accuracy and reliability are achieved.
Patent Information
- Application Number
- CN202510539893.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-08-08
AI Technical Summary
Existing deep learning models fail to fully utilize multi-scale feature combination, residual connection and attention mechanisms in prostate cancer image classification, resulting in insufficient capture and integration of key features in complex images, affecting diagnostic accuracy and reliability.
The multi-directional residual band pooling attention module and multi-scale feature-guided self-attention module are introduced in the DenseNet121 model. The attention to important features is enhanced through residual connections, and global self-attention is captured through multi-scale features to optimize feature representation.
It significantly improves the accuracy and reliability of prostate cancer image classification, enhances the model's ability to pay attention to key features, and improves the diagnostic performance.
Smart Images

Figure CN120451096A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of deep learning image classification algorithms and medical image classification, and in particular relates to a prostate cancer image classification method based on attention guidance. Background Art
[0002] Prostate cancer is a common malignancy in men, and its early diagnosis and treatment are crucial to the patient's prognosis. Medical imaging technologies such as magnetic resonance imaging (MRI) and ultrasound imaging play a key role in prostate cancer diagnosis. In recent years, the development of deep learning technology has brought new insights into cancer diagnosis. Deep learning algorithms can automatically identify and extract features of relevant lesions in medical images and output predictions. Compared to previous methods that relied on doctors' subjective judgment of images, this technology can provide important reference for doctors' judgments, greatly improving the accuracy and efficiency of their diagnoses.
[0003] Convolutional neural networks (CNNs), with their powerful ability to extract image features, have demonstrated outstanding performance in medical image classification tasks. DenseNet121, as an efficient deep learning model, is widely used in image classification tasks due to its dense connectivity. Other common classification networks include ResNet, AlexNet, and MobileNet. Some researchers have achieved promising results in medical image classification tasks by building and improving existing classification networks, further demonstrating the positive impact of deep learning algorithms on prostate cancer diagnosis. However, compared to other medical tasks, research on prostate cancer images still has considerable room for development.
[0004] In medical image classification tasks, attention mechanisms effectively assist convolutional neural networks (CNNs) in feature selection and information focusing by assigning different weights to input features. This helps highlight important features, such as lesion areas, and suppress irrelevant noise, thereby guiding the model to more accurately diagnose and analyze. Common attention mechanisms include spatial attention, channel attention, and self-attention, each of which focuses on important features of different types and dimensions within an image. By appropriately employing these attention mechanisms for prostate cancer images, the model's classification performance can be significantly improved, enhancing its understanding of key image features and providing strong technical support for the early detection and diagnosis of prostate cancer.
[0005] Multi-scale feature integration and residual connections are key strategies for improving model performance. By simultaneously capturing information at different scales within an image, the model can more comprehensively understand both image details and global structure, which is particularly important for identifying complex lesion features. Residual connections, by introducing shortcuts within the network, alleviate the vanishing gradient problem in deep networks, enhancing feature transfer and model training efficiency. Combining these two strategies, the model not only more effectively extracts and integrates multi-level feature information, but also maintains stable performance during deep learning, significantly improving the accuracy and reliability of diagnoses for diseases like prostate cancer.
[0006] While existing methods have made progress in medical image classification, they still fail to fully exploit the potential of techniques such as multi-scale feature integration, residual connections, and attention mechanisms. This deficiency results in models being unable to fully capture and integrate key features when processing complex prostate cancer images, which in turn affects the accuracy and reliability of diagnosis. Summary of the Invention
[0007] Purpose of the invention: In response to the related problems pointed out in the background technology, the present invention proposes a prostate cancer image classification method based on attention guidance. By introducing an improved multi-directional residual strip pooling attention module (MDRSPA) into the backbone network DenseNet121, it focuses on the long-distance contextual information in different directions of the feature map, and enhances the attention to important feature points through residual connections to highlight key features. At the same time, a multi-scale feature guided self-attention module (MFGSA) is introduced to guide the calculation of global self-attention by capturing local features of different scales, optimize global features, enhance the feature representation of the model, and thus improve the classification performance of the original backbone model.
[0008] Technical solution: The present invention proposes a prostate cancer image classification method based on attention guidance, comprising the following steps:
[0009] Step 1: Divide the prostate cancer image dataset into training and test sets in proportion, and each classification category is also divided in proportion;
[0010] Step 2: Improve the backbone network DenseNet121 model. Add a multi-faceted residual strip pooling attention module after the first two DenseBlocks of the DenseNet121 model. Add a multi-faceted residual strip pooling attention module branch and a multi-scale feature-guided self-attention module branch after the third and fourth Dense Blocks of the DenseNet121 model. Then, concatenate the output feature maps of the multi-faceted residual strip pooling attention module branch and the multi-scale feature-guided self-attention module branch after the third and fourth Dense Blocks.
[0011] Step 3: Train and optimize the improved backbone network DenseNet121 model, and use the optimized backbone network DenseNet121 model to classify and detect prostate cancer images.
[0012] Furthermore, the specific method of step 1 is:
[0013] Step 1.1: Collect relevant prostate cancer classification datasets;
[0014] Step 1.2: Split the dataset into training and test sets in a 4:1 ratio and create subfolders by class label, where the class labels are 0 for prostate hyperplasia and 1 for prostate cancer; or 0 for NC without cancer, 1 for GG3 with Gleason grade 3, 2 for GG4 with Gleason grade 4, and 3 for GG5 with Gleason grade 5.
[0015] Step 1.3: No pre-processing is done on the dataset. The images are only resized to 256×256 during model training and testing.
[0016] Furthermore, the multi-directional residual strip pooling attention module improves the original strip pooling module, uses residual connections to supplement information on feature maps in both horizontal and vertical directions, adds additional attention information, highlights key feature points with larger eigenvalues in the feature maps, and guides the module to focus on important feature areas.
[0017] Furthermore, the multi-directional residual strip pooling attention module is specifically as follows:
[0018] For the expanded horizontal and vertical feature maps in the strip pooling module, they are combined with the original input feature maps through residual connections, respectively. The formula is as follows:
[0019]
[0020] Where S h and S v Represent the feature results after residual connection in the horizontal and vertical directions respectively, and Represent the expanded feature maps in two directions respectively, and X represents the original input feature map;
[0021] The feature map S h and S v Combine, that is, S=S h +S v ,Then through the calculation of attention, the module is guided to pay more attention to the key features.
[0022] Furthermore, the multi-scale feature guided self-attention module combines multi-scale local feature information and uses the self-attention mechanism to capture global dependencies, thereby enabling local features to guide global features, as follows:
[0023] 1) During the self-attention calculation process, multi-scale local feature guided MSLFG operations are performed on the key K and value V respectively. Feature extraction is performed through different convolution operation branches. Then, these multi-scale features are fused to enhance the local feature representation. There are two selection strategies for input feature maps of different sizes. The specific methods are as follows:
[0024] Strategy 1: When the feature map size is small, it includes a 3x3 depth convolution branch and a dilation convolution branch with a dilation rate of 2:
[0025]
[0026] Strategy 2: When the feature map size is large, add an additional dilation rate of 4 to the dilation convolution branch:
[0027]
[0028] Where K dw and V dw , K dilated2 and V dilated2 and K extra4 and V extra4 Denote the output results of depthwise convolution and dilated convolution respectively. DepthwiseConv3×3(·), DilatedConv3×3(·) and ExtraDilatedConv3×3(·) represent different convolution operations respectively. Dilated2 and extra4 indicate that the dilation rates of dilated convolution are 2 and 4 respectively.
[0029] 2) Features of different scales are then concatenated. For different strategies, the process is expressed as follows:
[0030]
[0031] In the formula, Concat represents the concatenation operation, because V combined As a further operation on the value V, it contains important feature information of the input feature map, so K is connected through the residual combined Supplementary information:
[0032] K final =K combined +V combined
[0033] 3) Finally, the self-attention is calculated. The multi-head attention mechanism divides the input features into multiple heads. For each head, K final Calculate attention weights independently and weight features V combined Get the output features of each head Output i , the calculation formula is as follows:
[0034]
[0035] Among them, i represents a head, d k is the dimension of the key vector, used to scale the dot product;
[0036] 4) Connect the output features of all heads to get the final output features.
[0037] Furthermore, the feature map output by the third Dense Block of the DenseNet121 model uses strategy 2 for multi-scale feature extraction, and the feature map output by the fourth Dense Block uses strategy 1 for multi-scale feature extraction.
[0038] The present invention adopts the above technical solution and has the following beneficial effects:
[0039] While retaining the original feature extraction capabilities of the DenseNet121 backbone network, by combining innovative attention feature enhancement strategies (multi-directional residual strip pooling attention module, multi-scale feature guided self-attention module), it can more effectively guide the module to focus on areas that are critical to medical image classification, and enhance the model feature representation by combining multi-scale feature information.
[0040] The present invention improves the original strip pooling module, uses residual connections to supplement the feature maps in the horizontal and vertical directions, adds additional attention information, thereby highlighting the key feature points with larger eigenvalues in the feature map, and guiding the module to focus on important feature areas.
[0041] The present invention designs a multi-scale feature-guided self-attention module, which combines multi-scale local feature information through multi-scale local feature guidance operations and uses the self-attention mechanism to capture their global dependencies, thereby realizing the guidance of local features on global features, so that the module can better grasp the global feature representation.
[0042] The present invention comprehensively considers the performance and computational complexity of the model, and adds the proposed multi-directional residual strip pooling attention module and multi-scale feature-guided self-attention module to the backbone network DenseNet121 with a reasonable strategy to improve the classification performance of the DenseNet121 model. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 It is the overall flow chart of the present invention;
[0044] Figure 2 The overall model architecture proposed;
[0045] Figure 3 The proposed Multi-Directional Residual Strip Pooling Attention Module (MDRSPA);
[0046] Figure 4 The proposed multi-scale feature guided self-attention module (MFGSA);
[0047] Figure 5 Visualization of the attention of the proposed model on the input image. DETAILED DESCRIPTION
[0048] The present invention is further illustrated below with reference to specific examples. It should be understood that these examples are only used to illustrate the present invention and are not used to limit the scope of the present invention. After reading the present invention, modifications of various equivalent forms of the present invention made by those skilled in the art all fall within the scope defined by the claims attached to this application.
[0049] The present invention discloses a prostate cancer image classification method based on attention guidance, which specifically includes the following steps:
[0050] Step 1: Divide the prostate cancer image dataset into training and test sets in proportion, and each classification category is also divided in proportion.
[0051] Step 1.1: Collect relevant prostate cancer classification datasets.
[0052] Two datasets are used in this implementation: (1) the HY Prostate private dataset from Huai'an First People's Hospital, which contains 888 prostate CT scan images (431 for benign prostatic hyperplasia and 457 for prostate cancer). (2) the SICAPv2 public dataset from the Internet, which contains 12,081 prostate tissue pathology images, including 4,417 NC (no cancer), 2,222 GG3 (Gleason grade 3), 4,494 GG4 (Gleason grade 4), and 948 GG5 (Gleason grade 5).
[0053] Step 1.2: Divide the dataset into training and test sets in a ratio of 4:1, and create subfolders by category label. (1) HY Prostate: The training set contains 709 images, of which 344 are prostate hyperplasia (label: 0) and 365 are prostate cancer (label: 1). The test set contains 179 images, of which 87 are prostate hyperplasia and 92 are prostate cancer. (2) SICAPv2: The training set contains 9959 images, of which 3773 are NC (label: 0), 1829 are GG3 (label: 1), 3641 are GG4 (label: 2), and 716 are GG5 (label: 3). The test set contains 2122 images, of which 644 are NC, 393 are GG3, 853 are GG4, and 232 are GG5.
[0054] Step 1.3: No preprocessing is performed on the two datasets. The images are only resized to 256×256 during model training and testing to facilitate observation of the model performance.
[0055] Step 2: Improve the original strip pooling module, use residual connections to supplement the feature maps in the horizontal and vertical directions, add additional attention information, thereby highlighting the key feature points with larger eigenvalues in the feature map, and guide the module to focus on important feature areas. Figure 3 , the specific contents of the module are as follows:
[0056] For a given input feature map X∈R C×H×W , where C, H, and W represent the number of channels, height, and width of the input feature map, respectively. The input feature map is averagely pooled along the height and width dimensions into a strip feature through horizontal strip pooling and vertical strip pooling. It can be expressed as:
[0057]
[0058] Similarly, vertical strip pooling features It can be expressed as:
[0059]
[0060] The elements X(:,:,j) and X(:,i,:) represent the eigenvalues of a specific feature point, respectively. i and j are the horizontal and vertical positions of the feature point in the feature map.
[0061] Then, the strip features are expanded to the original image size through the 1D convolution kernel to obtain the expanded feature maps in two directions and Such manipulation can effectively grasp the important feature points of the feature map in a certain row or column.
[0062] For the expanded horizontal and vertical feature maps in the strip pooling module, they are combined with the original input feature maps through residual connections. Since the original input feature maps themselves contain rich feature information, this combination can effectively supplement the information of the expanded feature maps, making the feature representation richer and more comprehensive. The formula is as follows:
[0063]
[0064] Where S h and S v Represent the feature results after residual connection in the horizontal and vertical directions respectively, and They represent the expanded feature maps in two directions respectively, and X represents the original input feature map.
[0065] The feature map S supplemented by residual connection h and S v , which can effectively highlight the important feature points and regions in the horizontal and vertical directions, rather than just expanding the feature representation of the entire row or column in the feature map. h and S v Combine, that is, S=S h +S v , the important features in both directions can be comprehensively considered to further enhance the prominence of these features.
[0066] Finally, the attention weight is calculated and weighted on the input feature map to guide the module to pay more attention to these key features. The output result Z can be expressed as:
[0067] Z=Scale(X,Sig(f(S)))
[0068] Where Scale(,) represents the element-by-element multiplication operation, Sig(·) represents the Sigmoid function, and f(·) represents the 1×1 convolution operation.
[0069] Step 3: Design a multi-scale feature guided self-attention module, which combines multi-scale local feature information and uses the self-attention mechanism to capture its global dependency, so as to guide the global features from the local features and better grasp the global feature representation. Figure 4 , the module is implemented as follows:
[0070] For the input feature map X∈R C×H×W, using two-dimensional position encoding to add a fixed position encoding PE to each position of the feature map, ensuring that the model can use position information for self-attention calculation. The feature map X' with position encoding can be expressed as:
[0071] X′=X+PE
[0072] Then, a 1×1 convolution is performed on X' to perform a linear transformation to generate the query Q, key K and value V matrices:
[0073] Q=X'W Q
[0074] K=X'W K
[0075] V=X'W V
[0076] Where W Q , W K , W V Represents the weight matrix that needs to be learned.
[0077] During the self-attention calculation process, the multi-scale local feature guidance (MSLFG) operation is performed on the key K and value V. Specifically, the feature extraction is performed through different convolution branches, and then these multi-scale features are fused to enhance the local feature representation. There are two selection strategies for input feature maps of different sizes, as follows:
[0078] Strategy 1. When the feature map size is small, it includes a 3x3 depthwise convolution branch and a dilation convolution branch with a dilation rate of 2:
[0079]
[0080] Strategy 2. When the feature map size is large, add an additional dilation rate of 4 to the dilation branch:
[0081]
[0082] Where K dw and V dw , K dilated2 and V dilated2 and K extra4 and V extra4 They represent the output results of depthwise convolution and dilated convolution respectively. DepthwiseConv3×3(·), DilatedConv3×3(·) and ExtraDilatedConv3×3(·) represent different convolution operations respectively. dilatrd2 and extra4 indicate that the dilation rates of dilated convolution are 2 and 4 respectively.
[0083] Subsequently, features of different scales are concatenated. For different strategies, the process can be expressed as:
[0084]
[0085] In the formula, Concat represents the concatenation operation. combined As a further operation on the value V, it contains important feature information of the input feature map, so K is connected through the residual combined Supplementary information:
[0086] K final =K combined +V combined
[0087] Finally, the self-attention is calculated, and K and V containing multi-scale local feature information can play a role. The multi-head attention mechanism divides the input features into multiple heads, and for each head, K containing multi-scale local feature information is used. final Calculate attention weights independently and weight features V combined Get the output features of each head Output i , the calculation formula is as follows:
[0088]
[0089] Where i represents the number of heads, d k is the dimension of the key vector, used to scale the dot product.
[0090] Then, the output features of all heads are concatenated to obtain the final output features:
[0091] MultiHead(Q,K,V)=Concat(Output1,…,Output i )W o
[0092] Where i represents the number of heads, W o is the output linear transformation matrix.
[0093] By combining multi-scale features, the feature representation is enriched. Therefore, the multi-head self-attention mechanism can more effectively utilize these features, calculate global dependencies, and achieve precise guidance of global features.
[0094] Step 4: Add the proposed multi-directional residual strip pooling attention module and multi-scale feature guided self-attention module to the backbone network DenseNet121 to improve the classification performance of DenseNet121. Figure 2 , as follows:
[0095] Step 4.1: Add a multi-directional residual strip pooling attention module after each of the four dense blocks in the DenseNet121 model to enable it to highlight the key feature points and areas in each stage.
[0096] Step 4.2: To reduce the computational complexity of the self-attention mechanism, a multi-scale feature-guided self-attention module branch is added only after the third and fourth Dense Blocks in the DenseNet121 model, which have smaller output feature maps, to capture global feature representations. Given the size of the input feature map, the third Dense Block outputs a larger feature map, so Strategy 2 (described in Step 3) is used for multi-scale feature extraction. Strategy 1 (described in Step 3) is used for the smaller feature map output by the fourth Dense Block, achieving a reasonable balance between computational complexity and performance requirements.
[0097] Step 4.3: Concatenate the output feature maps of the multi-directional residual strip pooling attention module branch after the third and fourth Dense Blocks and the multi-scale feature guided self-attention module branch, combining the highlighted important feature information and the global features guided by the multi-scale features, so as to further refine the features, enrich and enhance the feature representation.
[0098] Step 4.4: For the four stages with the proposed module added, perform attention visualization on the final feature maps of each stage to observe how the added module focuses on and affects the features, reflecting the role of the proposed module.
[0099] Step 5: Verify the feasibility and effectiveness of the model through comparative experiments and ablation experiments, and perform attention visualization to further demonstrate the correctness of the model's classification decisions.
[0100] Comparative experiment:
[0101] The proposed algorithm model was compared with several state-of-the-art models on two prostate cancer datasets. The comparison models mainly included six classic basic models, three improved Transformer-based algorithm models, and one improved convolutional neural network-based algorithm model. The experimental results on the HY Prostate dataset and the SICAPv2 dataset are shown in Tables 1 and 2, respectively.
[0102] Table 1 Experimental results of HYProstate dataset
[0103]
[0104] Table 2 Experimental results of SICAPv2 dataset
[0105]
[0106]
[0107] The algorithm model proposed in this invention has achieved reasonable control of feature extraction, multi-scale feature fusion and highlighting of important features, and achieved excellent results on two data sets, confirming its feasibility.
[0108] Reference Figure 5 , calculate the attention weight of the final output feature map of the model, and then use the attention weight to generate an attention heat map. Finally, the attention heat map is superimposed with the original image to highlight the area that the model focuses on. Figure 5 The first row is the original image; the second row is the attention visualization map; (a) prostate cancer samples in HY Prostate dataset; (b) benign prostatic hyperplasia samples in HY Prostate dataset; (c) G3 samples in SICAPv2 dataset; (d) G4 samples in SICAPv2 dataset; (e) G5 samples in SICAPv2 dataset.
[0109] Ablation experiments: To further demonstrate that the proposed modules can effectively improve the classification performance of the backbone network, the performance of each module was evaluated. The experimental results on the two datasets are shown in Tables 3 and 4, respectively.
[0110] Table 3 Evaluation of the design module on the HYProstate dataset
[0111]
[0112] Table 4. Module evaluation of the designed modules on the SICAPv2 dataset
[0113]
[0114] The multi-directional residual strip pooling attention module MDRSPA and the multi-scale feature guided self-attention module MFGSA proposed in this paper achieved higher accuracy on both datasets, and the overall accuracy of the combination of the two reached the optimal result, fully demonstrating the effectiveness of the proposed modules.
[0115] In order to further demonstrate the applicability and feasibility of the module proposed in this invention, it is considered to be applied to DenseNet as well as another backbone model ResNet, as shown in Tables 5 and 6.
[0116] Table 5 Improved performance of different backbone models on HYProstate dataset
[0117]
[0118] Table 6 Improved performance of different backbone models on SICAPv2 dataset
[0119]
[0120] The two improved backbone models achieved significant improvements in accuracy on both datasets. This demonstrates that the two module algorithms proposed in this invention can not only significantly improve the classification accuracy of the backbone models, but also have good applicability.
[0121] The present invention can be combined with a medical image computer diagnosis system to assist doctors in the classification and diagnosis of medical images in clinical practice and provide valuable reference.
[0122] The above embodiments are intended only to illustrate the technical concepts and features of the present invention. Their purpose is to enable those skilled in the art to understand and implement the present invention accordingly. They are not intended to limit the scope of protection of the present invention. Any equivalent modifications or variations based on the spirit of the present invention are intended to fall within the scope of protection of the present invention.
Claims
1. A prostate cancer image classification method based on attention guidance, characterized in that: The steps include: Step 1: Divide the prostate cancer image dataset into training and test sets in proportion, and each classification category is also divided in proportion; Step 2: Improve the backbone network DenseNet121 model. Add a multi-faceted residual strip pooling attention module after the first two Dense Blocks of the DenseNet121 model. Add a multi-faceted residual strip pooling attention module branch and a multi-scale feature-guided self-attention module branch after the third and fourth Dense Blocks of the DenseNet121 model. Then, concatenate the output feature maps of the multi-faceted residual strip pooling attention module branch and the multi-scale feature-guided self-attention module branch after the third and fourth Dense Blocks. Step 3: Train and optimize the improved backbone network DenseNet121 model, and use the optimized backbone network DenseNet121 model to classify and detect prostate cancer images.
2. The attention-guided dense network for prostate cancer image classification according to claim 1, characterized in that The specific method of step 1 is: Step 1.1: Collect relevant prostate cancer classification datasets; Step 1.2: Split the dataset into training and test sets in a 4:1 ratio and create subfolders by class label, where the class labels are 0 for prostate hyperplasia and 1 for prostate cancer; or 0 for NC without cancer, 1 for GG3 with Gleason grade 3, 2 for GG4 with Gleason grade 4, and 3 for GG5 with Gleason grade 5. Step 1.3: No pre-processing is done on the dataset. The images are only resized to 256×256 during model training and testing.
3. The method for prostate cancer image classification based on attention guidance according to claim 1, characterized in that: The multi-directional residual strip pooling attention module improves the original strip pooling module, uses residual connections to supplement information in feature maps in both horizontal and vertical directions, adds additional attention information, highlights key feature points with larger eigenvalues in the feature maps, and guides the module to focus on important feature areas.
4. The method for prostate cancer image classification based on attention guidance according to claim 3, characterized in that: The multi-directional residual strip pooling attention module is specifically as follows: For the expanded horizontal and vertical feature maps in the strip pooling module, they are combined with the original input feature maps through residual connections, respectively. The formula is as follows: Where S h and S v Represent the feature results after residual connection in the horizontal and vertical directions respectively, and Represent the expanded feature maps in two directions respectively, and X represents the original input feature map; The feature map S h and S v Combine, that is, S=S h +S v ; Then, through attention calculation, the module is guided to pay more attention to key features.
5. The method for prostate cancer image classification based on attention guidance according to claim 1, characterized in that: The multi-scale feature guided self-attention module combines multi-scale local feature information and uses the self-attention mechanism to capture global dependencies, thereby enabling local features to guide global features. Specifically: 1) During the self-attention calculation process, multi-scale local feature guided MSLFG operations are performed on the key K and value V respectively. Feature extraction is performed through different convolution operation branches. Then, these multi-scale features are fused to enhance the local feature representation. There are two selection strategies for input feature maps of different sizes. The specific methods are as follows: Strategy 1: When the feature map size is small, it includes a 3x3 depth convolution branch and a dilation convolution branch with a dilation rate of 2: Strategy 2: When the feature map size is large, add an additional dilation rate of 4 to the dilation convolution branch: Where K dw and V dw , K dilated2 and V dilated2 and K extra4 and V extra4 Denote the output results of depthwise convolution and dilated convolution respectively. DepthwiseConv3×3(·), DilatedConv3×3(·) and ExtraDilatedConv3×3(·) represent different convolution operations respectively. Dilated2 and extra4 indicate that the dilation rates of dilated convolution are 2 and 4 respectively. 2) Features of different scales are then concatenated. For different strategies, the process is expressed as follows: In the formula, Concat represents the concatenation operation, because V combined As a further operation on the value V, it contains important feature information of the input feature map, so K is connected through the residual combined Supplementary information: K final =K combined +V combined 3) Finally, the self-attention is calculated. The multi-head attention mechanism divides the input features into multiple heads. For each head, K final Calculate attention weights independently and weight features V combined Get the output features of each head Output i , the calculation formula is as follows: Where i represents the number of heads, d k is the dimension of the key vector, used to scale the dot product; 4) Connect the output features of all heads to get the final output features.
6. The method for prostate cancer image classification based on attention guidance according to claim 5, characterized in that: The feature map output by the third Dense Block of the DenseNet121 model uses strategy 2 for multi-scale feature extraction, and the feature map output by the fourth Dense Block uses strategy 1 for multi-scale feature extraction.