Tumor pathological image typing and grading method and device based on deep learning and residual network technology

CN122049494BActive Publication Date: 2026-09-15JIANG SU AI YING YI LIAO KE JI YOU XIAN GONG SI +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202610141924.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-02-02
Publication Date
2026-09-15
Estimated Expiration
2046-02-02

AI Technical Summary

Technical Problem

[0002]随着数字病理技术的快速发展,基于深度学习的肿瘤病理图像分析已成为辅助医学诊断的重要研究方向,通过对肿瘤组织病理切片进行自动化分型与分级,可为临床治疗决策提供客观、可重复的依据,有效减轻病理医生的工作负担并提高诊断一致性;然而,现有方法多集中于单一任务,如分类或分割,难以在同一框架下实现分型与分级这两个密切相关但又具有不同语义层次任务的协同处理,导致模型泛化能力有限、可解释性不足

Benefits of technology

[0040] This invention constructs a deep convolutional feature extraction network based on residual networks to extract features from both shallow and deep layers. After upsampling and weighted fusion, it not only preserves the detailed information of the tissue region but also enhances the global semantic features. By using parallel convolutional branches to extract specific features and combining feature deviation to generate a spatial attention map, it achieves feature enhancement for hierarchical tasks on classification tasks, promotes knowledge complementarity between tasks, and makes the classification results more consistent with the essential characteristics of tumor pathology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122049494B_ABST
    Figure CN122049494B_ABST
Patent Text Reader

Abstract

The application provides a tumor pathological image typing and grading method and device based on deep learning and residual network technology, relates to the technical field of image processing, and comprises the following steps: collecting and pre-processing tumor pathological images, segmenting tissue regions, performing color normalization, and dividing into image blocks; a deep convolution feature extraction network based on a residual network is constructed to extract shallow and deep feature maps; the deep feature maps are up-sampled and weightedly fused with the shallow feature maps, input into a parallel convolution branch to obtain typing and grading feature maps; a spatial attention map is generated by calculating feature deviation based on the grading feature maps to determine the region focusing degree; the attention weight is used to enhance the typing feature map; the enhanced feature map and the grading feature map are respectively subjected to global pooling and classification to output typing and grading results; and a threshold is set according to the region focusing degree to screen high attention points and generate a key region heat map; and the application realizes collaborative processing of typing and grading, and improves the efficiency and accuracy of diagnosis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, specifically to a method and apparatus for classifying and grading tumor pathological images based on deep learning and residual network technology. Background Technology

[0002] With the rapid development of digital pathology technology, deep learning-based tumor pathology image analysis has become an important research direction for auxiliary medical diagnosis. By automatically classifying and grading tumor tissue pathology slides, objective and reproducible evidence can be provided for clinical treatment decisions, effectively reducing the workload of pathologists and improving diagnostic consistency. However, existing methods are mostly focused on single tasks, such as classification or segmentation, and it is difficult to achieve the collaborative processing of two closely related but semantically different tasks of classification and grading within the same framework, resulting in limited model generalization ability and insufficient interpretability.

[0003] In the prior art, a fine-grained brain tumor classification method based on residual channel attention, disclosed in CN117173521A, improves the fine-grained classification performance of brain tumor MRI images by introducing a residual channel attention module and a bilinear attention pooling algorithm. However, this method mainly focuses on a single classification task and does not involve tumor grading tasks, and lacks a deep fusion mechanism for multi-scale features. In addition, its attention generation method does not fully integrate the semantic interaction between tasks, which limits the model's ability to locate key regions and its interpretability, making it difficult to adapt to the actual diagnostic needs of parallel typing and grading tasks in pathological images.

[0004] The information disclosed in the background section is only intended to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0005] The purpose of this invention is to provide a method and system for classifying and grading tumor pathological images based on deep learning and residual network technology, so as to solve the problems mentioned in the background art.

[0006] To achieve the above objectives, the present invention provides the following technical solution:

[0007] A tumor pathology image typing and grading method and system based on deep learning and residual network technology, the specific steps of which include:

[0008] Step 1: Acquire digital pathological images of tumor tissue and perform preprocessing. Segment the tissue region based on pixel grayscale and color features. Normalize the color of the segmented tissue region and divide it into multiple image blocks of the same size.

[0009] Step 2: Construct a deep convolutional feature extraction network based on residual network. Input the image patch into the network for forward propagation and extract the first intermediate feature map and the second intermediate feature map from the shallow and deep layers of the network, respectively. The spatial size of the second intermediate feature map is smaller than that of the first intermediate feature map.

[0010] Step 3: Upsample the second intermediate feature map, calculate the fusion weights of corresponding pixels in the upsampled feature map and the first intermediate feature map, and perform weighted fusion of the two based on the fusion weights to obtain the fused feature map. Input the fused feature map into two parallel convolution branches for processing, and output the first task feature map and the second task feature map.

[0011] Step 4: Calculate the feature deviation of each pixel in the second task feature map, generate a spatial attention map containing the spatial attention weights of each pixel in the second task feature map based on the feature deviation, and determine the region focus based on the statistical distribution of the attention weights; multiply the attention weights of each pixel in the spatial attention map element-wise with the feature vectors of the corresponding pixels in the first task feature map to obtain the enhanced first task feature map.

[0012] Step 5: Perform global pooling and classification processing on the enhanced first task feature map and the second task feature map respectively to obtain the tumor subtyping and grading results; determine the critical threshold based on the regional focus degree, select high attention points from the spatial attention map according to the critical threshold, and generate a tumor key region heat map based on the distribution of high attention points.

[0013] Furthermore, the preprocessing specifically refers to image noise reduction preprocessing to eliminate light and shadow noise;

[0014] The tissue region is segmented based on pixel grayscale and color features. Specifically, the grayscale value and RGB color channel value of each pixel in the image are first obtained. The pixel region that conforms to the tissue features is identified according to the preset threshold range. After removing noise and holes through morphological operations, a continuous tissue region mask is obtained.

[0015] The logic for dividing the image blocks is as follows: the original image area covered by the tissue region mask is aligned with the standard pathological template image in terms of color distribution. The brightness and chromaticity channels are adjusted in the Lab color space through a color transfer algorithm to eliminate color differences caused by different staining batches. Finally, the normalized tissue region image is divided into grids according to a preset fixed size to generate multiple image blocks of the same size.

[0016] Furthermore, a deep convolutional feature extraction network is constructed based on the residual network. The deep convolutional feature extraction network uses a pre-trained ResNet as its backbone and removes its terminal global pooling layer and classification layer.

[0017] The segmented image blocks are input into the deep convolutional feature extraction network for forward propagation. At a shallow layer of the network, specifically at the output of the second residual module, a first intermediate feature map is extracted, the size of which is denoted as [size not specified]. At a deeper level in the network, specifically at the output of the fourth residual module, the second intermediate feature map is extracted, and its size is denoted as... And satisfy , , ;in, , , These represent the height, width, and number of channels of the first intermediate feature map, respectively. , , These represent the height, width, and number of channels of the second intermediate feature map, respectively.

[0018] Furthermore, the second intermediate feature map is upsampled using bilinear interpolation to make it the same size as the first intermediate feature map, thereby obtaining an upsampled feature map;

[0019] For any pixel in the upsampled feature map Its pixel value It is calculated as follows: The pixel is determined in the second intermediate feature map. The four closest pixels, based on pixel location Calculate the weight coefficients of the four pixels based on their relative positions to the four nearest pixels in both the horizontal and vertical directions; then sum the pixel values ​​of the four pixels after multiplying them by their corresponding weight coefficients to obtain the upsampled feature map pixels. Pixel values;

[0020] Concatenate the upsampled feature map with the first intermediate feature map in channel-dimensional order, and input a... The convolutional layer generates an initial fusion weight map and performs Softmax normalization on the initial fusion weight map along the channel dimension to obtain a fusion weight map. Each spatial location corresponds to a normalized fusion weight vector, which contains two weight values, corresponding to the fusion weights of the upsampled feature map and the first intermediate feature map, respectively.

[0021] The upsampled feature map and the first intermediate feature map are weighted and fused based on the fusion weight map. Specifically, for each spatial location, the fusion weight vector corresponding to that location is obtained from the fusion weight map. The feature vector of the upsampled feature map at that location is multiplied by its fusion weight, and the feature vector of the first intermediate feature map at that location is multiplied by its fusion weight. The two are then added together, and the result is the feature vector of the fusion feature map at that location. The feature vector refers to the vector composed of the feature values ​​of the pixel in all channels.

[0022] Perform the above operations on all spatial locations to obtain a complete fused feature map;

[0023] The fused feature map is simultaneously input into two parallel convolutional branches. The first convolutional branch consists of two 3×3 convolutional layers and one 1×1 convolutional layer, used to extract feature representations for tumor subtyping tasks and output a first task feature map. The second convolutional branch consists of two 3×3 convolutional layers and one 1×1 convolutional layer, with its convolutional kernel parameters independent of the first convolutional branch, used to extract feature representations for tumor grading tasks and output a second task feature map. The spatial dimensions of the first and second task feature maps are the same as those of the fused feature map, and the number of channels is determined by the number of output channels of the 1×1 convolutional layers in each branch.

[0024] Furthermore, for the second task feature map, the feature deviation of each pixel is calculated along the channel dimension. Specifically, for each pixel, the mean of all channel feature values ​​is first calculated; then, the difference between the feature value of the pixel in each channel and the mean is squared and summed. The difference between the total number of channels and one is calculated. The sum is divided by the difference to obtain the quotient. The square root of the quotient is then taken. The result is the feature deviation of the pixel.

[0025] The feature deviation of each pixel is normalized to between 0 and 1 using the Sigmoid activation function, and used as the initial attention weight for that pixel to generate an initial spatial attention map. The initial spatial attention map is then subjected to Gaussian smoothing filtering to suppress noise, thus obtaining the spatial attention map.

[0026] The mean and variance of the attention weights of all pixels in the spatial attention map are statistically analyzed. The region focus is determined based on the mean and variance. Specifically, the mean and variance are multiplied by their respective preset weights, and the product of the two is added together to obtain the region focus.

[0027] Furthermore, the attention weight of each pixel in the spatial attention map is multiplied element-wise with the feature vector of the pixel at the same spatial position in the first task feature map. Specifically, for each pixel in the spatial attention map, its attention weight is multiplied with the feature value of each channel in the feature vector of the corresponding pixel in the first task feature map to obtain the enhanced feature vector.

[0028] The feature vector refers to the vector composed of the feature values ​​of a pixel across all channels;

[0029] By iterating through all pixels in the spatial attention map, an enhanced first task feature map is obtained, whose spatial size is consistent with that of the first task feature map, and whose number of channels remains unchanged.

[0030] Furthermore, the enhanced first task feature map is subjected to global average pooling, compressing its spatial dimension into a one-dimensional vector, which is then input into the first classifier for classification to obtain the tumor subtyping result; the second task feature map is subjected to global average pooling, compressing its spatial dimension into a one-dimensional vector, which is then input into the second classifier for classification to obtain the tumor grading result; wherein, the first classifier and the second classifier are respectively composed of a fully connected layer and a Softmax activation function, and the parameters of the two classifiers are independent of each other.

[0031] Furthermore, the region focus is multiplied by a preset baseline threshold to obtain a critical threshold; for each pixel in the spatial attention map, its attention weight is compared with the critical threshold, and if its attention weight is greater than or equal to the critical threshold, the pixel is marked as a high attention point.

[0032] All high-attention points are mapped back to their corresponding spatial locations in the fused feature map, and a color gradient heatmap, i.e., a heatmap of key tumor regions, is generated based on their attention weights.

[0033] The present invention also provides a tumor pathology image typing and grading device based on deep learning and residual network technology, for performing the above-described tumor pathology image typing and grading method based on deep learning and residual network technology, comprising:

[0034] The preprocessing module acquires digital pathological images of tumor tissue and performs preprocessing. It segments the tissue region based on pixel grayscale and color features, normalizes the color of the segmented tissue region, and divides it into multiple image blocks of the same size.

[0035] The feature extraction module is used to construct a deep convolutional feature extraction network based on residual networks. The image patch is input into the network for forward propagation, and the first intermediate feature map and the second intermediate feature map are extracted from the shallow and deep layers of the network, respectively. The spatial size of the second intermediate feature map is smaller than that of the first intermediate feature map.

[0036] The fusion module is used to upsample the second intermediate feature map, calculate the fusion weight of the corresponding pixels of the upsampled feature map and the first intermediate feature map, and perform weighted fusion of the two based on the fusion weight to obtain the fused feature map. The fused feature map is then input into two parallel convolutional branches for processing, and the first task feature map and the second task feature map are output.

[0037] The enhancement module is used to calculate the feature deviation of each pixel in the second task feature map, generate a spatial attention map containing the spatial attention weights of each pixel in the second task feature map based on the feature deviation, and determine the region focus based on the statistical distribution of the attention weights; multiply the attention weights of each pixel in the spatial attention map element-wise with the feature vectors of the corresponding pixels in the first task feature map to obtain the enhanced first task feature map.

[0038] The classification and grading module is used to perform global pooling and classification processing on the enhanced first task feature map and the second task feature map respectively to obtain tumor subtyping and grading results; a critical threshold is determined based on the regional focus degree, high attention points are selected from the spatial attention map according to the critical threshold, and a heat map of key tumor regions is generated based on the distribution of high attention points.

[0039] Compared with the prior art, the beneficial effects of the present invention are:

[0040] This invention constructs a deep convolutional feature extraction network based on residual networks to extract features from both shallow and deep layers. After upsampling and weighted fusion, it not only preserves the detailed information of the tissue region but also enhances the global semantic features. By using parallel convolutional branches to extract specific features and combining feature deviation to generate a spatial attention map, it achieves feature enhancement for hierarchical tasks on classification tasks, promotes knowledge complementarity between tasks, and makes the classification results more consistent with the essential characteristics of tumor pathology.

[0041] Furthermore, the introduction of regional focus in this invention allows for adaptive adjustment of the high-interest screening threshold, and the generated heatmap can accurately identify key tumor regions, providing doctors with intuitive references. This invention balances automation and interpretability, avoiding the limitations of single-task models, and eliminating interference caused by differences between imaging and staining through preprocessing such as color normalization and noise reduction, significantly improving diagnostic consistency, effectively reducing the workload of pathologists, and providing objective and reliable auxiliary evidence for clinical treatment decisions. Attached Figure Description

[0042] Figure 1 This is a schematic diagram of the overall method flow of the present invention;

[0043] Figure 2 A scatter plot showing the sum of squared deviations versus the degree of characteristic deviation when the total number of channels is 32, 64, and 128;

[0044] Figure 3 A 3D scatter plot of the total number of channels, sum of squared deviations, and feature deviation.

[0045] Figure 4 This is a schematic diagram of the overall device module of the present invention. Detailed Implementation

[0046] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments.

[0047] It should be noted that, unless otherwise defined, the technical or scientific terms used in this invention should have the ordinary meaning understood by one of ordinary skill in the art to which this invention pertains. The terms "first," "second," and similar terms used in this invention do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0048] Example:

[0049] Please see Figures 1-3 The present invention provides a technical solution:

[0050] A tumor pathology image typing and grading method based on deep learning and residual network technology includes the following steps:

[0051] Step 1: Acquire digital pathological images of tumor tissue and perform preprocessing. Segment the tissue region based on pixel grayscale and color features. Normalize the color of the segmented tissue region and divide it into multiple image blocks of the same size.

[0052] In this embodiment, the preprocessing specifically refers to image noise reduction preprocessing to eliminate light and shadow noise. Specifically, it involves: First, using Gaussian filtering to smooth the entire image of the acquired digital pathological image of tumor tissue, suppressing Gaussian noise and salt-and-pepper noise introduced by the image sensor; Next, addressing the light and shadow interference caused by uneven illumination or staining in the pathological image, using a brightness correction method based on local adaptive thresholds to equalize the image and eliminate excessive differences between bright and dark areas; Then, converting the image from RGB space to Lab space through color space conversion, and performing histogram equalization or adaptive contrast enhancement on the brightness channel to improve the contrast between the tissue area and the background and further suppress artifacts caused by non-uniform illumination; Finally, combining morphological opening and closing operations to perform post-processing on the image, removing tiny isolated noise points and filling small holes, thereby effectively eliminating light and shadow noise in the image while maintaining sharp tissue edges and texture details.

[0053] After preprocessing, tissue regions are segmented based on pixel grayscale and color features. Specifically, the grayscale value and RGB color channel value of each pixel in the image are first obtained, and the pixel regions that conform to tissue features are identified according to the preset threshold range. After removing noise and holes through morphological operations, a continuous tissue region mask is obtained.

[0054] The threshold range is set based on the color distribution characteristics of tissue regions in typical pathological staining. Specifically, the pixel grayscale values ​​and RGB channel values ​​of tissue regions in a batch of standard stained sample images are statistically analyzed, and an initial threshold is set based on expert experience. Then, the K-means clustering algorithm is used to perform preliminary segmentation of the images, and the threshold range is dynamically adjusted to adapt to different staining batches. In this embodiment, the threshold range is set as follows: grayscale values ​​are between 50 and 200, and the red channel value in RGB is higher than 100 and the blue channel value is lower than 150, so as to effectively distinguish tissue regions from background and non-tissue regions.

[0055] The logic for dividing the image into blocks is as follows: The original image region covered by the tissue region mask is aligned with the standard pathological template image in terms of color distribution. Specifically, both are first converted from the RGB color space to the Lab color space, and the mean and standard deviation of their luminance and chrominance channels are calculated respectively. Then, through linear transformation, the mean and standard deviation of the L, a, and b channels of the original image region are adjusted to match those of the standard template image, achieving color distribution alignment and effectively eliminating color differences caused by different staining batches or scanning devices. After color normalization, the normalized tissue region image is divided into grids: starting from the upper left corner of the image, the tissue region is cut using a non-overlapping sliding window with a fixed size of 224×224 pixels, following a left-to-right and top-to-bottom order. If the edges of the tissue region cannot be completely covered, overlapping sliding windows or zero-padding are used in the edge areas to ensure that each image block is the same size, ultimately generating a series of image blocks, each 224×224 pixels in size, for input into the subsequent deep convolutional feature extraction network for processing.

[0056] The 224×224 pixel size is the standard input size for pre-trained residual networks. It can directly utilize the weights trained on large-scale image datasets without adjusting the network structure or performing complex size resampling. At the same time, 224×224 pixels can preserve key details such as cell nuclei and stained areas in tumor pathology images, while also being sufficient to cover local tissue structures. This is beneficial for the model to capture both fine-grained features and local semantics, achieving a good balance between computational efficiency, information integrity, and model compatibility.

[0057] Step 2: Construct a deep convolutional feature extraction network based on residual network. Input the image patch into the network for forward propagation and extract the first intermediate feature map and the second intermediate feature map from the shallow and deep layers of the network, respectively. The spatial size of the second intermediate feature map is smaller than that of the first intermediate feature map.

[0058] In this embodiment, a deep convolutional feature extraction network is constructed based on a residual network. The deep convolutional feature extraction network uses a pre-trained ResNet as its backbone and removes its terminal global pooling layer and classification layer.

[0059] The segmented image blocks are input into the deep convolutional feature extraction network for forward propagation. At a shallow layer of the network, specifically at the output of the second residual module, a first intermediate feature map is extracted, the size of which is denoted as [size not specified]. At a deeper level in the network, specifically at the output of the fourth residual module, the second intermediate feature map is extracted, and its size is denoted as... And satisfy , , ;in, , , These represent the height, width, and number of channels of the first intermediate feature map, respectively. , , These represent the height, width, and number of channels of the second intermediate feature map, respectively.

[0060] The deep convolutional feature extraction network receives image patches of size 224×224×3 as input, where 3 represents the three RGB color channels. The network uses a pre-trained ResNet-50 as its backbone, and its basic structure includes an initial convolutional layer, a batch normalization layer, and a ReLU activation function, followed by four consecutive residual modules. Each module consists of several stacked residual units. Each residual unit contains a convolutional layer, a batch normalization layer, and a ReLU activation function, and skip connections are used to directly add the input to the output to ensure effective gradient propagation. After removing the global average pooling layer and fully connected classification layer at the end of the original ResNet, the output of the second residual module is used as the shallow feature extraction point, with an output dimension of [missing information]. The output of the fourth residual module is used as the deep feature extraction point, and its output dimension is... ,in , , and , , These are the height, width, and number of channels of the corresponding feature map, respectively, and satisfy the following conditions: , , All convolutional layers in the network use ReLU as the activation function, batch normalization layers are used to accelerate training and improve generalization ability, and the weights of the entire network are initialized with ImageNet pre-trained parameters to achieve effective transfer learning on tumor pathology images.

[0061] The purpose of constructing a deep convolutional feature extraction network based on residual networks is to efficiently and accurately extract multi-level and multi-scale semantic features from tumor pathology images. Residual Networks (ResNet) are an architecture that solves the gradient vanishing and degradation problems in deep neural network training by introducing skip connections. Its core idea is to directly pass input features to the output of deeper layers, enabling the network to learn residual mappings rather than direct mappings, thereby supporting the stable training of deeper networks. In tumor pathology image analysis, visual features at different levels have different semantics: shallow features usually contain detailed information such as edges and textures, which are suitable for capturing cell morphology and local structures; deep features carry more abstract semantic information, such as tissue arrangement, lesion areas and other global patterns. By constructing a deep feature extraction network based on residual networks, the model can not only use its strong representational ability of deep structure to capture high-level pathological semantics, but also use the gradient flow of residual connections to maintain training stability. Thus, it can simultaneously acquire detailed and global information in a single forward propagation, providing a robust and distinguishable feature foundation for subsequent classification and grading tasks, and improving the model's understanding and diagnostic accuracy of complex pathological structures.

[0062] Step 3: Upsample the second intermediate feature map, calculate the fusion weights of corresponding pixels in the upsampled feature map and the first intermediate feature map, and perform weighted fusion of the two based on the fusion weights to obtain the fused feature map. Input the fused feature map into two parallel convolution branches for processing, and output the first task feature map and the second task feature map.

[0063] In this embodiment, the second intermediate feature map is upsampled using bilinear interpolation to make it the same size as the first intermediate feature map, thereby obtaining the upsampled feature map.

[0064] For any pixel in the upsampled feature map Its pixel value It is calculated as follows: The pixel is determined in the second intermediate feature map. The four closest pixels, based on pixel location Calculate the weight coefficients of the four pixels based on their relative positions to the four nearest pixels in both the horizontal and vertical directions; then sum the pixel values ​​of the four pixels after multiplying them by their corresponding weight coefficients to obtain the upsampled feature map pixels. Pixel values;

[0065] This method employs bilinear interpolation upsampling, and the calculation of its weight coefficients is derived from the natural extension of linear interpolation in a two-dimensional plane: linear interpolation is performed in both the horizontal and vertical directions, and the final pixel value of the interpolation point is obtained by weighting the four nearest neighbor pixels according to their relative distances. The formula for calculating the weight coefficients directly reflects the spatial relative positional relationship between the interpolation point and its neighboring pixels: the closer the pixel is, the higher its weight, and vice versa, thus achieving a smooth and continuous spatial transition.

[0066] This method can maintain the spatial continuity and smoothness of feature distribution while enlarging the spatial size of feature maps, effectively avoiding block artifacts or information loss caused by direct enlargement. This weighted method based on spatial distance is not only simple and efficient to calculate, but also preserves the geometric structure of local features, providing a foundation for the accurate fusion of subsequent multi-level features and helping to improve the model's ability to express detailed textures and semantic information.

[0067] Concatenate the upsampled feature map with the first intermediate feature map in channel-dimensional order, and input a... The convolutional layer generates an initial fusion weight map, and then performs Softmax normalization on the initial fusion weight map along the channel dimension to obtain a fusion weight map. Each spatial location corresponds to a normalized fusion weight vector, which contains two weight values, corresponding to the fusion weights of the upsampled feature map and the first intermediate feature map, respectively.

[0068] The 1×1 convolutional layer automatically learns and fuses the cross-scale feature interaction information between the upsampled feature map and the first intermediate feature map along the channel dimension, generating an initial fusion weight map. Specifically, the 1×1 convolution dynamically captures the complementarity and importance differences between features of different scales by performing linear combination and nonlinear mapping between channels on the concatenated multi-scale features, thus generating an initial weight allocation for each spatial location. Subsequently, Softmax normalization is performed along the channel dimension to ensure that the sum of the weights corresponding to the two scales at each spatial location is 1, which reflects the competition between weights and ensures the numerical stability of the fusion process. This method adaptively learns the fusion ratio of multi-scale features through a data-driven approach, avoiding the limitations of manually setting fixed weights. This allows the model to dynamically adjust the contribution of shallow detail features and deep semantic features according to the specific image content, thereby enhancing the global semantic representation while preserving local structural information and improving the effectiveness and task adaptability of feature fusion.

[0069] The upsampled feature map and the first intermediate feature map are weighted and fused based on the fusion weight map. Specifically, for each spatial location, the fusion weight vector corresponding to that location is obtained from the fusion weight map. The feature vector of the upsampled feature map at that location is multiplied by its fusion weight, and the feature vector of the first intermediate feature map at that location is multiplied by its fusion weight. The two are then added together, and the result is the feature vector of the fusion feature map at that location. The feature vector refers to the vector composed of the feature values ​​of the pixel in all channels.

[0070] Perform the above operations on all spatial locations to obtain a complete fused feature map;

[0071] This method achieves the fusion of upsampled feature maps and first intermediate feature maps through a spatially adaptive weight allocation mechanism. For each spatial location in the image, the model can dynamically adjust the fusion ratio of the two based on the complementarity and importance differences of multi-scale features at that location: in regions with rich texture details and complex edge structures, shallow features can be given higher weights to preserve fine structures; in regions with prominent semantic information and significant global patterns, the contribution of deep features is increased to enhance high-level semantic representation.

[0072] The fused feature map is simultaneously input into two parallel convolutional branches. The first convolutional branch consists of two 3×3 convolutional layers and one 1×1 convolutional layer, used to extract feature representations for tumor subtyping tasks and output a first task feature map. The second convolutional branch consists of two 3×3 convolutional layers and one 1×1 convolutional layer, with its convolutional kernel parameters independent of the first convolutional branch, used to extract feature representations for tumor grading tasks and output a second task feature map. The spatial dimensions of the first and second task feature maps are the same as those of the fused feature map, and the number of channels is determined by the number of output channels of the 1×1 convolutional layers in each branch.

[0073] The first convolutional branch extracts features for tumor subtyping, while the second convolutional branch extracts features for tumor grading. The key lies in the fact that both branches employ the same structure but independent parameters, allowing each branch to adaptively learn and optimize features for the specific semantic requirements of the task. Specifically, the first convolutional branch extracts fine-grained structural features such as local texture, cell morphology, and tissue arrangement through two 3×3 convolutional layers. These features are crucial for distinguishing tumor subtypes, such as adenocarcinoma and squamous cell carcinoma. Then, a 1×1 convolutional layer reorganizes and compresses these features along the channel dimension, forming a high-dimensional feature representation specifically for subtyping tasks—the first task feature map. Simultaneously, the second convolutional branch, with its independent convolutional kernel parameters, captures features related to tumor malignancy, such as nuclear atypia, mitotic density, and stromal invasion patterns, through the same two 3×3 convolutional layers. These features focus on quantifying the biological behavior and invasive potential of the lesion. Finally, an independent 1×1 convolutional layer maps these features to a dedicated feature representation for grading tasks—the second task feature map.

[0074] Through this parallel and parameter-decoupled branch design, the model can learn and optimize feature representations that adapt to different semantic levels of classification and gradation, based on sharing the same fused feature map, thereby achieving collaborative processing of the two tasks and targeted feature extraction under a unified framework.

[0075] Step 3 effectively combines deep semantic features with shallow detail features through upsampling and weighted fusion, which not only preserves the local texture and edge information of the image, but also incorporates the global tissue structure and lesion semantics. Then, task-specific features are extracted through parallel convolution branches, realizing the decoupling and optimization of classification and hierarchical features, providing accurate and complementary feature representations for subsequent steps, and improving the model's ability to distinguish and recognize complex pathological patterns.

[0076] Step 4: Calculate the feature deviation of each pixel in the second task feature map, generate a spatial attention map containing the spatial attention weights of each pixel in the second task feature map based on the feature deviation, and determine the region focus based on the statistical distribution of the attention weights; multiply the attention weights of each pixel in the spatial attention map element-wise with the feature vectors of the corresponding pixels in the first task feature map to obtain the enhanced first task feature map.

[0077] In this embodiment, for the second task feature map, the feature deviation of each pixel is calculated along the channel dimension, based on the following formula:

[0078]

[0079] In the formula, Indicates the first Feature deviation of each pixel The index of the pixel in the feature map of the second task; This represents the total number of channels in the feature map of the second task. The index of the channels in the feature map of the second task; Indicates the first The pixel at the th point The characteristic values ​​of each channel; Indicates the first The mean of the feature values ​​of each pixel across all channels;

[0080] For this formula, the dependent variable Used to characterize the second task feature map The degree of consistency of the feature value distribution of a pixel across all channels, that is, the average degree of deviation of all channel feature values ​​of that pixel from its mean. The larger the value, the greater the feature difference of the pixel in each channel, and the more drastic the information change; The smaller the value, the more consistent the feature values ​​of each channel are, and the more stable and concentrated the local features of the pixel are.

[0081] The more channels there are, the smoother the average contribution per channel becomes, making the distribution of eigenvalues ​​more concentrated, thereby reducing... value; The sum of squares represents the deviation of a feature value of a certain channel from the mean. The larger the difference, the more uneven the expression of the feature among the channels. In real physical scenarios, channels usually correspond to different feature extraction filters or semantic attributes. If the response values ​​of a certain pixel differ significantly in different channels, it indicates that the pixel contains more semantic changes, which may correspond to the edge, texture or pathological structure transformation area in the image.

[0082] The formula borrows the statistical concept of sample standard deviation. By calculating the sum of squares of the relative means of the feature values ​​of each channel, dividing by the number of channels minus one, and then taking the square root, it effectively quantifies the dispersion of feature distribution within a pixel. This calculation method can adapt to network structures with different numbers of channels, and the normalization factor... This avoids artificially amplifying fluctuations due to an increase in the number of channels, thus ensuring the dependent variable... It possesses comparability and stability, laying the foundation for subsequent applications based on... The generation of attention weights provides an objective and interpretable statistical basis, enhancing the model's ability to capture key regions.

[0083] Table 1: Feature Deviation Statistics

[0084]

[0085] It should be noted that the pixel index, total number of channels, feature value of each channel, feature mean, and sum of squared deviations in Table 1 correspond to... , , , , item.

[0086] Based on the 15 sets of feature deviation data in Table 1 and Figure 2 -(1) to Figure 2 The scatter plot analysis of -(3) shows that the total number of channels in the second task feature map has a significant impact on the feature deviation: when the total number of channels is 32, the feature deviation is concentrated between 0.07 and 0.098, with large overall fluctuations; when the total number of channels is increased to 64, the feature deviation drops to the range of 0.04 to 0.075, with a more concentrated distribution; when the total number of channels is further increased to 128, the feature deviation stabilizes between 0.033 and 0.056, with the smallest dispersion; this indicates that as the total number of channels increases, the consistency of the feature values ​​of each channel of the pixel is enhanced, and the calculation results of the feature deviation are more reliable;

[0087] further Figure 3 It can be seen that when the number of channels is fixed, the increase of the sum of squared deviations will directly lead to an increase in feature deviation, which reflects the degree of feature fluctuation of pixels in different channels. This statistical characteristic provides a basis for the subsequent generation of spatial attention weights, enabling the model to identify pixel regions with significant feature changes based on feature deviation, thereby enhancing the representation ability of key regions in the classification task and improving the synergy and interpretability between classification and hierarchical tasks.

[0088] The feature deviation of each pixel is normalized to between 0 and 1 using the Sigmoid activation function, and used as the initial attention weight for that pixel to generate an initial spatial attention map. The initial spatial attention map is then subjected to Gaussian smoothing filtering to suppress noise, thus obtaining the spatial attention map.

[0089] The mean and variance of the attention weights of all pixels in the spatial attention map are used to determine the region focus using the following formula:

[0090]

[0091] In the formula, For regional focus; and These are the mean and variance of the attention weights for all pixels in the spatial attention map, respectively. and The preset weights for the corresponding indicators, and .

[0092] In tumor pathology analysis, identifying key regions requires not only a high overall level of attention but also a model capable of clearly distinguishing between different regions; that is, the attention should exhibit significant spatial variability. It can more sensitively capture local anomalies or significant changes, therefore it should be given greater weight when calculating region focus. It relies more on the uneven distribution of attention, thereby enhancing the model's ability to capture key structures, such as tumor margins and areas of cellular atypia.

[0093] For this formula, the dependent variable Used to characterize the concentration of overall attention distribution and the significance of spatial variations in spatial attention maps; The larger the value, the more concentrated the attention is in the image, the stronger the salient change in the local region, and the clearer the model's spatial focus; The smaller the value, the more evenly the attention is distributed or the changes are gradual, indicating a lack of a clear focus area.

[0094] The mean value reflects the overall attention level of the entire attention map; a higher mean value indicates that the model pays more attention to the image as a whole. This is used to measure the spatial dispersion of attention distribution. A larger variance indicates that attention is unevenly distributed in the image, meaning that some areas are highlighted. Physically... and The overall data reflects the consistency of the model's response to organizational features in space: if the mean is high and the variance is large, it indicates that the model focuses on a concentrated area and has obvious regional bias; if the mean is low and the variance is small, it indicates that the model pays less attention to the overall image and the areas of focus are scattered.

[0095] The formula combines the concentration and variability of attention distribution, comprehensively reflecting the spatial attention characteristics of the model through a weighted summation method; this linear weighted form is simple and intuitive, easy to calculate and adjust, and... It can reflect both the overall intensity of attention and the salience of spatial structure; in addition, and The settings can be fine-tuned according to the specific task characteristics to ensure that the regional focus is adapted to the visual salience of the actual pathological structure, and to provide a reliable statistical basis for dynamically generating heatmap thresholds.

[0096] The attention weight of each pixel in the spatial attention map is multiplied element-wise with the feature vector of the pixel at the same spatial location in the first task feature map.

[0097]

[0098] In the formula, The first in the spatial attention graph Enhanced feature vector of each pixel The index of the pixel in the spatial attention map; The first in the spatial attention graph Attention weights for each pixel; This indicates that the first task feature map corresponds to the first... Feature vector of each pixel;

[0099] This represents the feature vector enhanced by spatial attention weighting. Its significance lies in selectively enhancing or suppressing the feature representation of different spatial locations in the feature map of the first task by introducing feature deviation information from the hierarchical task. This improves the sensitivity and discriminative ability of the classification task to key pathological regions and enhances the interpretability and effectiveness of the model's feature representation in multi-task collaborative processing.

[0100] In real-world physical scenarios, attention weights This reflects the dispersion and significance of the feature distribution at this location in the feature map of the hierarchical task. That is, regions with high feature deviation often correspond to key areas with obvious pathological structural changes and rich information, such as tumor edges and atypical cells. Multiplying this weight with the feature vector of the classification task is equivalent to introducing a spatial guidance mechanism at the feature level. This allows the classification task to focus more on the pathological regions with high information content indicated by the hierarchical task while maintaining its own semantic features. This enables feature complementarity and knowledge transfer between tasks, thereby improving the model's overall understanding of complex pathological images.

[0101] The feature vector refers to the vector composed of the feature values ​​of a pixel across all channels;

[0102] By iterating through all pixels in the spatial attention map, an enhanced first task feature map is obtained, whose spatial size is consistent with that of the first task feature map, and whose number of channels remains unchanged.

[0103] By calculating the feature deviation of each pixel in the feature map of the second task and generating a spatial attention map based on this, the discreteness of feature distribution in the tumor grading task is transformed into spatial guidance information for the subtyping task. This method introduces the weights of key regions identified in the grading task, such as tumor edges or atypical cell regions with significant feature changes, into the feature map of the subtyping task through a spatial attention mechanism, realizing feature complementarity and knowledge transfer between tasks. This not only enhances the subtyping task's ability to focus on important pathological regions and improves the model's discrimination accuracy and robustness, but also further improves the model's interpretability through the visualization of attention weights, providing doctors with clear guidance on pathological areas of concern, thereby strengthening the clinical practical value of decision support while ensuring diagnostic accuracy.

[0104] Step 5: Perform global pooling and classification processing on the enhanced first task feature map and the second task feature map respectively to obtain the tumor subtyping and grading results; determine the critical threshold based on the regional focus degree, select high attention points from the spatial attention map according to the critical threshold, and generate a tumor key region heat map based on the distribution of high attention points.

[0105] In this embodiment, a global average pooling operation is first performed on the enhanced first task feature map. Specifically, each channel in the enhanced first task feature map is traversed, and the arithmetic mean of all spatial location feature values ​​of that channel is calculated. This compresses the original three-dimensional feature map, which contains height, width, and the number of channels dedicated to the classification task, into a one-dimensional feature vector with a length equal to the number of its channels. This vector integrates all spatial information related to the classification task. Subsequently, this one-dimensional vector is input into the first classifier, which consists of at least one fully connected layer. The classifier extracts high-level semantic representations through linear transformation and nonlinear activation, and uses the Softmax activation function in the output layer to map the features to a probability distribution belonging to each tumor classification category. Finally, the category with the highest probability is selected as the tumor classification result output.

[0106] Simultaneously, the same global average pooling process is performed independently on the feature map of the second task, compressing it into a one-dimensional vector with a length equal to its own number of channels, and inputting it into a second classifier with the same structure but independent parameters for processing. This classifier also uses a fully connected layer and a softmax function to convert the features into a probability distribution of tumor grading categories, and outputs the final grading result.

[0107] The first and second classifiers are composed of fully connected layers and Softmax activation functions, respectively, and their parameters are independent. The training process is as follows: Based on the pre-trained residual network, a phased training strategy is adopted. First, the parameters of the backbone network are frozen. Pathological image blocks labeled with tumor typing and grading are used as input. The fused feature maps are extracted for task-specific features through independent parallel convolutional branches. Then, the independent parameters of the first and second classifiers are input. The typing loss and grading loss are calculated simultaneously using the cross-entropy loss function. Backpropagation is performed using the multi-task weighted sum as the total loss to optimize the parameters of the two classifiers and parallel branches. After the task-specific modules have initially converged, some layers of the backbone network are unfrozen, and end-to-end joint fine-tuning is performed with a low learning rate to optimize the feature extraction, fusion, enhancement, and classification modules in a coordinated manner, ultimately obtaining a model that can simultaneously output high-precision typing and grading results.

[0108] The critical threshold is obtained by multiplying the region focus by a preset baseline threshold. For each pixel in the spatial attention map, its attention weight is compared with the critical threshold. If its attention weight is greater than or equal to the critical threshold, the pixel is marked as a high attention point.

[0109] Pixels with attention weights greater than or equal to the critical threshold are marked as high-attention points because these pixels exhibit significant anomalies in feature deviation, indicating that the region shows high information variation or structural abnormalities in the grading task, and may correspond to key pathological structures such as tumor margins or atypical cells. By setting a threshold for screening, the system can automatically focus on regions that are of great significance for classification and grading tasks, improving the interpretability of the model and providing doctors with intuitive diagnostic basis.

[0110] The baseline threshold was determined by analyzing the attention weight distribution of the spatial attention map in historical training data and selecting the weight value corresponding to the 85th percentile as the baseline threshold. This baseline threshold can effectively distinguish pixels that are significantly higher than the general level in the attention distribution, covering approximately 15% of the highly significant regions. These regions typically correspond to key pathological structures such as tumor margins and cellular atypia. Furthermore, this threshold is automatically determined based on the distribution of real data, avoiding the arbitrariness of subjective experience and ensuring a balance between sensitivity and specificity in the selection of high-attention points.

[0111] All high-interest points are mapped back to their corresponding spatial locations in the fused feature map, and a color gradient heatmap is generated based on their attention weights. Specifically: First, the two-dimensional coordinates of each selected high-interest point are scaled back to their original spatial location in the fused feature map according to the ratio of their spatial dimensions to the fused feature map. Next, for each mapped high-interest point, a corresponding color value is found in a continuous color gradient from blue to red based on its corresponding attention weight value in the spatial attention map. Points with higher weights are assigned a color closer to red, and points with lower weights are assigned a color closer to blue. Then, each high-interest point is colored using the determined color value at its spatial location in the original fused feature map. Locations not marked as high-interest points are left transparent or filled with the background color. Finally, all colored points are superimposed to form a pseudo-color heatmap where the color gradient is positively correlated with the attention weight, i.e., a tumor key region heatmap, which intuitively identifies the key pathological regions determined by the model that significantly contribute to tumor classification and grading tasks.

[0112] Please see Figure 4 A tumor pathology image typing and grading device based on deep learning and residual network technology, comprising:

[0113] The preprocessing module acquires digital pathological images of tumor tissue and performs preprocessing. It segments the tissue region based on pixel grayscale and color features, normalizes the color of the segmented tissue region, and divides it into multiple image blocks of the same size.

[0114] The feature extraction module is used to construct a deep convolutional feature extraction network based on residual networks. The image patch is input into the network for forward propagation, and the first intermediate feature map and the second intermediate feature map are extracted from the shallow and deep layers of the network, respectively. The spatial size of the second intermediate feature map is smaller than that of the first intermediate feature map.

[0115] The fusion module is used to upsample the second intermediate feature map, calculate the fusion weight of the corresponding pixels of the upsampled feature map and the first intermediate feature map, and perform weighted fusion of the two based on the fusion weight to obtain the fused feature map. The fused feature map is then input into two parallel convolutional branches for processing, and the first task feature map and the second task feature map are output.

[0116] The enhancement module is used to calculate the feature deviation of each pixel in the second task feature map, generate a spatial attention map containing the spatial attention weights of each pixel in the second task feature map based on the feature deviation, and determine the region focus based on the statistical distribution of the attention weights; multiply the attention weights of each pixel in the spatial attention map element-wise with the feature vectors of the corresponding pixels in the first task feature map to obtain the enhanced first task feature map.

[0117] The classification and grading module is used to perform global pooling and classification processing on the enhanced first task feature map and the second task feature map respectively to obtain tumor subtyping and grading results; a critical threshold is determined based on the regional focus degree, high attention points are selected from the spatial attention map according to the critical threshold, and a heat map of key tumor regions is generated based on the distribution of high attention points.

[0118] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.

[0119] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented in software, the above embodiments can be implemented, in whole or in part, as a computer program product. Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution.

[0120] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.

[0121] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.

Claims

1. A tumor pathological image typing and grading method based on deep learning and residual network technology, characterized in that, Specifically, it includes: Digital pathological images of tumor tissue were acquired and preprocessed. The tissue region was segmented based on pixel grayscale and color features. The segmented tissue region was then normalized by color and divided into multiple image blocks of the same size. A deep convolutional feature extraction network based on residual network is constructed. Image patches are input into the network for forward propagation. First intermediate feature maps and second intermediate feature maps are extracted from the shallow and deep layers of the network, respectively. The spatial size of the second intermediate feature map is smaller than that of the first intermediate feature map. Upsample the second intermediate feature map, calculate the fusion weights of corresponding pixels in the upsampled feature map and the first intermediate feature map, and perform weighted fusion of the two based on the fusion weights to obtain a fused feature map. Input the fused feature map into two parallel convolutional branches for processing, and output the first task feature map and the second task feature map. Calculate the feature deviation of each pixel in the second task feature map, generate a spatial attention map containing the spatial attention weights of each pixel in the second task feature map based on the feature deviation, and determine the region focus based on the statistical distribution of the attention weights; multiply the attention weights of each pixel in the spatial attention map element-wise with the feature vectors of the corresponding pixels in the first task feature map to obtain the enhanced first task feature map. Global pooling and classification are performed on the enhanced first task feature map and the second task feature map respectively to obtain tumor subtyping and grading results; a critical threshold is determined based on the regional focus degree, high attention points are selected from the spatial attention map according to the critical threshold, and a heat map of key tumor regions is generated based on the distribution of high attention points. For the second task feature map, the feature deviation of each pixel is calculated along the channel dimension. Specifically, for each pixel, firstly, the mean of all channel feature values ​​is calculated; then, the difference between the feature value of the pixel in each channel and the mean is squared and summed. The difference between the total number of channels and one is calculated. The summation result is divided by the difference to obtain the quotient. The square root of the quotient is then taken. The result is the feature deviation of the pixel. The feature deviation of each pixel is normalized to between 0 and 1 using the Sigmoid activation function, and used as the initial attention weight for that pixel to generate an initial spatial attention map. The initial spatial attention map is then subjected to Gaussian smoothing filtering to suppress noise, thus obtaining the spatial attention map. The mean and variance of the attention weights of all pixels in the spatial attention map are statistically analyzed. The region focus is determined based on the mean and variance. Specifically, the mean and variance are multiplied by their respective preset weights, and the product of the two is added together to obtain the region focus.

2. The tumor pathological image typing and grading method based on deep learning and residual network technology according to claim 1, characterized in that: The preprocessing specifically refers to image noise reduction preprocessing to eliminate light and shadow noise; The tissue region is segmented based on pixel grayscale and color features. Specifically, the grayscale value and RGB color channel value of each pixel in the image are first obtained. The pixel region that conforms to the tissue features is identified according to the preset threshold range. After removing noise and holes through morphological operations, a continuous tissue region mask is obtained. The logic for dividing the image blocks is as follows: the original image area covered by the tissue region mask is aligned with the standard pathological template image in terms of color distribution. The brightness and chromaticity channels are adjusted in the Lab color space through a color transfer algorithm to eliminate color differences caused by different staining batches. Finally, the normalized tissue region image is divided into grids according to a preset fixed size to generate multiple image blocks of the same size.

3. The tumor pathological image typing and grading method based on deep learning and residual network technology according to claim 2, characterized in that: A deep convolutional feature extraction network is constructed based on a residual network. The deep convolutional feature extraction network uses a pre-trained ResNet as its backbone, and removes its terminal global pooling layer and classification layer. The segmented image blocks are input into the deep convolutional feature extraction network for forward propagation. At a shallow layer of the network, specifically at the output of the second residual module, a first intermediate feature map is extracted, the size of which is denoted as [size not specified]. At a deeper level in the network, specifically at the output of the fourth residual module, the second intermediate feature map is extracted, and its size is denoted as... And satisfy , , ;in, , , These represent the height, width, and number of channels of the first intermediate feature map, respectively. , , These represent the height, width, and number of channels of the second intermediate feature map, respectively.

4. The tumor pathological image typing and grading method based on deep learning and residual network technology according to claim 3, characterized in that: The second intermediate feature map is upsampled using bilinear interpolation to make it the same size as the first intermediate feature map, thus obtaining the upsampled feature map. For any pixel in the upsampled feature map Its pixel value It is calculated as follows: The pixel is determined in the second intermediate feature map. The four closest pixels, based on pixel location Calculate the weight coefficients of the four pixels based on their relative positions to the four nearest pixels in both the horizontal and vertical directions; then sum the pixel values ​​of the four pixels after multiplying them by their corresponding weight coefficients to obtain the upsampled feature map pixels. Pixel values; Concatenate the upsampled feature map with the first intermediate feature map in channel-dimensional order, and input a... The convolutional layer generates an initial fusion weight map and performs Softmax normalization on the initial fusion weight map along the channel dimension to obtain a fusion weight map. Each spatial location in the fusion weight map corresponds to a normalized fusion weight vector, which contains two weight values, corresponding to the fusion weights of the upsampled feature map and the first intermediate feature map at that spatial location. The upsampled feature map and the first intermediate feature map are weighted and fused based on the fusion weight map. Specifically, for each spatial location, the fusion weight vector corresponding to that spatial location is obtained from the fusion weight map. The feature vector of the pixel at that spatial location in the upsampled feature map is multiplied by its fusion weight, and the feature vector of the pixel at that spatial location in the first intermediate feature map is multiplied by its fusion weight. The two are then added together, and the result is the feature vector of the fusion feature map at that spatial location. The feature vector refers to the vector composed of the feature values ​​of the pixel in all channels. Perform the above operations on all spatial locations to obtain a complete fused feature map; The fused feature map is simultaneously input into two parallel convolutional branches. The first convolutional branch consists of two 3×3 convolutional layers and one 1×1 convolutional layer, used to extract feature representations for tumor subtyping tasks and output a first task feature map. The second convolutional branch consists of two 3×3 convolutional layers and one 1×1 convolutional layer, with its convolutional kernel parameters independent of the first convolutional branch, used to extract feature representations for tumor grading tasks and output a second task feature map. The spatial dimensions of the first and second task feature maps are the same as those of the fused feature map, and the number of channels is determined by the number of output channels of the 1×1 convolutional layers in each branch.

5. The tumor pathological image typing and grading method based on deep learning and residual network technology according to claim 1, characterized in that: The attention weight of each pixel in the spatial attention map is multiplied element-wise with the feature vector of the pixel at the same spatial position in the first task feature map. Specifically, for each pixel in the spatial attention map, its attention weight is multiplied with the feature value of each channel in the feature vector of the corresponding pixel in the first task feature map to obtain the enhanced feature vector. The feature vector refers to the vector composed of the feature values ​​of a pixel across all channels; By iterating through all pixels in the spatial attention map, an enhanced first task feature map is obtained, whose spatial size is consistent with that of the first task feature map, and whose number of channels remains unchanged.

6. The tumor pathological image typing and grading method based on deep learning and residual network technology according to claim 5, characterized in that: The first task feature map is enhanced by global average pooling, which compresses its spatial dimension into a one-dimensional vector. This vector is then input into the first classifier for classification to obtain the tumor subtyping result. The second task feature map is also enhanced by global average pooling, which compresses its spatial dimension into a one-dimensional vector. This vector is then input into the second classifier for classification to obtain the tumor grading result. The first and second classifiers are each composed of a fully connected layer and a Softmax activation function, and the parameters of the two classifiers are independent of each other.

7. The tumor pathological image typing and grading method based on deep learning and residual network technology according to claim 1, characterized in that: The critical threshold is obtained by multiplying the region focus by a preset baseline threshold. For each pixel in the spatial attention map, its attention weight is compared with the critical threshold. If its attention weight is greater than or equal to the critical threshold, the pixel is marked as a high attention point. All high-attention points are mapped back to their corresponding spatial locations in the fused feature map, and a color gradient heatmap, i.e., a heatmap of key tumor regions, is generated based on their attention weights.

8. A tumor pathology image typing and grading device based on deep learning and residual network technology, used to execute the tumor pathology image typing and grading method based on deep learning and residual network technology as described in any one of claims 1-7, characterized in that, include: The preprocessing module acquires digital pathological images of tumor tissue and performs preprocessing. It segments the tissue region based on pixel grayscale and color features, normalizes the color of the segmented tissue region, and divides it into multiple image blocks of the same size. The feature extraction module is used to construct a deep convolutional feature extraction network based on residual networks. The image patch is input into the network for forward propagation, and the first intermediate feature map and the second intermediate feature map are extracted from the shallow and deep layers of the network, respectively. The spatial size of the second intermediate feature map is smaller than that of the first intermediate feature map. The fusion module is used to upsample the second intermediate feature map, calculate the fusion weight of the corresponding pixels of the upsampled feature map and the first intermediate feature map, and perform weighted fusion of the two based on the fusion weight to obtain the fused feature map. The fused feature map is then input into two parallel convolutional branches for processing, and the first task feature map and the second task feature map are output. The enhancement module is used to calculate the feature deviation of each pixel in the second task feature map, generate a spatial attention map containing the spatial attention weights of each pixel in the second task feature map based on the feature deviation, and determine the region focus based on the statistical distribution of the attention weights; multiply the attention weights of each pixel in the spatial attention map element-wise with the feature vectors of the corresponding pixels in the first task feature map to obtain the enhanced first task feature map. The classification and grading module is used to perform global pooling and classification processing on the enhanced first task feature map and the second task feature map respectively to obtain tumor subtyping and grading results; a critical threshold is determined based on the regional focus degree, high attention points are selected from the spatial attention map according to the critical threshold, and a heat map of key tumor regions is generated based on the distribution of high attention points.

Citation Information

Patent Citations

  • Fine-grained brain tumor classification method based on residual channel attention

    CN117173521A

  • Prostate cancer positioning system

    CN117456228A

  • Multi-scale residual error brain tumor image segmentation method based on attention mechanism

    CN118840552A