Pathological image classification system based on cross-scale grading self-attention
By introducing a cross-scale hierarchical self-attention mechanism in pathological image classification, segmenting the image into an attention window and calculating cross-scale attention, the problem of insufficient multi-scale context information capture in pathological image diagnosis is solved, and more efficient pathological image classification is achieved.
Patent Information
- Application Number
- CN202510190702.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art is difficult to effectively use deep learning models for pathological image diagnosis, especially due to the huge size of the pathological images and the lack of fine-grained annotation, which leads to insufficient models in capturing multi-scale context information.
A pathological image classification method based on the cross-scale hierarchical self-attention mechanism is proposed. By segmenting the entire pathological image into small-size attention windows, calculating cross-scale attention in each window, aggregating the features of different windows, and finally realizing global feature extraction and classification of multi-scale images.
This method can effectively capture cross-scale and different particle size context information within the global scope of pathological images, improving the accuracy of pathological image classification, especially in tasks such as molecular typing that require global information.
Smart Images

Figure CN120147698A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to the fields of deep learning, computer vision, pathological image diagnosis, etc. Background Art
[0002] Pathological diagnosis is one of the important criteria for cancer diagnosis. Whole Slide Images (WSI) is a digital image processing technology widely used in the field of pathological diagnosis. Traditional pathological sections are scanned by a digital scanner to obtain high-resolution digital images, and doctors perform pathological diagnosis by analyzing and evaluating the WSI. However, this method requires a large amount of manpower and time. Utilizing deep learning computer vision models and computer-aided pathological image classification models can provide more efficient and accurate disease diagnosis assistance tools for pathologists.
[0003] However, the main difficulties in applying deep learning models to pathological images lie in the excessively large image size of pathological images and the lack of fine-grained annotations. The image resolution of WSI is usually at the level of hundreds of millions of pixels and cannot directly apply conventional deep learning network models designed for natural images. Therefore, generally, the WSI is segmented into small pieces by a sliding window method and encoded as feature vectors for analysis, which can reduce the size of the input data and enable the deep learning model to process it. Due to the lack of fine-grained annotations in WSI, that is, the labels of each small image block cannot be obtained, weak supervision learning methods are usually used for pathological image classification.
[0004] Many current studies use the Multi-Instance Learning (MIL) method to handle the image classification and disease diagnosis problems of Whole Slide Images (WSIs). The pathological image classification model based on the MIL method uses the classification label of each high-resolution pathological image as the annotation used during training. During the model training process, these methods do not require annotating each small image patch in the whole pathological image. After training, the model can analyze the association between each small image patch and the final classification result, so as to discover the image parts and features with high correlation. However, the methods based on the MIL architecture are all based on the strong assumption that these small image patches are independent and identically distributed, which is inconsistent with the actual situation of WSI images in clinical practice: in actual WSI images, there is a certain association between each image patch, and referring to the context information around the reference target image patch is beneficial during the model prediction process, and doctors also refer to the information of surrounding images to make a judgment on the actual classification of the current image patch during the actual clinical diagnosis process. At the same time, the existing pathological image diagnosis methods also ignore the multi-scale characteristics of pathological images: each pathological image file directly stores the pathological image scanning results obtained at different magnification levels, and the change of these magnification levels is achieved by the combined switching of the magnification of the microscope eyepiece and objective lens, thus realizing lossless, optically-based multi-scale image acquisition, which is different from the natural image processing where the image size change is achieved based on algorithms, or different scale feature maps are extracted from different stages of deep learning models to achieve multi-scale image processing.
[0005] In order to model the cross-scale context information in multi-scale pathological images during the pathological image diagnosis process, the present invention proposes a pathological image classification method based on a cross-scale hierarchical attention mechanism. The method can capture the association information between each small image patch and the whole image, as well as the context information at multiple magnification levels, and can achieve different pathological image classification diagnosis tasks such as benign and malignant judgment, lymph node metastasis judgment, tumor typing (squamous cell carcinoma, adenocarcinoma, etc.), and molecular typing (such as POLE) on WSI images. Summary of the Invention
[0006] The present invention first proposes a cross-scale hierarchical self-attention method (CHSA) for multi-scale images. Its design is based on the multi-scale characteristics of multi-scale images, that is, in multi-scale images of huge sizes such as pathological images and remote sensing images, multiple images at different scanning magnifications are contained in the same image and the same image file. The cross-scale hierarchical self-attention method adopts the design idea of the windowed attention mechanism, uses the multi-scale characteristics of the image itself to divide the entire huge-scale pathological image into several small-sized attention windows, calculates the attention between different scales within each window respectively, aggregates the features of different windows, and calculates the attention between different windows again. Through the cross-scale hierarchical self-attention method, the problems faced by the traditional self-attention mechanism in pathological image classification, such as excessive image resolution, excessive number of divided image patches, and the resulting overly large attention matrix, can be solved. The pathological image diagnosis model based on the multi-instance architecture is built on the strong assumption that the individual image patches in the pathological image are independently and identically distributed. However, this does not conform to the actual situation of pathological image diagnosis, especially for tasks such as molecular typing where there is no clear judgment basis in the pathological image, and multi-faceted features such as cell morphology, tissue structure, and cell-cell interaction within the lesion area are all related to image classification. The relevance between image patches is more obvious and more important for making molecular typing judgments. The cross-scale hierarchical self-attention method can capture the potential correlations between all image patches within the global scope, provide richer information for the pathological image classification task, and better solve classification tasks with high requirements for global information such as molecular typing. The application scenarios of the cross-scale hierarchical self-attention method include, but are not limited to, any image classification, detection, segmentation, etc. tasks of images containing multi-scale characteristics such as pathological images and remote sensing images.
[0007] Based on the above cross-scale hierarchical self-attention CHSA algorithm, this patent proposes a multi-scale image classification model HiViT based on cross-scale hierarchical self-attention, which realizes dividing the entire huge-scale multi-scale image into several small-sized attention windows by using the multi-scale characteristics of the multi-scale image itself. First, cross-scale attention is calculated between different-scale images within the same window to fuse cross-scale features, then global attention is calculated between different windows at the smallest scale to fuse global image features at a lower computational cost, and finally a global feature vector representing the entire image is obtained, and classification of the image is realized on this vector.
[0008] The cross-scale hierarchical self-attention CHSA algorithm on the multi-scale image processes the multi-scale pathological images in the form of an image patch pyramid and a feature pyramid. Specifically, first, all images at all scales are cut into non-overlapping and non-gapped image patches using a specified image patch size. Then, for each image patch at the lowest magnification, all image patches at higher magnifications in the same relative position as this image patch are taken out and constructed into an image patch pyramid. After that, all image patches are encoded into feature vectors respectively through a feature encoding model, and a feature pyramid is constructed according to the original spatial position of the image patches, forming a multi-scale feature map describing the entire multi-scale image.
[0009] The cross-scale hierarchical self-attention CHSA algorithm on the multi-scale image takes the multi-scale feature map and the global classification query vector after feature mapping and position embedding as inputs, and each CHSA module is composed of multiple cross-scale self-attention (CSA) modules. The specific number and calculation process of the CSA modules it contains are related to the scanning magnification of the pathological image to be classified itself.
[0010] The multi-scale pathological image classification model HiViT is based on the cross-scale hierarchical self-attention method and is stacked by multiple consecutive CHSA modules, so as to capture the context information between different scales in the input multi-scale pathological image, and by adding a global classification query vector, the judgment of the pathological image classification is completed.
[0011] The cross-scale hierarchical self-attention CHSA algorithm is implemented as a cross-scale hierarchical self-attention module in the HiViT model, and each CHSA module is composed of multiple cross-scale self-attention (CSA) modules. The specific number and calculation process of the CSA modules it contains are related to the scanning magnification of the pathological image to be classified itself. The specific number and calculation process of the cross-scale self-attention are as follows: First, a cross-scale self-attention realizes feature aggregation between the global classification query vector and the feature vector with the lowest magnification and the largest receptive field in each feature pyramid through the multi-head self-attention method; the remaining cross-scale self-attention respectively performs feature aggregation again between two adjacent scales inside each feature pyramid by means of the cross-scale self-attention method; in the calculation process of the cross-scale hierarchical self-attention method layer, the order of calculation of each cross-scale self-attention is sequential, and in the next cross-scale hierarchical self-attention method layer, the calculation order is reverse calculation.
[0012] The cross-scale hierarchical self-attention CHSA module is divided into a Zoom-In CHSA module CHSA inZoom-Out Cross-Scale Hierarchical Self-Attention (CHSA) Module CHSA out Calculate the self-attention between the features of two adjacent scale levels in the reverse order respectively. CHSA in The CHSA module calculates the cross-scale attention from the classification query vector at the highest level to the image features at the lowest level (the most numerous and highest magnification), CHSA out Calculate the cross-scale attention in the opposite direction (from the image features at the lowest level to the classification query vector at the highest level).
[0013] A pathological image assisted diagnosis system applies the cross-scale hierarchical self-attention CHSA algorithm on the multi-scale image and the multi-scale image classification model HiViT. The present invention proposes a pathological image assisted diagnosis system, which is composed of a patient data management module, a pathological image classification model, and a visualization and assisted diagnosis module.
[0014] The patient data management module consists of a pathological image encoding module and a storage database. The pathological image encoding module includes a multi-size pathological image encoding module and a feature serialization input modeling module, which perform operations such as segmenting, encoding into feature vectors, and serializing into one-dimensional feature vector sequences on the patient's pathological images, and store them in the database together with the original pathological images, patient basic information, medical history, diagnosis results, etc.
[0015] The pathological image classification module is based on the multi-scale image classification model and performs the user-specified pathological image classification tasks on the pathological images feature-encoded by the patient data management module, specifically including but not limited to judging the presence or absence of lymph node metastasis, classifying the benign and malignant of tumors, molecular typing, etc.
[0016] The visualization and assisted diagnosis module includes a classification result visualization module and a result display interface. Among them, the classification result visualization module extracts the attention weights at each level from the multi-scale image classification model based on the visualization method of the cross-scale hierarchical self-attention CHSA, and calculates the correlation between the image patches at each scale and the global classification query vector. These attention weights are plotted as heatmaps and displayed in the result display interface. In addition, the image patches with the highest attention weights are also extracted by this module and displayed in the result display interface.
[0017] The technical effects to be achieved by the present invention are as follows:
[0018] 1. The present invention proposes a multi-scale image classification model HiViT (Histopathological Image Vision Transformer) and a cross-scale hierarchical self-attention algorithm Cross-scale Hierarchical Self-Attention, CHSA, for multi-scale images. The HiVIT model and the CHSA algorithm can be applied to tasks such as image classification, detection, and segmentation of any images with multi-scale characteristics, including but not limited to pathological images and remote sensing images.
[0019] The HiViT model takes the cross-scale hierarchical self-attention algorithm CHSA as the core, can establish associations between all image patches of different scales, capture cross-scale and different-granularity context information within the global range of pathological images, thereby improving the performance of pathological image classification tasks. Especially compared with traditional multi-instance learning methods, this method is particularly suitable for tasks such as tumor molecular typing, which are difficult for doctors to identify, have no clear key discriminant basis, and require referring to the comprehensive performance of images within the global range. The cross-scale hierarchical self-attention CHSA algorithm can, through the visualization of attention weights, achieve the localization of fine tumor regions, classification-related image regions, and image features, thereby explaining the basis for the model's judgment.
[0020] 2. Applying the above method and model, this patent proposes a pathological image-assisted diagnosis system. The pathological image-assisted diagnosis system consists of a patient data management module, a pathological image classification model, and a visualization and assisted diagnosis module.
[0021] Among them, the patient data management module consists of a pathological image encoding module and a storage database. The pathological image encoding module includes a multi-size pathological image encoding module and a feature serialization input modeling module, which perform processes such as segmenting, encoding into feature vectors, and serializing into one-dimensional feature vector sequences for the patient's pathological images, and store them in the database together with the original pathological images, patient basic information, medical history, diagnosis results, etc.
[0022] The pathological image classification module, based on the multi-scale image classification model, performs user-specified pathological image classification tasks on the pathological images feature-encoded by the patient data management module, specifically including but not limited to judging whether there is lymph node metastasis, classifying the benign and malignant of tumors, molecular typing, etc.
[0023] The visualization and auxiliary diagnosis module includes a classification result visualization module and a result display interface. Among them, the classification result visualization module extracts attention weights at all levels from the multi-scale image classification model based on the visualization method of the cross-scale hierarchical self-attention (CHSA), and calculates the correlation between image patches at each scale and the global classification query vector. These attention weights are plotted as heatmaps and displayed in the result display interface. In addition, the image patches with the highest attention weights are also extracted by this module and displayed in the result display interface.
[0024] The pathological image auxiliary diagnosis system can help doctors determine the benign and malignant nature of pathological images, different types of tumors, and even the genotype classification of tumors during actual clinical diagnosis. Moreover, the system can visualize the obtained diagnosis results by analyzing and displaying the correlation between the final classification category and each small image patch and image feature, and prompt doctors to pay special attention to image regions and image features. Brief Description of the Drawings
[0025] Figure 1 The architecture of the pathological image classification method based on the cross-scale hierarchical self-attention method;
[0026] Figure 2 The principle of the cross-scale hierarchical self-attention method (CHSA) for multi-scale images;
[0027] Figure 3 The display principle of the visualization module;
[0028] Figure 4 Examples of visualization results. Detailed Embodiments
[0029] The following are the preferred embodiments of the present invention in combination with the accompanying drawings, and the technical solutions of the present invention will be further described, but the present invention is not limited to this embodiment.
[0030] The present invention proposes a new cross-scale hierarchical self-attention algorithm, Cross-scale Hierarchical Self-Attention (CHSA), for multi-scale images. This method divides the attention window based on the multi-scale characteristics inherent in pathological images, restricting the self-attention calculation on high magnification, large-scale pathological images, which has the highest computational cost, within the attention window, while the global self-attention calculation on the entire pathological image is only restricted to the image with the lowest magnification and lowest scale, thereby significantly reducing the computational overhead of the self-attention mechanism on multi-scale and huge-sized pathological images and enabling the application of the self-attention mechanism in the field of pathological image classification. In addition, through the above design method, the multi-scale characteristics inherent in pathological images, that is, the same pathological image contains scanned images at different magnifications, are naturally introduced into the model design. Based on the cross-scale hierarchical self-attention algorithm CHSA, the present invention proposes a pathological image diagnosis model, HiViT (Histopathological Image Vision Transformer). The HiViT model captures the correlation between all image patches of different sizes and at different magnifications in the entire pathological image through the cross-scale hierarchical self-attention algorithm CHSA, thereby realizing the modeling of cross-scale global context information and ultimately improving the performance of the pathological image classification task, especially for tasks such as tumor molecular typing that are difficult for doctors to identify, lack clear key discriminant criteria, and require referring to the comprehensive performance of the image globally.
[0031] Based on the pathological image classification method, this patent proposes a pathological image auxiliary diagnosis system. The auxiliary diagnosis system is oriented to the actual clinical pathological diagnosis process and consists of a patient data management module, a pathological image classification module, and a visualization and auxiliary diagnosis module. Among them, the pathological image classification module is based on the pathological image classification method; the visualization and auxiliary diagnosis module, based on the visualization of the cross-scale hierarchical self-attention proposed in this patent, can realize the visualization of the obtained diagnosis results by analyzing and displaying the correlation between the finally classified category and each small image patch and image feature, prompting the image regions and image features that doctors need to pay special attention to. The auxiliary diagnosis system can help doctors determine the presence or absence of cancer cells, tumor classification, and tumor gene typing, etc. in the pathological whole slide image (WSI) during the actual clinical diagnosis process. In addition, the system can also analyze and display the image regions and image features based on which the model makes judgments.
[0032] Multi-scale pathological image classification model HiViT
[0033] The architecture of the multi-scale pathological image classification model HiViT is as Figure 1As shown in the figure, it includes the following four components: 1. A multi-scale pathological image encoding module, which is used to convert multi-scale pathological images into a set of multi-scale high-dimensional feature maps with spatial correspondence; 2. A multi-scale image feature enhancement module, which is used to randomly shuffle the obtained multi-scale high-dimensional feature maps within two different ranges of global and local, so as to introduce global feature permutation invariance while retaining the spatial correspondence at the fine-grained level, and achieve feature enhancement; 3. A feature serialization input modeling module with position information preservation, which is used to flatten the two-dimensional feature map into a one-dimensional input vector sequence, and in this process, the two-dimensional spatial structure and blank image block features are removed to reduce the video memory consumption and computational overhead, while retaining the position information of each image block and the correlation between image blocks; 4. A multi-scale pathological image classification model HiViT, based on the cross-scale hierarchical self-attention mechanism CHSA, which is used to perform feature aggregation and classification on the obtained encoded pathological image feature sequence, and this model has the ability to capture and analyze the multi-scale global context.
[0034] Multi-scale pathological image encoding module
[0035] In order to process high-resolution multi-scale pathological images with reasonable computational overhead, the first step of the pathological image classification method is to perform image tiling and feature encoding operations on large-sized pathological images. For this process, the present invention proposes a method of "image block pyramid" and "feature pyramid" for modeling multi-scale pathological images. For each pathological image, this module, in an offline (independent of model training, as image preprocessing) form, crops the images at different scanning magnifications in the pathological image into small image blocks of the same size. These image blocks with different scanning magnifications are arranged in the form of an "image block pyramid" and encoded into a "feature pyramid" through a feature encoding model, so as to convert the multi-scale pathological images into a set of multi-scale high-dimensional feature maps that occupy less video memory space and have spatial correspondence.
[0036] Specifically, all the image blocks with higher magnifications corresponding to each image block with the lowest magnification (corresponding to the largest image range) cut out from the whole pathological image will be taken out and form an image block pyramid. Taking the pathological image X =
[0037] X 1 ,X 2 ,X 3 as an example (where, is the image scanned at 5×, and (which are images scanned at 10× and 20× respectively). The cropping and arrangement method of the image patch pyramid is as follows: First, the X is cropped into image patches {a 1 of the same size (this size can be arbitrarily specified, including 224×224, 256×256, and 512×512, etc., hereinafter represented by w×h) in a non-overlapping and non-gapped manner, 11 , …, a ij , a CR}. Among them, represents the image patch located at the (i, j) position, represents the total number of rows of the feature map, represents the number of columns of the feature map. The field of view (FoV) of the feature a ij at 5× magnification is w×h in size, and this field of view is 2w×2h at 10× magnification. The 10× image within this field of view is cropped into image patches {b ij(1) , …, b ij(4)} of the same w×h size. Similarly, the 20× image within this field of view is cropped into image patches {c ij(1) , …, c ij(16)}. We define x = {a ij , b ij(1) , c ij(1) , …, c ij(4) , b ij(2) , c ij(5) , …, c ij(16)} as an image patch pyramid. These image pyramids {x ij} are then encoded by a pre-trained feature encoding model into a feature pyramid {f ij} = {α 11 , β 11(1) , γ 11(1) , …, γ 11(4) , β 11(2) , γ 11(5) , …, γ 11(16)}, and represent the features extracted from images a, b, and c respectively. The specific implementation of the pre-trained feature encoding model can include but is not limited to ImageNet or ResNet, EfficientNet, or Vision Transformer pre-trained on pathological image patches. Through the above method, the features of the entire large-scale multi-scale pathological image are encoded into a high-dimensional feature map F = {f 11 , f 12 , …, f CR}. Additionally, during this process, threshold filtering is performed on the image patches respectively to discover blank image patches that do not contain tissues and cells. The features of these blank image patches are directly set to 0 and will be removed in the subsequent feature serialization input modeling module.
[0038] The multi-scale pathological image encoding module processes multi-scale pathological images using the "image patch pyramid" and "feature pyramid" methods proposed in this patent, thereby achieving dimensionality reduction of the pathological images. Specifically, taking pathological images with 5×, 10×
[0039] , 20× as an example, first, all pathological images at all scales are cut into non-overlapping and non-gapped image patches using an arbitrarily specified image patch size. Each image patch at the lowest magnification (5×) and all image patches at higher magnifications (10×, 20×) at the same relative position as this 5× image patch are taken out and a image patch pyramid is constructed. Each image patch pyramid includes 21 image patches (1 5×, 4 10×, 16 20×).
[0040] All image patches at different scales cut from the pathological image are encoded into feature vectors through a feature encoding model (this feature encoding model can adopt different pre-trained models such as the ResNet-50 or EfficientNet model pre-trained on ImageNet). The 21 feature vectors from image patches at different scales within each image patch pyramid form a feature pyramid. For each multi-scale pathological image, by splicing different feature pyramids two-dimensionally according to the spatial position, a multi-scale feature map describing the entire multi-scale pathological image is formed.
[0041] Multi-scale image feature enhancement module
[0042] The multi-scale image feature enhancement module proposed in the present invention is based on the feature enhancement method (Multi-Scale Feature Augmentation, MSFeatAug) on the multi-scale image features, aiming to make up for the deficiencies of the long-term lack of data and feature enhancement methods in the pathological image diagnosis model. The feature enhancement method MSFeatAug includes the following two specific feature enhancement operations, introducing randomness to the data at different scales while maintaining the fine-grained local context: 1. Coarse-grained random shuffling implemented on the multi-scale high-dimensional feature map, shuffling each feature pyramid f 11 , f 12 , …, f CRRearrange in a random order; 2. Fine-grained feature pyramid shuffling implemented inside each feature pyramid. Inside each feature pyramid f ij inside, shuffle β ij ( 1) , …, β_ij(4) and γ ij(1) , …, γ ij(4) in order. By this method, transformation invariance can be introduced between global features, and fine-grained local cross-scale context information can be retained, which helps to improve the performance of the HiViT model. By introducing randomness to the input data in this way, the spatial structure and correspondence between levels in the feature pyramid are retained at the same time.
[0043] That is, the multi-scale image feature enhancement module is based on the feature enhancement method on the multi-scale image features proposed in this patent. First, by taking each feature pyramid as a whole, randomly shuffle the relative positions of these feature pyramids, but the inside of each feature pyramid is not affected, and features are not exchanged between feature pyramids. By introducing randomness to the input data in this way, the spatial structure and correspondence between levels in the feature pyramid are retained at the same time. Then, the module randomly arranges features between each scale level inside each feature pyramid respectively.
[0044] Feature serialization input modeling module with position information preserved
[0045] The feature serialization input modeling module with position relationship preserved is based on the Position Preserving Sequential Input Modeling (PPSIM) method for multi-scale image serialization feature modeling with spatial position preserved, and is composed of a feature mapping and a feature serialization modeling sub-module. PPSIM removes unnecessary two-dimensional spatial structures and redundant blank image blocks while retaining cross-scale position correlations, so as to reduce the computational overhead and speed up the calculation speed.
[0046] The feature mapping sub-module IP uses multiple 1×1 convolutional kernels and normalization layers to map the input feature sequence and reduce the spatial dimension of each input feature vector:
[0047] IP(F′) = conv dd (ReLU(LN(conv Dd (LN(F′)))))
[0048] where F′ represents the multi-scale high-dimensional feature map after feature enhancement, and conv DdIt is indicated that the input dimension of this convolutional layer is D and the output dimension is d. Usually, d is set to be smaller than D to compress the spatial dimension of the features. Since even for different pathological image classification tasks, the above-mentioned pathological image features are extracted by a frozen and non-learnable feature extraction model, it may not be optimal for each task. Therefore, compared with the feature mapping adopted in the existing methods, the feature mapping of this method is more complex and is placed before the position encoding to be used as an adapter for the input features.
[0049] In the feature serialization modeling sub-module, the projected feature map is added to the position encoding (including but not limited to simple sinusoidal positions) to preserve the position information, and then the feature map is flattened:
[0050] S = Flatten(IP(F') + PE(F')),
[0051] At the highest feature level, "position" is the coordinate of the randomly shuffled image patch in the whole image. And at lower levels, "position" refers to the relative position within each feature pyramid. The feature sequence is flattened in a specific order to maintain the hierarchical relationship between the image patches, as shown specifically in Figure 1 (c).
[0052] The feature input modeling module is based on the serialized feature modeling method of the multi-scale image with spatial position retention proposed in this patent and consists of a feature serialization modeling sub-module and a feature mapping module.
[0053] The feature serialization modeling sub-module first adds spatial encoding to the feature vectors according to the positions and scale levels of the image patches corresponding to the feature vectors in the original pathological image to preserve the position information of the image patches and the correlation between each image patch. The multi-scale feature map with position encoding is flattened into a one-dimensional feature vector sequence with specific corresponding relationships between different scales, and at the same time, the feature vectors corresponding to the blank image patches are directly removed to reduce the video memory overhead. Through feature serialization modeling, the two-dimensional spatial structure of the input feature map is removed, thus effectively reducing the video memory overhead, but the spatial relationship between each image patch is retained.
[0054] The feature mapping sub-module uses multiple 1×1 convolutional kernels and normalization layers to map the input feature sequence and reduce the spatial dimension of each input feature vector. At the same time, since the feature extraction of image patches is implemented in an offline preprocessing form that is not trainable and independent of model training, the feature mapping sub-module can serve as a learnable adapter for these features, while keeping the pre-model weights in feature extraction unchanged and retaining the knowledge of the pre-trained model, enabling the extracted features to adapt to different tumor types (such as lung cancer, melanoma, colorectal cancer, gastric cancer, etc.) and classification task types (such as primary tumor differential diagnosis, lymph node metastasis judgment, molecular typing, etc.).
[0055] Multi-scale Pathological Image Classification Model HiViT
[0056] The multi-scale pathological image classification model HiViT proposed by the present invention is based on the Cross-Scale Hierarchical Self-Attention (CHSA) mechanism, establishing cross-scale direct connections between image patches at different magnifications and in different ranges, so as to reduce the computational memory overhead and achieve the capture of context information between every two adjacent scales in the input multi-scale pathological image.
[0057] The HiViT model takes the image patch feature sequence after feature mapping and position embedding as input, and feeds a global classification query vector to the feature sequence to achieve the classification of the entire pathological image. During the calculation process of the HiViT model, multiple CHSA modules are applied to these image patch features. Among these features, the classification query vector cls is defined as the highest-level feature vector because its field of view (FoV) is the entire pathological image, followed by the image features of 5×, 10×, and 20×. The dimension of each feature vector remains unchanged before and after each CHSA block. The CHSA module is divided into two types, the Zoom-In CHSA module CHSA in and the Zoom-Out CHSA module CHSA out, are alternately placed in the model in sequence. The HiViT model is consistent with the architecture of the Vision Transformer (ViT) model. However, the key difference is that the standard Multi-head Self-Attention (MSA) module in ViT is replaced by the CHSA module we proposed to achieve the processing of multi-scale images and reduce the computational overhead for pathological images with huge resolutions, avoiding the huge global self-attention matrix that existing GPUs cannot fully handle. Compared with existing pathological image classification models based on the Transformer architecture, the HiViT model does not rely on the two-dimensional spatial structure and can utilize the multi-scale features of pathological images to learn cross-scale features of images from different ranges. The HiViT model includes a total of three pairs of focused and dispersed CHSA modules, with a total of six layers.
[0058] The multi-scale pathological image classification model HiViT is based on the cross-scale hierarchical self-attention mechanism (CHSA) proposed in this patent and is stacked by multiple consecutive CHSA layers, so as to capture the context information between different scales in the input multi-scale pathological images, and complete the judgment of pathological image classification by adding a global classification query vector.
[0059] The cross-scale hierarchical self-attention mechanism is divided into multiple cross-scale self-attention (CSA). Specifically, the number and calculation process of the cross-scale self-attention are related to the scanning magnification of the pathological image itself to be classified. Taking pathological images at three different scanning magnifications, 5×, 10×, and 20× as an example, the cross-scale hierarchical self-attention mechanism CHSA contains three cross-scale self-attention CSAs 1 , CSA 2 and CSA 3 . Its calculation process is that CSA 1 realizes feature aggregation between the global classification query vector and the feature vector (5×) with the lowest magnification and the largest receptive field in each feature pyramid through the multi-head self-attention mechanism; CSA 2 and CSA 3Feature aggregation is performed again between two adjacent scales (5× and 10×, as well as 10× and 20×) within each feature pyramid using the cross-scale self-attention mechanism CSA. Compared with the method of constructing a huge attention matrix between feature vectors at all scales, the CHSA method splits the attention calculation into three small cross-scale self-attention CSA calculation processes that are only between two adjacent scales and are confined within each attention window, thus greatly reducing the video memory overhead during calculation. These three CSA modules are placed in the opposite order in alternating CHSA layers. During the calculation process of the above cross-scale hierarchical self-attention mechanism CHSA, the calculation order of each cross-scale self-attention CSA is CSA 1 , CSA 2 and CSA 3 . In the next CHSA layer of the model, the calculation order is CSA 3 , CSA 2 and CSA 1 .
[0060] Cross-scale Hierarchical Self-Attention (CHSA)
[0061] The cross-scale hierarchical self-attention mechanism (Cross-Scale Hierarchical Self-Attention, CHSA) proposed by the present invention is divided into two types, the Zoom-In CHSA module CHSA in and the Zoom-Out CHSA module CHSA out , which calculate the self-attention between the features at two adjacent scale levels in the opposite order respectively. The CHSA in module calculates the cross-scale attention from the highest-level classification query vector to the lowest-level (the most numerous and largest magnification) image features, and the CHSA out calculates the cross-scale attention in the opposite direction (from the lowest-level image features to the highest-level classification query vector). Taking a pathological image that scans the image at magnifications of 5×, 10×, and 20× as an example, the Zoom-In CHSA module CHSA in includes three sub-modules based on cross-scale attention (Cross-Scale Self-Attention, CSA): CSA 1 , CSA 2 and CSA 3 . First, CSA 1 calculates the attention between the highest-level global classification query vector and all 5× image features in the whole pathological image:
[0062] CSA 1 (seq 12 ) = MSA(seq 12 ) = Concat(head 1 , head 2 , …, head h )W O ,
[0063]
[0064] where seq 12 = {cls, α 11 , …, α CR}}, is the mapping matrix in self-attention calculation, h = 8 represents the number of attention heads, and d k = d / h represents the feature dimension in each self-attention head.
[0065] Next, CSA 2 and CSA 3 calculate cross-scale attention between 5× and 10× image patches, and between 10× and 20× image patches inside each feature pyramid respectively:
[0066] CSA 2 (seq 23 ) = Concat(MHA(α 11 , β 11(1) , …, β 11(4) ), …, MHA(α CR , β CR ( 1) , …, β CR(4) ))),
[0067] CSA 3 (seq 34 ) = Concat(MHA(β 11(1) , γ 11(1) , …, γ 11(4) ), …, MHA(β 11(4) , γ 11(13) , …, γ 11(16) ), …), where seq 23 = {α 11 , β 11(1) , …, β 11(4) , α 12 , …, β CR(4)}}, seq 34 = {β 11(1) , …, β 11(4) , γ 11(1),…,γ 11(16) ,β 12(1) ,…,βCR (16)}。Dispersed (Zoom-On) CHSA module CHSA out Consisting of cross-scale attention CSA sub-modules placed in the reverse order of the CHSA in module: CSA 3 , CSA 2 and CSA 1 . In addition to the above-mentioned CSA attention calculation, each CSA sub-module is followed by two linear mapping layers and two intermediate GELU non-linear layers after each CSA attention calculation. Compared with the method of constructing a huge attention matrix between feature vectors at all scales, the CHSA method splits the attention calculation into three small cross-scale self-attention CSA calculation processes that are only calculated between adjacent two scales and limited within each attention window, thus greatly reducing the video memory overhead during calculation. These three CSA modules are placed in the reverse order in mutually alternating CHSA layers.
[0068] Pathological image classification and auxiliary diagnosis system based on cross-scale hierarchical self-attention mechanism
[0069] Based on the above-mentioned cross-scale hierarchical attention-based image classification method, the present invention designs a pathological image classification and auxiliary diagnosis system, which specifically includes: 1. Patient data management module; 2. Pathological image classification model; 3. Visualization and auxiliary diagnosis module.
[0070] Patient data management module
[0071] The patient data management module consists of a pathological image encoding module and a storage database. The pathological image encoding module directly adopts the multi-size pathological image encoding module and the feature serialization input modeling module in the above-mentioned cross-scale hierarchical attention mechanism-based image classification method to convert the patient's pathological image into a multi-scale feature sequence, and stores it in the database together with the original pathological image, patient basic information, medical history, diagnosis results, etc.;
[0072] Pathological image classification module
[0073] The pathological image classification module directly uses the multi-scale pathological image classification model HiViT in the above-mentioned cross-scale hierarchical attention mechanism CHSA-based image classification method to perform user-specified classification tasks on the pathological images selected by the user of the auxiliary diagnosis system and feature-encoded by the patient data management module, including judgment of lymph node metastasis, classification of tumor benign and malignant, molecular typing, etc.;
[0074] Visualization and auxiliary diagnosis module
[0075] The visualization and auxiliary diagnosis module includes a classification result visualization sub-module and a result display interface. The classification result visualization sub-module is based on the visualization method of the cross-scale hierarchical self-attention,
[0076] extracts the attention weights at each level in the calculation process of the cross-scale hierarchical self-attention mechanism CHSA from the HiViT model in the pathological image classification module, and calculates the association between the image patches at each scale and the global classification query vector, as Figure 2 shown. These attention weights are plotted as heatmaps and displayed in the result display interface, so as to reflect the strength of the association relationship between each image patch and the classification vector at each level, and interpret the image parts and image features on which the model classification is based. In addition, the image patch with the highest attention weight is also extracted by this module and displayed in the result display interface. The image patch with the highest attention weight is often the image patch or representative image patch on which the classification is based, that is, the tumor cell image in the benign and malignant classification. However, for the molecular typing task, there are often no representative image patches or representative image features.
[0077] The patient data management module consists of a pathological image encoding module and a storage database. The pathological image encoding module directly uses the multi-scale pathological image encoding module and the feature serialization input modeling module in the image classification method based on the cross-scale hierarchical attention mechanism to convert the patient's pathological image into a multi-scale feature sequence, and stores it in the database together with the original pathological image, patient's basic information, medical history, diagnosis results, etc.;
[0078] The pathological image classification module directly uses the multi-scale pathological image classification model HiViT in the image classification method based on the cross-scale hierarchical attention mechanism CHSA to perform user-specified classification tasks on the pathological images selected by the user of the auxiliary diagnosis system and feature-encoded by the patient data management module, including the judgment of whether there is lymph node metastasis, the benign and malignant classification of tumors, molecular typing, etc.;
[0079] The visualization and auxiliary diagnosis module includes a classification result visualization sub-module and a result display interface. The classification result visualization sub-module, based on the visualization method of cross-scale hierarchical self-attention, can extract the attention weights at each level in the calculation process of the cross-scale hierarchical self-attention mechanism CHSA from the HiViT model, and thereby calculate the correlation between the image patches at each scale and the global classification query vector. These attention weights are plotted as heatmaps and displayed in the result display interface, so as to reflect the strength of the correlation between each image patch and the classification vector at each level, and interpret the image parts and image features on which the model classification is based. In addition, the image patch with the highest attention weight is also extracted by this module and displayed in the result display interface. The image patch with the highest attention weight is often the image patch or representative image patch on which the classification is based, that is, the tumor cell image in the benign and malignant classification. However, for the molecular typing task, there are often no representative image patches or representative image features.
Claims
1. A pathological image classification system based on cross-scale hierarchical self-attention, characterized in that: The pathological image is input and the cross-scale hierarchical self-attention algorithm on the multi-scale image is applied to process the image. Based on the multi-scale characteristics of the multi-scale image itself, the whole image is divided into several small-sized attention windows. First, the cross-scale attention is calculated between the different scale images in the same window to fuse the cross-scale features. Then, the global attention is calculated between the different windows of the smallest scale to fuse the global image features. Finally, the global feature vector representing the whole image is obtained, and the image is classified on this vector.
2. A pathological image classification system based on cross-scale hierarchical self-attention as described in claim 1, characterized in that: Multi-scale pathological images are processed using image block pyramid and feature pyramid.
3. A pathological image classification system based on cross-scale hierarchical self-attention as described in claim 2, characterized in that: The specific method for processing multi-scale pathological images in the form of image block pyramid and feature pyramid is as follows: first, the images at all scales are cut into image blocks without overlap and gaps using the specified image block size, and then for each image block with the lowest magnification, all image blocks with higher magnifications that are in the same relative position as the image block are taken out and constructed into an image block pyramid; then, all image blocks are encoded into feature vectors through a feature encoding model, and constructed into a feature pyramid according to the original spatial positions of the image blocks, forming a multi-scale feature map that describes the entire multi-scale image.
4. A pathological image classification system based on cross-scale hierarchical self-attention as described in claim 3, characterized in that: The multi-scale feature map and the global classification query vector that have been feature mapped and position embedded are taken as input, and each cross-scale hierarchical self-attention module that implements the cross-scale hierarchical self-attention algorithm is composed of multiple cross-scale self-attention modules, and the specific number and calculation process of the multiple cross-scale self-attention modules included are related to the scanning multiples of the pathological image itself that needs to be classified. Taking the pathological images including scanned images at 5×, 10×, and 20× magnification as an example, the cross-scale hierarchical self-attention module includes three cross-scale self-attention modules CSA1, CSA2, and CSA3. The calculation process is that CSA1 calculates multi-head self-attention between the global classification query vector and the feature vector (5×) with the lowest magnification and the largest receptive field in each feature pyramid: CSA1(seq 12 )=MSA(seq 12 )=Concat(head1,head2,…,head h )W O , Among them, seq 12 ={cls,α 11 ,…,α CR }, is the mapping matrix in the self-attention calculation, h = 8 represents the number of attention heads, d k =d / h represents the feature dimension in each self-attention head. Next, CSA2 and CSA3 calculate cross-scale attention between 5× and 10× image blocks, and between 10× and 20× image blocks within each feature pyramid: CSA2(seq 23 )=Concat(MHA(a 11 ,b 11(1) ,…,b 11(4) ),…,MHA(a CR ,b CR(1) ,…,b CR(4) )), CSA3(seq 34 )=Concat(MHA(β 11(1) ,c 11(1) ,…,c 11(4) ),…,MHA(β 11(4) ,c 11(13) ,…,c 11(16) ),…,), among them, seq 23 ={a 11 ,b 11(1) ,…,b 11(4) ,a 12 ,…,b CR(4) },seq 34 ={β 11(1) ,…,b 11(4) ,c 11(1) ,…,c 11(16) ,b 12(1) ,…,c CR(16) }。 5. A pathological image classification system based on cross-scale hierarchical self-attention as described in claim 4, characterized in that: The cross-scale hierarchical self-attention module is divided into two types, focusing on the cross-scale hierarchical self-attention module CHSA in and decentralized cross-scale hierarchical self-attention module CHSA out , respectively, calculate the self-attention between the features of two adjacent scale levels in reverse order; CHSA in Compute cross-scale attention from the highest-level classification query vector to the lowest-level image features, CHSA out Compute cross-scale attention from the lowest-level image features to the highest-level classification query vector; Taking pathological images including scanned images at 5×, 10×, and 20× magnification as an example, the focused cross-scale hierarchical self-attention module CHSA in The method comprises three sequentially placed cross-scale attention CSA modules CSA1, CSA2 and CSA3 as claimed in claim 3, wherein the decentralized cross-scale hierarchical self-attention module CHSA out By and the CHSA in The cross-scale attention CSA modules CSA3, CSA2 and CSA1 are placed in reverse order.
Citation Information
Cited By
Image interpretable classification method and device, computer equipment and storage medium
CN120783103A
Method and device for detecting pan-tumor lymph node metastasis based on bidirectional reciprocity cross-scale attention fusion
CN120976643A