Hierarchical Representation Method of Pathological Images Based on Dual-Stream Cross-Attention Network
The low-high resolution blocking features of pathological images are processed through the dual-current cross-attention network, and the semantic gap and computational cost problems in the global representation learning of pathological images are solved, and efficient pathological image hierarchical feature extraction and differentiation improvement are achieved.
Patent Information
- Application Number
- CN202211439389.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-17
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-11-17
AI Technical Summary
When processing gigapixel pathological images, it is difficult to effectively deal with semantic gaps between different scale features and high computational cost, resulting in inefficient learning of global representation of pathological images.
The dual-current cross-attention network is adopted to process the feature learning of low-resolution and high-resolution image blocking through two branches, and the feature fusion is used to reduce the number of features in high-resolution blocking and improve the effectiveness and distinction of pathological image hierarchical features.
Improves the effectiveness and distinction of pathological image hierarchical features, reduces computational costs, and performs well in downstream prognostic tasks.
Smart Images

Figure CN115908997B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of pathological image processing. More specifically, it relates to a method for hierarchical representation of pathological images based on a two-stream cross-attention network. Background Art
[0002] Representation learning based on pathological images has always been a very challenging problem, which aims to study how to extract effective global image representations from pathological images (Whole-Slide Image, abbreviated as WSI) stained with special stains (such as hematoxylin-eosin staining) for downstream prediction tasks. Different from ordinary images, pathological images at the maximum magnification are generally gigapixel-level, which can provide rich pathological visual information, such as tissue phenotypes and tissue microenvironments, and even individual cells. These multi-level information can provide important bases for pathologists to evaluate diseases, but at the same time, it also brings great challenges to the global representation learning of pathological images. On the one hand, classical computer vision models can only process ordinary-sized images and cannot be directly applied to gigapixel-level pathological images. On the other hand, due to the extremely high resolution but low information density of pathological images, the representation learning models based on pathological images need to be able to efficiently and automatically extract discriminative visual features from ultra-high-resolution images.
[0003] To solve the above problems, the representation learning methods based on pathological images often include three necessary steps: patch division of pathological images, patch feature extraction, and multi-instance learning based on patches. Among them, multi-instance learning based on patches aims to learn a global bag-level (or whole-image-level) representation from the patch bags of each pathological image. Early technologies almost all only used single-resolution pathological images. They usually first learned the patch dependencies of specific structures to obtain new patch-level features, and then aggregated these features to obtain the final bag-level representation. For example, Li et al. first used a graph to describe the patch structure, then learned the graph node representations where the patches are located with the help of a graph neural network, and finally used the node pooling operation to obtain the global representation of the pathological image. Later, more patch structures, such as clustering and sequences, were tried for global representation learning of pathological images. These technologies have all achieved good results in common downstream tasks of pathological images (such as cancer subtype classification, cancer prognosis risk prediction).
[0004] In recent years, scholars have gradually started to shift their attention from single-scale images to multi-scale visual objects, such as the image pyramid structure of pathological images, aiming to make full use of the inherent hierarchical semantic information in pathological images, such as cells, tissue microenvironments, and tissue phenotypes, to better represent pathological images. For example, Li et al. and Hou et al. respectively used sequences and heterogeneous graphs to describe the patch structure and explored learning a hierarchical global representation of pathological images using two-resolution images. Figure 1Schematic diagram of a hierarchical representation method for pathological images based on sequences and heterogeneous graphs. As Figure 1 shown, in this method, images of two resolutions are used to extract features based on sequences and heterogeneous graphs respectively, and then fused to obtain the global representation of pathological images. Although these methods have achieved better results than single-resolution images, there are still two main problems with this type of pathological image pyramid-based scheme. First, in the feature fusion process, the inherent semantic gap between features of different scales is not properly processed, but is limited to simple multi-scale vector splicing or direct message passing between features of different scales. Second, a large number of high-resolution blocks explicitly participate in the entire block vector learning process, resulting in high computational costs. Summary of the Invention
[0005] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a hierarchical representation method for pathological images based on a two-stream cross-attention network, which decouples the feature learning processes of different scales, separately performs feature learning on the image blocks of two resolutions using two branches, and then combines the features of the two branches to obtain a multi-scale hierarchical representation of pathological images, thereby improving the effectiveness and discriminability of hierarchical features of pathological images.
[0006] To achieve the above object of the invention, the hierarchical representation method for pathological images based on a two-stream cross-attention network of the present invention includes the following steps:
[0007] S1: Collect a number of pathological images and their corresponding labels, and determine two resolutions K l , K h in the image pyramid of the pathological images according to actual needs, where the side length magnification factor of the high-resolution K h relative to the low-resolution K l is λ. Obtain the images of these two resolutions from the image pyramid of the collected pathological images and perform non-overlapping block partitioning, ensuring that the blocks of the low-resolution image are regionally aligned with λ×λ high-resolution blocks in a one-to-one form; use the low-resolution image blocks and high-resolution image blocks of each pathological image as inputs, and their corresponding labels as outputs to form a training sample, thereby obtaining a training sample set;
[0008] S2: Construct a two-stream cross-attention network, including a feature extraction module, a convolutional layer, a low-resolution stream encoder, a feature reconstruction module, a square pooling module, a high-resolution stream encoder, a feature splicing module, a global attention pooling module, and a feature projection module, where:
[0009] The feature extraction module is used to extract features from the low-resolution image blocks and high-resolution image blocks of the pathological images respectively, and denote the obtained low-resolution features as and the high-resolution features as where m represents the number of blocks obtained by dividing the low-resolution image, d represents the dimension of the feature vector extracted from each block, and the low-resolution feature X l is sent to the convolutional layer and the square pooling module, and the high-resolution feature X h is sent to the feature reconstruction module;
[0010] The convolutional layer is used to perform convolution on the low-resolution feature X l to obtain the feature and send it to the low-resolution stream encoder, where d e represents the dimension of the feature vector of the low-resolution block obtained by the convolution operation;
[0011] The low-resolution stream encoder is used to encode the feature E l to obtain the feature and send it to the feature concatenation module;
[0012] The feature reconstruction module is used to reconstruct the high-resolution feature X h to obtain the high-resolution feature and send it to the square pooling module;
[0013] The square pooling module is used to perform square pooling on the high-resolution feature O h to obtain the pooled high-resolution feature and send it to the high-resolution stream encoder; The specific method of the square pooling operation is as follows:
[0014] E h = F(O h ; Q l )
[0015] where, represents the projected low-resolution feature, which is obtained by the following formula:
[0016] Q l = X l W l
[0017] represents the projection matrix to be trained;
[0018] F represents the cross-attention pooling function, and the expression is as follows:
[0019]
[0020] where, E h (i), O l (i) respectively represent the i-th row feature vectors in the feature E h , the feature Q l , i = 1, 2,..., m, Denote the high-resolution feature O h and the low-resolution feature O l (i) There is a sub-feature matrix composed of λ×λ eigenvectors with aligned block regions, W q 、W k 、W v respectively represent the projection matrices to be trained for query, key, and value;
[0021] The high-resolution stream encoder is used to encode the feature E h to obtain the feature and send it to the feature concatenation module;
[0022] The feature concatenation module is used to concatenate the feature V l and the feature V h to obtain the concatenated feature and send it to the global attention pooling module;
[0023] The global attention pooling module is used to perform global attention pooling on the concatenated feature S to obtain the feature and send it to the feature projection module;
[0024] The feature projection module is used to perform projection mapping on the feature H to obtain the hierarchical feature of the pathological image where D represents the preset hierarchical feature dimension;
[0025] S3: Determine a prediction model according to actual needs, then form a pathological image prediction network by combining the two-stream cross-attention network and the prediction model. Use the low-resolution image patches and high-resolution image patches of each pathological image in step S1 as inputs, and their corresponding labels as the expected outputs. Train the pathological image prediction network, and extract the trained two-stream cross-attention network from the trained pathological image prediction network;
[0026] S4: For the pathological image to be hierarchically characterized, extract the low-resolution K l and high-resolution K h images from its image pyramid, then perform non-overlapping block partitioning on the two-resolution images according to the method in step S1. Input the obtained low-resolution image patches and high-resolution image patches into the two-stream cross-attention network trained in step S3 to obtain the hierarchical feature of the pathological image.
[0027] The hierarchical representation method of pathological images based on the dual-stream cross-attention network of the present invention constructs a dual-stream cross-attention network. In this network, the low-resolution image patches and the high-resolution image patches are separately subjected to feature learning on two branches. In the feature learning of the high-resolution image patches, square pooling guidance is performed based on the features of the low-resolution image patches using the cross-attention mechanism. Then, the features learned on the two branches are combined to obtain a multi-scale hierarchical representation of the pathological images; a pathological image prediction network is constructed by combining the dual-stream cross-attention network and a prediction model, and it is trained using training samples. The trained dual-stream cross-attention network is extracted from the trained pathological image prediction network; for the pathological image to be hierarchically represented, its low-resolution image patches and high-resolution image patches are input into the trained dual-stream cross-attention network to obtain the hierarchical features of the pathological image.
[0028] The present invention decouples the feature learning processes at different scales, separately performs feature learning on image patches of two resolutions using two branches, and then combines the features of the two branches to obtain a multi-scale hierarchical representation of pathological images, thereby improving the effectiveness and distinctiveness of the hierarchical features of pathological images. Verification shows that the present invention has excellent performance in downstream prognosis tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 is a schematic diagram of the hierarchical representation method of pathological images based on sequences and heterogeneous graphs;
[0030] Figure 2 is a flowchart of the specific implementation of the hierarchical representation method of pathological images based on the dual-stream cross-attention network of the present invention;
[0031] Figure 3 is a structural diagram of the dual-stream cross-attention network in the present invention;
[0032] Figure 4 is a structural diagram of the feature reconstruction module in this embodiment;
[0033] Figure 5 is a processing flowchart of the dual-stream cross-attention network in this embodiment;
[0034] Figure 6 is a heat map of the attention regions of the hierarchical features of different cancer pathological images extracted using the present invention in this embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0035] The following describes the specific implementation of the present invention in conjunction with the accompanying drawings so that those skilled in the art can better understand the present invention. It should be particularly noted that in the following description, when the detailed description of known functions and designs may dilute the main content of the present invention, these descriptions will be omitted here.
[0036] Embodiment
[0037] Figure 2 is the flowchart of the specific implementation manner of the pathological image hierarchical representation method based on the dual-stream cross-attention network of the present invention. As Figure 2 shown, the specific steps of the pathological image hierarchical representation method based on the dual-stream cross-attention network of the present invention include:
[0038] S201: Obtain training samples:
[0039] Collect a number of pathological images and their corresponding labels, and determine two resolutions K l , K h in the image pyramid of the pathological images according to actual needs, where the magnification factor of the image side length of the high-resolution K h relative to the low-resolution K l is λ. Obtain the images of these two resolutions from the image pyramid of the collected pathological images and perform non-overlapping block partitioning, ensuring that the blocks of the low-resolution image are regionally aligned with λ×λ blocks of the high-resolution image in a one-to-one form. Take the low-resolution image blocks and high-resolution image blocks of each pathological image as inputs, and their corresponding labels as outputs to form a training sample, thereby obtaining a training sample set.
[0040] S202: Construct a dual-stream cross-attention network:
[0041] In order to be able to process the images of two resolutions in the image pyramid of pathological images and be able to properly solve the potential semantic gap in the process of fusing two-scale features, and extract effective and hierarchical global representations of pathological images, a dual-stream cross-attention network is constructed in the present invention. Figure 3 is the structural diagram of the dual-stream cross-attention network in the present invention. As[[ID=~]] Figure 3 shown, the dual-stream cross-attention network in the present invention includes a feature extraction module, a convolutional layer, a low-resolution stream encoder, a feature reconstruction module, a square pooling module, a high-resolution stream encoder, a feature splicing module, a global attention pooling module, and a feature projection module. Next, each module will be described in detail.
[0042] The feature extraction module is used to extract features from the low-resolution image blocks and high-resolution image blocks of pathological images respectively, and denote the obtained low-resolution features as and the high-resolution features as where m represents the number of blocks obtained by dividing the low-resolution image, d represents the dimension of the feature vector extracted from each block, and send the low-resolution feature X l to the convolutional layer and the square pooling module, and send the high-resolution feature X hSend to the feature reconstruction module. The feature extraction module can be selected according to actual needs. In this embodiment, the first three residual blocks in the ResNet50 model pre-trained on ImageNet are used as the feature extraction module.
[0043] The convolutional layer is used to perform convolution on the low-resolution feature X l to obtain the feature and send it to the low-resolution stream encoder, where d e represents the dimension of the feature vector of the low-resolution block obtained by the convolution operation. The convolution operation can be expressed by the following formula:
[0044] E l = Conv1D(X l ; k, s)
[0045] where, represents the one-dimensional convolution operation, k represents the size of the convolution kernel of the one-dimensional convolution, and s represents the convolution stride of the one-dimensional convolution. Through the convolution operation, a spatially overlapping block vector representation can be learned.
[0046] The low-resolution stream encoder is used to encode the feature E l to obtain the feature and send it to the feature concatenation module. The encoding operation can be expressed by the following formula:
[0047] V l = T h (E l )
[0048] where, represents the low-resolution stream encoder function.
[0049] Through the encoder, the dependence relationship between the learned blocks can be utilized to obtain a non-local feature representation of the blocks. The specific structure of the low-resolution stream encoder can be set according to actual needs. In this embodiment, the Transformer encoder is used as the low-resolution stream encoder.
[0050] The feature reconstruction module is used to reconstruct the high-resolution feature X h to obtain the high-resolution feature and send it to the square pooling module. Figure 4 is the structural diagram of the feature reconstruction module in this embodiment. As Figure 4 shown, the feature reconstruction module in this embodiment includes a spatial alignment module, a fully connected layer, and an activation function layer, where:
[0051] The spatial alignment module is used to align the high-resolution feature X h with the low-resolution feature X lPerform block alignment to obtain the aligned high-resolution features And send them to the fully connected layer.
[0052] The fully connected layer is used to reduce the dimension of the high-resolution feature f h to obtain the high-resolution feature And send them to the activation function layer.
[0053] The activation function layer is used to process the high-resolution feature g h according to a preset non-linear activation function to obtain the high-resolution feature And send them to the square pooling module.
[0054] In this embodiment, the feature reconstruction operation can be expressed by the following formula:
[0055] O h = σ(ρ(resize(X h )))
[0056] Where, represents the spatial alignment of the high-resolution block feature and the low-resolution block feature, that is, one low-resolution block feature corresponds to λ×λ high-resolution block features. represents the dimensionality reduction operation of the fully connected layer, and σ() represents the preset non-linear activation function.
[0057] The square pooling module is used to perform square pooling on the high-resolution feature O l according to the low-resolution feature X h to obtain the pooled high-resolution feature And send them to the high-resolution stream encoder. The square pooling can be expressed by the following formula:
[0058] E h = F(O h )
[0059] Where, represents the square pooling operation. Through the square pooling operation, it is possible to significantly reduce the number of features of the high-resolution blocks while learning the block features, so as to efficiently utilize the fine-grained image features to enhance the global representation of the pathological image.
[0060] Since two-scale features are used in the present invention, an inevitable problem is how to properly handle the fusion between different-scale features. However, the existing hierarchical representation methods of pathological images based on image pyramids do not fully consider this problem. The present invention proposes to implement the square pooling of high-resolution block features through a cross-attention mechanism, and this square pooling is guided by each low-resolution block aligned with the high-resolution block region. Specifically, the square pooling operation in the present invention is as follows:
[0061] E h = F(O h ; Q l )
[0062] Among them, represents the low-resolution feature after projection and is obtained by the following formula:
[0063] Q l = X l W l
[0064] represents the projection matrix to be trained, which is used to transform the feature dimension of the low-resolution block to the same as that of the high-resolution feature Q h .
[0065] F represents the cross-attention pooling function, and the expression is as follows:
[0066]
[0067] Among them, the superscript T represents the transpose, E h (i), O l (i) respectively represent the i-th row feature vectors of the features E h , the feature Q l , where i = 1, 2,..., m, represents the sub-feature matrix composed of λ×λ feature vectors in the high-resolution feature O h that are block-region aligned with the low-resolution feature O l (i), and W q , W k , W v respectively represent the projection matrices to be trained for query, key, and value.
[0068] Through the cross-attention pooling function, λ×λ high-resolution block features can be weighted and averaged under the guidance of a low-resolution block feature with a corresponding region alignment, so as to smoothly achieve the fusion of features at different scales.
[0069] The high-resolution stream encoder is used to encode the feature E h to obtain the feature and send it to the feature splicing module. Similar to the low-resolution stream encoder, the encoding operation of the high-resolution stream encoder can be expressed by the following formula:
[0070] V h = T h (E h )
[0071] Among them, Represents the high-resolution stream encoder function. In this embodiment, the high-resolution stream encoder uses the same Transformer encoder structure as the low-resolution stream encoder.
[0072] The feature splicing module is used to splice feature V l and feature V h to obtain the spliced feature and send it to the global attention pooling module.
[0073] The global attention pooling module is used to perform global attention pooling on the spliced feature S to obtain the feature and send it to the feature projection module.
[0074] The global attention pooling can be expressed by the following formula:
[0075] H = GAP(S)
[0076] where represents the global attention pooling function.
[0077] The feature projection module is used to perform projection mapping on the feature H to obtain the hierarchical features of the pathological image where D represents the preset hierarchical feature dimension. The feature projection can be expressed by the following formula:
[0078] G = HW map
[0079] where represents the projection matrix to be trained.
[0080] S203: Train the dual-stream cross-attention network:
[0081] Determine a prediction model according to actual needs, then form a pathological image prediction network with the dual-stream cross-attention network and the prediction model, use the low-resolution image patches and high-resolution image patches of each pathological image in step S201 as inputs, and their corresponding labels as the expected outputs, train the pathological image prediction network, and extract the trained dual-stream cross-attention network from the trained pathological image prediction network.
[0082] S204: Obtain the hierarchical representation of the pathological image:
[0083] For the pathological image to be hierarchically represented, extract the low-resolution K l and high-resolution K h from its image pyramid.The images are then non-overlappingly partitioned into blocks for the two-resolution images according to the method in step S201, and the obtained low-resolution image blocks and high-resolution image blocks are input into the two-stream cross-attention network trained in step S203 to obtain the hierarchical features of the pathological images.
[0084] To better illustrate the technical effects of the present invention, a specific example is used to experimentally verify the present invention. In this embodiment, the TCGA-BRCA dataset is used. First, training samples are obtained. First, low-resolution images (5 times) and high-resolution images (20 times) are selected from the WSI pyramid, and then the two-resolution images are respectively divided into a number of non-overlapping 256*256 pixel blocks. The low-resolution image obtains 256 blocks, and the high-resolution image obtains 1024 blocks.
[0085] Figure 5 is the processing flow chart of the two-stream cross-attention network in this embodiment. As Figure 5 shown, the magnification factor λ of the side length of the high-resolution image relative to the low-resolution image is 4. The feature extraction module uses the first three residual blocks in the ResNet50 model pre-trained on ImageNet to extract low-resolution features from the low-resolution blocks extract high-resolution features from the high-resolution blocks Both the low-resolution stream encoder and the high-resolution encoder use Transformer encoders, and both the feature dimensionality reduction module and the feature projection module use MLP (multi-layer perceptron).
[0086] Figure 6 is the attention area heat map of the hierarchical features of different cancer pathological images extracted by the present invention in this embodiment. Figure 6 The area delimited by the black line in the pathological image in Figure 6 is the region of interest marked by the pathologist. As
[0087] shown, the regions concerned by the hierarchical feature map extracted by the present invention are almost all within the regions of interest marked by the doctor. This shows that the two-stream cross-attention network proposed by the present invention can automatically extract important clues from pathological images.
[0088] Deep Sets, see the literature "Zaheer, M., et al. "Deep Sets." (2017).";
[0089] Attention-MIL, see the literature "Ilse, M., J. M. Tomczak, and M. Welling. "Attention-based Deep Multiple Instance Learning." (2018).";
[0090] Deep Attn MISL, see the literature "Yao, J., et al. "Whole Slide Images based CancerSurvival Prediction using Attention Guided Deep Multiple Instance LearningNetworks." Medical Image Analysis 65(2020).";
[0091] Deep Graph Conv, see the literature "Li, R., et al. “Graph CNN for Survival Analysison Whole Slide Pathological Images.” (2018).";
[0092] Patch-GCN, see the literature "Chen, R. J., et al. “Whole Slide Images Are 2D PointClouds:Context-aware Survival Prediction using Patch-based GraphConvolutional Networks.” (2021).";
[0093] Se TranSurv, see the literature "Huang, Z., et al. “Integration of Patch Featuresthrough Self-supervised Learning and Transformer for Survival Analysis onWhole Slide Images.” (2021).";
[0094] TransMIL, see the literature "Shao, Z., et al. “TransMIL: Transformer based Correlated Multiple Instance Learning for Whole Slide Image Classification.” (2021).”;
[0095] H2MIL, see the literature "Hou, W., et al. “H2MIL: Exploring Hierarchical Representation with Heterogeneous Multiple Instance Learning for Whole Slide Image Analysis.” (2022).”;
[0096] Table 1 is a comparison table of the C-Index evaluation index and standard deviation of the present invention and eight comparison methods in three data sets in this embodiment.
[0097]
[0098] Table 1
[0099] As shown in Table 1, the C-Index evaluation indexes of the present invention in the three pathological image data sets all exceed the existing eight comparison methods. It can be seen that the hierarchical representation method of the present invention is effective in the downstream pathological image prognosis task. At the same time, the standard deviations of the present invention in the three pathological image data sets also remain at a low level, which can prove the stability of the present invention.
[0100] Although the above describes the illustrative specific embodiments of the present invention for the convenience of those skilled in the art to understand the present invention, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.
Claims
1. A hierarchical representation method for pathological images based on a dual-stream cross-attention network, characterized in that Including the following steps: S1: Collect a number of pathological images and their corresponding labels, and determine two resolutions K l , K h in the image pyramid of the pathological images according to actual needs, where the high-resolution K h has an image side length magnification factor of λ relative to the low-resolution K l . Obtain the images of these two resolutions from the image pyramid of the collected pathological images and perform non-overlapping block partitioning, ensuring that the blocks of the low-resolution image are regionally aligned with λ×λ high-resolution blocks in a one-to-one form; Taking the low-resolution image patches and high-resolution image patches of each pathological image as inputs, and their corresponding labels as outputs, to form a training sample, thereby obtaining a training sample set; S2: Constructing a dual-stream cross-attention network, including a feature extraction module, a convolutional layer, a low-resolution stream encoder, a feature reconstruction module, a square pooling module, a high-resolution stream encoder, a feature splicing module, a global attention pooling module, and a feature projection module, where: The feature extraction module is used to extract features from the low-resolution image patches and high-resolution image patches of the pathological image respectively. Denote the obtained low-resolution features as and the high-resolution features as where m represents the number of patches obtained by dividing the low-resolution image, d represents the dimension of the feature vector extracted from each patch, and send the low-resolution feature X l to the convolutional layer and the square pooling module, and send the high-resolution feature X h to the feature reconstruction module; The convolutional layer is used to perform convolution on the low-resolution feature X l to obtain a feature and send it to the low-resolution stream encoder, where d e represents the dimension of the feature vector of the low-resolution block obtained by the convolution operation; The low-resolution stream encoder is used to encode feature E l to obtain feature and send it to the feature splicing module; The feature reconstruction module is used to reconstruct the high-resolution feature X h to obtain the high-resolution feature and send it to the square pooling module; The square pooling module is used to perform square pooling on the high-resolution feature O h to obtain the pooled high-resolution feature and send it to the high-resolution stream encoder; the specific method of the square pooling operation is as follows: E h = F(O h ; Q l ) Among them, represents the low-resolution feature after projection and is obtained by the following formula: Q l = X l W l Denote the projection matrix to be trained; F represents the cross-attention pooling function, and the expression is as follows: Among them, E h (i), O l (i) respectively represent the eigenvector of the i-th row in the feature E h , feature Q l , where i = 1, 2,..., m, represents the high-resolution feature O h in which there is a sub-feature matrix composed of λ × λ eigenvectors that are block-region aligned with the low-resolution feature O l (i), and W q , W k , W v respectively represent the projection matrices to be trained for query, key, and value; The high-resolution stream encoder is used to encode feature E h to obtain a feature and send it to the feature splicing module; The feature splicing module is used to splice feature V l and feature V h to obtain the spliced feature and send it to the global attention pooling module; The global attention pooling module is used to perform global attention pooling on the concatenated feature S to obtain a feature and send it to the feature projection module; The feature projection module is used to perform projection mapping on the feature H to obtain the hierarchical features of the pathological image where D represents the preset dimension of the hierarchical feature; S3: Determining a prediction model according to actual needs, then combining the dual-stream cross-attention network and the prediction model to form a pathological image prediction network, taking the low-resolution image patches and high-resolution image patches of each pathological image in step S1 as inputs, and their corresponding labels as the expected outputs, training the pathological image prediction network, and extracting the trained dual-stream cross-attention network from the trained pathological image prediction network; S4: For the pathological image to be hierarchically represented, extract the low-resolution K l and high-resolution K h images from its image pyramid, and then perform non-overlapping block partitioning on the images of the two resolutions according to the method in step S1. Input the obtained low-resolution image blocks and high-resolution image blocks into the two-stream cross-attention network trained in step S203 to obtain the hierarchical features of the pathological image.
2. The hierarchical representation learning method for pathological images according to claim 1, wherein The feature reconstruction module includes a spatial alignment module, a fully connected layer, and an activation function layer, where: The spatial alignment module is used to perform block alignment on the high-resolution feature X h and the low-resolution feature X l according to the image block division method, so as to obtain the aligned high-resolution feature and send it to the fully-connected layer; The fully connected layer is used to reduce the dimension of the high-resolution feature f h to obtain the high-resolution feature and send it to the activation function layer; The activation function layer is used to process the high-resolution feature g according to a preset non-linear activation function h to obtain a high-resolution feature and send it to the square pooling module.
Citation Information
Patent Citations
Eye movement attention feature vector determination method for children ADHD screening evaluation system
CN111627526A
Convolutional neural network based on four-branch attention mechanism, and image segmentation method
CN112949838A