Attention-based learning for digital histopathology analysis
By pre-training deep learning networks through hierarchical encoding and self-supervised learning, the problem of large data volumes in digital histopathology images was solved, and efficient image analysis and tumor subtyping prediction were achieved.
Patent Information
- Application Number
- CN202480014075.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-02-23
- Filing Date
- 2024-02-22
- Publication Date
- 2025-10-10
AI Technical Summary
The large amount of pixel data in digital histopathology images leads to low efficiency of computerized analysis and difficulty in performing effective regression analysis and segmentation.
We use layered encoding and self-supervised learning (SSL) objectives to pre-train a deep learning network, use methods such as DINOv2 and simCLR to train patch-level and region-level encoders, and combine them with an attention-based encoder for image processing.
It significantly reduces storage requirements and improves the efficiency and accuracy of image analysis, especially showing high performance in tumor subtyping prediction tasks.
Smart Images

Figure CN120770044A_ABST
Abstract
Description
[0001] Cross Reference to Related Applications
[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 486,625, filed February 23, 2023. The contents of this application are incorporated herein by reference. BACKGROUND
[0003] The present disclosure relates generally to techniques for computerized inference from patient medical images. SUMMARY
[0004] The amount of pixel data in a typical whole slide digital histopathology image (WSI) is significant. This presents challenges for efficient and practical computerized analysis for inference from the WSI, regression analysis performed on the WSI, and segmentation of the WSI.
[0005] Embodiments of the present disclosure utilize hierarchical encoding and self-supervised learning (SSL) objectives to provide pre-training portions of deep learning networks for digital histopathology analysis. Some embodiments of the present disclosure utilize different training methods to train different portions of the network. Some embodiments use the DINOv2 objective to pre-train a patch-level encoder and / or to pre-train a region-level encoder. Some embodiments use a contrastive learning objective, such as simCLR, to pre-train a patch-level encoder, but use a different learning objective to pre-train a region-level encoder. These and other aspects of the present disclosure are described in greater detail below. BRIEF DESCRIPTION OF DRAWINGS
[0006] Figure 1 A deep learning network for processing digital histopathology images according to embodiments of the present disclosure is illustrated.
[0007] Figure 2 is a flowchart illustrating a method for training a digital histopathology deep learning network, such as, for example, the deep learning network of Figure 1 According to embodiments of the present disclosure.
[0008] Figure 3 Alternative learning methods for each stage of training a hierarchical deep learning network, such as, for example, the deep learning network of Figure 1 are illustrated.
[0009] Figure 4 is a data table illustrating results of testing various embodiments of the present disclosure.
[0010] Figure 5 An example of a computer system in which one or more of the devices, systems, and methods illustrated herein can be implemented is shown.
[0011] While the embodiments are described with reference to the above drawings, the drawings are intended to be illustrative, and other embodiments are consistent with the spirit of the disclosure and within the scope of the disclosure. DETAILED DESCRIPTION
[0012] Various embodiments will now be described more fully with reference to the accompanying drawings in which some embodiments are illustrated. The various embodiments are not intended to be limiting. The embodiments can be used in combination with each other as well as in other combinations than those explicitly provided. Other embodiments are provided which are within the scope of the embodiments. Various embodiments can be implemented in hardware, software or a combination thereof. The various embodiments can be implemented in one or more computing devices or other processing systems.
[0013] Figure 1 A deep learning network 1000 for processing digital histopathology images, such as, for example, whole slide images 11, is illustrated. In digital histopathology, whole slide images of hematoxylin and eosin (H&E) stained tissue are often referred to simply as “whole slide images” or “WSIs”.
[0014] In the illustrated example, the system is assumed to be fully trained and fine-tuned for the end use of generating predictions or otherwise making inferences related to a particular task (e.g., the presence of a particular type of cancer based on a histological sample). In this example, the patch-level and region-level encoders have been pre-trained for a more general task using self-supervised learning (SSL) techniques, also referred to in the art as “self-supervision.” An example of hierarchical pre-training of a Vision Transformer (ViT) encoder for WSI is described in Chen et al., “Scaling Vision Transformers to Gigapixel Images via Hierarchical Self-Supervised Learning” (2022) (arXiv:2206.02647). This paper is incorporated by reference in its entirety. The whole slide level encoder has been trained with labeled data with a classification network to fine-tune for the particular prediction task.
[0015] The pre-processing block 110 receives a WSI 11 (e.g., digital data corresponding to an image of H&E stained tissue on a histopathology slide, the tissue corresponding to a sample taken from a patient) and performs pre-processing, which includes dividing the image 11 into regions and patches, where each region includes multiple patches. In one non-limiting example, a WSI is divided into several regions, each region having a size of 4480 x 4480 pixels. In this example, each region is divided into 400 patches, each patch having a size of 224 x 224 pixels. The number of regions in a WSI will depend on the size of the WSI and can be expected to vary across a set of WSIs. It can be 4 regions, 8 regions, 10 regions, 16 regions, or some other number of regions.
[0016] WSIs are known to have various types of artifacts, ranging from large background to handwriting, blurry regions, and air bubbles. The attention mechanism in the visual transformer (ViT) network can automatically focus on salient regions and thus does not require very fine-grained quality control (QC) output. In one embodiment, background and handwriting are detected from the WSI.
[0017] In one embodiment, the pre-processing block 110 uses HistoQC (https: / / github.com / choosehappy / HistoQC / wiki), which is an open-source quality control tool for digital histopathology, to extract tissue regions of size 4480 x 4480 from the WSI. Each region is processed by the tool and it sequentially produces metrics and tracks the location of background and handwriting in the full slide image. Based on these metrics, in one example, the pre-processing block 110 filters out patches containing artifacts and selects regions in the image that have rich features.
[0018] The pre-processing block 110 outputs patch-level digital images 12. In one example, the patch-level images 12 have a size of 224 x 224 pixels. The first pre-trained encoder 120 processes each patch image 12 to generate a patch-level representation 21. In one example, the patch-level representation 21 is a vector having 768 dimensions. Thus, in this example, the resulting representation data for each patch-level representation 21 has a size of 1 x 768, which consumes significantly less memory than the 224 x 224 pixel data of the patch 12. Of course, those skilled in the art will appreciate that the dimension size will depend on the type of encoder used and its configuration for a particular application. Thus, the dimensions given for this embodiment are merely by way of example.
[0019] The patch-level representations 21 are used to represent each patch in a region, providing a region-level pseudo-image 20. The image 20 is referred to as a “pseudo” image (or representative image) because the patches of the region image are not pixel space representations, but rather representations in a space whose dimensions are defined by the first patch-level encoder 120. In one example, each region includes 40 patches, and thus the size of the region-level pseudo-image 20 is 40 x 768.
[0020] The region-level pseudo-image 20 is processed by a second pre-trained encoder 130. The encoder 130 processes each region-level pseudo-image 20 and generates a region-level representation 31. In one example, the region-level representation 31 is a vector having 384 dimensions. Thus, in this example, the size of the resulting representation data for each region-level representation 31 is 1 x 384, which consumes significantly less memory than the 4480 x 4480 pixel data from the corresponding region of the WSI 11. Of course, one skilled in the art will appreciate that the dimension size will depend on the type of region-level encoder used and its configuration for a particular application. Thus, the dimensions given for this embodiment are merely examples.
[0021] The region-level representations 31 are used to represent each region of the whole slide image (one representation 31 per region), providing a whole slide-level pseudo 30 image comprising the region-level representations 31. The number of regions can vary depending on the slide. But in one example, the WSI is divided into approximately 4 regions, so the size of the whole slide pseudo-image 30 can be 4 x 384.
[0022] The region-level representations 31 corresponding to the whole slide pseudo-image 30 are processed by an attention-based encoder 140. In one embodiment, the attention-based encoder 140 is a Vision Transformer (ViT) encoder. In one example, the ViT has 4 attention heads, 2 layers (i.e., depth = 2) and an output dimension of 384. The attention-based encoder 140 generates a whole slide-level representation 41 (which, in the context of a ViT encoder, is the processed CLS token). The representation 41 is then processed by a classifier network 150, which in one example can include one or more typical feed-forward hidden layers (i.e., multi-layer perceptron or “MLP”) followed by a softmax layer to obtain a prediction (e.g., class probabilities) of the classification of the tissue corresponding to the relevant WSI 11 processed by the system 1000.
[0023] Figure 2 is a flowchart illustrating a method 2000 for training a digital histopathology deep learning network, such as, for example, the network 1000 of Figure 1
[0024] Step 201 performs self-supervised learning on the patch-level training images to obtain a pre-trained patch-level encoder. Preferably, a large training set of histopathology images is used at step 201. Those images are preferably not task-specific, but rather include a wide variety of tissue samples, e.g., from many different potential disease sites. In one example, over 11,677 H&E stained whole slide images from the TCGA public data are used for pre-training. In one example, the training slides are from 33 different potential tumor sites. In one example, approximately 33 million 224x224 pixel patches are extracted from the WSI training set and used at step 201.
[0025] After obtaining the pre-trained patch-level encoder from step 201, step 202 encodes the training image patches (pixel data) into patch-level representations using the pre-trained patch-level encoder. In one example, each patch-level representation is a vector representation including 768 dimensions. In one example, the training image patches encoded at step 202 are the same training patches used in step 201. In another example, the training image patches encoded at step 202 are different image patches from a different training set than the training set used for pre-training at step 201.
[0026] Step 203 uses the patch-level representations to form region-level pseudo images, each corresponding to a region of a WSI, the pseudo images including a plurality of patch representations for patches in the corresponding region. In one example, the region-level pseudo images include 400 patch-level representations.
[0027] Step 204 performs self-supervised learning using the region-level pseudo images to pre-train a region-level encoder.
[0028] Once the patch-level and region-level encoders are pre-trained, they are available for use in hierarchically generating whole slide pseudo images from a labeled training set. These whole slide pseudo images, which include region-level representations, can then be used to train a whole slide level encoder along with a classifier for a particular prediction task corresponding to the labeled training set. This downstream “fine-tuning” of the encoders and classifiers for a particular task using supervised learning is further described below in the context of steps 205-209.
[0029] Step 205 generates patch-level representations from patches of WSIs from the labeled training set using the pre-trained patch-level encoder. Step 206 then incorporates the patch-level representations into region-level pseudo images including patch representations.
[0030] Step 207 generates a region-level representation from the region-level pseudo image comprising patch-level representations using the pre-trained region-level encoder. Step 208 incorporates the region-level representation into a whole slide-level pseudo image. Each whole slide-level pseudo image corresponds to a WSI in the labeled training set.
[0031] However, because the whole slide-level pseudo image is generated by hierarchical encoding, first at the patch level (patch-level representations based on patch-level pixel data), and then at the region level (region-level representations generated using patch-level representations instead of patch-level pixel data), the whole slide-level pseudo image contains a small fraction of the data (in representation space) compared to the data contained in the pixel space data volume of the corresponding whole slide image. For example, if each region comprises data from 400 patches, and each patch is 224 x 224 pixels, and for example, the WSI has 4 regions, the entire WSI can contain a data volume of 80,281,600 pixels (or multiplied by 3 in RGB pixel space). In contrast, the whole slide-level pseudo image can contain 4 x 384 or 1,536 values (a small fraction of the pixel space version of the image) to represent the entire slide (assuming 4 regions). Figure 1 In embodiments of the
[0032] Continuing with the description of Figure 2 Step 209 uses the whole slide-level pseudo image for supervised learning to train a whole slide-level attention-based encoder to effectively encode whole slide-level representations of the WSIs in the training set for a particular prediction task. In one embodiment, as described above in the context of Figure 1 In one embodiment, as described above in the context of Figure 3 The whole slide-level encoder generates whole slide-level representations from the whole slide-level pseudo image during supervised training. The whole slide-level representations can then be submitted to a classifier network (e.g., a multilayer perceptron) to generate predictions. During training, the computed loss is backpropagated through both the one or more classifier layers and the whole slide-level encoder to adjust the learnable parameters. In other words, the whole slide-level encoder and the classifier are trained together using supervised learning.
[0033] Figure 3 Various alternatives for each stage of the training process are illustrated. Alternative 301 illustrates a pre-training process for the patch-level encoder (such as, for example, the patch-level encoder described above in the context of Figure 1alternative self-supervised learning (SSL) method for the first SSL pre-trained encoder 120) of the illustrated embodiment. In some embodiments, the “SimCLR” method can be used, as described in Chen, et al., “A Simple Framework for Contrastive Learning of Visual Representation” (2020) (arXiv:2002.05709). In some embodiments, the ViT Masked Autoencoder (ViTMAE) method can be used, as described in He, K., et al., “Masked Autoencoder Are Scalable Vision Learners” (2021) (arXiv:2111.06377). In other embodiments, the “DINOv1” method can be used, as described in Caron, et al., “Emerging Properties in Self-Supervised Vision Transformers” (2021) (arXiv:2104.14294). In some embodiments, the “DINOv2” method can be used, as described in Oquab, et al., “DINOv2: Learning Robust Visual Features without Supervision” (2023) (arXiv:2304.07193). All four papers are hereby incorporated by reference in their entireties.
[0034] Alternative 302 illustrates an alternative SSL method for the pre-trained region-level encoder (such as, for example, the pre-trained region-level encoder 120) of the illustrated embodiment. Figure 1 Alternative 302 illustrates an alternative SSL method for the pre-trained region-level encoder (such as, for example, the pre-trained region-level encoder 120) of the illustrated embodiment.
[0035] Alternative 303 illustrates an alternative SSL method for the full-slide-level encoder (such as, for example, the full-slide-level encoder 140) of the illustrated embodiment along with the classifier (such as, for example, the classifier 150) of the illustrated embodiment. Figure 1 Alternative 303 illustrates an alternative SSL method for the full-slide-level encoder (such as, for example, the full-slide-level encoder 140) of the illustrated embodiment along with the classifier (such as, for example, the classifier 150) of the illustrated embodiment. Figure 1alternative to the DINOv2 objective for pre-training the ViT for both the patch-level encoder (e.g., 120) and the region-level encoder (e.g., 130) of the classifier 150) of some embodiments. In some embodiments, the full-slide-level encoder is trained with a supervised learning along with a classifier. As shown, one embodiment uses a ViT encoder. When using a ViT encoder, the region-level representations for the full-slide-level pseudo image are position encoded and provided to the multi-headed attention layers of the ViT encoder. An additional “classification” token (CLS) is processed along with the region-level representations, and because the CLS token participates in every region-level representation it processes, the output value of this token provides an effective “aggregated” representation that can be used downstream as input to a classifier to classify the entire slide. The learnable parameters of the ViT can be fine-tuned along with the learnable parameters of the classifier (e.g., MLP) via supervised learning.
[0036] In some embodiments, during the fine-tuning phase, a convolutional neural network (CNN) such as a residual CNN (ResNet) is used to further encode each region representation of the full-slide pseudo image, then an attention mechanism assigns attention weights to each region encoding, which can then be used to aggregate the encodings to obtain an attention-weighted representation of the entire slide. This approach is described, for example, in Ilse et al., “Attention-based Deep Multiple Instance Learning” (2018) (arXiv:1802.04712v4). The learnable parameters of the CNN and the attention mechanism along with the learnable parameters of the classifier (e.g., MLP) can be fine-tuned via supervised learning.
[0037] In one specific embodiment, the DINOv2 objective is used to pre-train a ViT for both the patch-level encoder (e.g., 120) and the region-level encoder (e.g., 130) of some embodiments. Figure 1 In one specific embodiment, the DINOv2 objective is used to pre-train a ViT for both the patch-level encoder (e.g., 120) and the region-level encoder (e.g., 130) of some embodiments. Figure 1 In some embodiments, the simCLR objective (contrastive learning) is used to pre-train a ResNet encoder as a patch-level encoder, and the DINOv2 objective is used to pre-train a ViT as a region-level encoder. In some embodiments, the ViTMAE objective is used to pre-train a region-level encoder.
[0038] In some embodiments, any combination of the listed alternatives can be selected for each phase. That is, any one or more of the 301 pre-training objective alternatives can be selected for pre-training the patch-level encoder, and any one or more of the 302 pre-training objective alternatives can be selected for pre-training the region-level encoder.
[0039] Experimental Results
[0040] Various implementations have been tested for downstream fine-tuning on three selected tasks: (1) NSCLC (non-small cell lung cancer) subtyping: 869 H&E images from LUAD (lung adenocarcinoma) and LUSC (lung squamous cell carcinoma); (2) RCC subtyping (renal cell carcinoma); 810 images from KIRP (renal papillary cell carcinoma), KICH (kidney pigmentation), KIRC (renal clear cell carcinoma), and (3) BRCA (breast cancer) subtyping: 883 H&E images.
[0041] For each of these tasks, the experiments divided the images into an 80 / 20 training / test split. The 80% split was then used for 10-fold cross validation, and performance on the test split was reported. Note that the 20% test split was not included in the pre-training of the patch-level and region-level models. For these experiments, the patch-level and region-level encoders were not fine-tuned (modified) during the downstream training of the full-slide-level encoder and classifier. All models used were trained on the training set using 10-fold cross validation. The model of each fold was selected on the epoch with the lowest validation loss. Performance was reported on the test set.
[0042] Figure 4 is a table summarizing the example results using the area under the curve (AUC) metric. Figure 3 Results of downstream prediction following possible combinations of pre-training methods for the patch-level encoder alternative 301 and the region-level encoder alternative 302 are shown.
[0043] refer to Figure 4 DINOv2 pre-training yields good performance at both the patch and region levels. Furthermore, performance is high for region-level pre-training methods using the patch-level DINOv2 model. Similarly, at the region level, DINOv2 can extract meaningful representations from models trained with lower-level patch-level performance (such as ViTMAE and DINOv1) and still provide high-level performance on prediction tasks. Furthermore, combinations of region-level pre-training using simCLR at the patch level and using any of the other objectives achieve high performance.
[0044] Unlike natural images, H&E images appear very uniform in pixel space and they have subtle variations. Therefore, in some embodiments, the mean squared error (MSE) loss applied to the images may not capture the granular nuances of the tumor microenvironment in these images. At the region level, the ViTMAE region-level model performs very well for models with high-level patch-level performance, such as SimCLR and DINOv2. Therefore, the mask modeling type pre-training objective can perform well in scenarios where the data has a high level of intra-image variation.
[0045] Figure 5 An example of a computer system 5000 is shown, one or more of which can be used to implement one or more of the devices, systems, and methods exemplified herein. The computer system 5000 executes instruction code contained in a computer program product 560. The computer program product 560 includes executable code in an electronically readable medium that can instruct one or more computers, such as the computer system 5000, to perform the processing that implements the exemplary method steps performed.
[0046] The electronically readable medium can be any transitory or non-transitory medium that electronically stores information and can be accessed locally or remotely via a network connection. The medium can include multiple geographically dispersed media each configured to store different portions of the executable code at different locations and / or at different times. The executable instruction code in the electronically readable medium directs the exemplified computer system 5000 to perform the various exemplary tasks described herein. The executable code for directing the performance of the tasks described herein will typically be implemented in software. However, those skilled in the art will appreciate that the computer or other electronic device can utilize code implemented in hardware to perform many or all of the identified tasks. Those skilled in the art will appreciate that many variations of the executable code implementing the exemplary methods can be found within the spirit and scope of the present disclosure.
[0047] The code or copies of the code contained in the computer program product 560 can reside in one or more storage persistent media (not shown separately) communicatively coupled to the system 5000 for loading into and storing to persistent storage 570 and / or memory 510 for execution by the processor 520. The computer system 500 also includes an I / O subsystem 530 and peripheral devices 540. The I / O subsystem 530, peripheral devices 540, processor 520, memory 510, and persistent storage 570 are coupled via a bus 550. Like the persistent storage 570 and any other persistent storage devices that can contain the computer program product 560, the memory 510 is a non-transitory medium (even if implemented as a typical volatile computer memory device). Moreover, those skilled in the art will appreciate that the memory 510 and / or the persistent storage 570 can be configured to store, in addition to the computer program product 560 for performing the processes described herein, the various data elements referenced and exemplified herein.
[0048] Those skilled in the art will appreciate that the computer system 5000 is merely one example of a system that can implement a computer program product according to the present disclosure. With reference to only one example, the execution of the instructions contained in the computer program product can be distributed across multiple computers, such as, for example, across computers of a distributed computing network.
[0049] Instructions for implementing an artificial neural network or other deep learning network can reside in the computer program product 560. When the processor 520 is executing the instructions of the computer program product 560, the instructions or a portion thereof are typically loaded into a working memory 510 that is easily accessible to the processor 520 from which the instructions can be accessed.
[0050] The processor 520 can include multiple processors, which can include respective additional working memories (additional processors and memories not shown separately) including one or more graphics processing units (GPUs) including at least thousands of arithmetic logic units that support massive parallel computing. GPUs are often used for deep learning applications because they can perform related processing tasks more efficiently than typical general-purpose processors (CPUs). The processor 520 can additionally or alternatively include one or more specialized processing units including systolic arrays and / or other hardware arrangements that support efficient parallel processing. Such specialized hardware can work in conjunction with CPUs and / or GPUs to perform the various processing described herein. Such specialized hardware can include application-specific integrated circuits or the like (which can refer to a portion of an application-specific integrated circuit), field-programmable gate arrays or the like, or combinations thereof. However, a processor such as the processor 520 can be implemented as one or more general-purpose processors, preferably with multiple cores, without necessarily departing from the spirit and scope of the present disclosure.
[0051] While the present disclosure has been described with respect to the illustrated embodiments, it will be understood that various changes, modifications and adaptations can be made based on the disclosure, and it is intended to cover such changes, modifications and adaptations in the scope of the disclosure. While the present disclosure has been described with respect to the presently preferred and alternative embodiments, it is to be understood that the disclosure is not limited to the disclosed embodiments, but is instead intended to cover any and all alternatives, modifications and equivalents within the scope of the underlying principles as described herein and in the appended claims.
[0052] Additional Embodiments
[0053] Some examples include a method of obtaining a pre-trained portion of a deep learning network configured to execute on one or more computers to process a histopathology image to extract a representation usable to analyze tissue corresponding to the histopathology image, the method comprising: obtaining a first pre-trained encoder (an “encoder” can also be referred to as a feature extraction network because it extracts a representation of input data), pre-training the first encoder using the first encoder to pre-train the first encoder by processing a plurality of patches obtained from a first plurality of histopathology images using a first self-supervised learning (SSL) objective; using the first pre-trained encoder to obtain a plurality of pseudo-images, each pseudo-image corresponding to an image region and comprising a plurality of patch representations obtained by processing a plurality of image patches in the image region, the plurality of pseudo-images corresponding to image regions from a second plurality of histopathology images; and obtaining a second pre-trained encoder using a second encoder by pre-training the second encoder by processing the plurality of pseudo-images using a second self-supervised learning objective; wherein the pre-trained portion of the deep learning network comprises the first pre-trained encoder and the second pre-trained encoder.
[0054] In some examples, the first SSL learning objective and the second SSL learning objective comprise DINOv2, and the first encoder and the second encoder comprise a visual transformer (ViT) encoder. In some examples, the first SSL learning objective comprises simCLR. In some examples, the first SSL learning objective comprises simCLR, and the first encoder comprises a convolutional neural network (CNN). In some examples, the CNN is a ResNet.
[0055] In some examples, the first plurality of histopathology images and the second plurality of histopathology images are the same. In some examples, the first plurality of histopathology images and the second plurality of histopathology images of histopathology images are the same.
[0056] In some examples, the image region corresponds to a portion of the image that is at least 100 times larger than the image patch. In some examples, the image region corresponds to a portion of the image that is at least 200 times larger than the image patch. In some examples, the image region corresponds to a portion of the image that is at least 400 times larger than the image patch.
[0057] Some embodiments include a computer-executable deep learning network stored in a non-transitory computer-readable medium and configured to execute on one or more computers to process a histopathology image and analyze tissue corresponding to the histopathology image, the computer-executable deep learning network comprising: a first pre-trained encoder; a second pre-trained encoder; and a machine learning analyzer configured to receive, from the second pre-trained encoder, a representation of the histopathology image and generate an analysis result of the tissue corresponding to the histopathology image, wherein the analysis result comprises at least one of a classification, a regression, and a segmentation.
[0058] In some examples, the machine learning analyzer includes one or more feedforward neural network layers. In some examples, the machine learning analyzer includes a visual transformer encoder. In some examples, the machine learning analyzer includes a non-neural network-based analyzer. In some examples, the machine learning analyzer includes a logistic regression-based analyzer. In some examples, the machine learning analyzer includes a machine learning classifier and the analysis result includes a classification expressed as a class probability.
Claims
1. A method for obtaining a pre-trained portion of a deep learning network configured to execute on one or more computers to process a histopathology image to extract a representation that can be used to analyze tissue corresponding to the histopathology image, the method comprising: obtaining a first pre-trained encoder by processing a plurality of patches obtained from a first plurality of histopathology images using a first self-supervised learning (SSL) objective to pre-train the first encoder; obtaining a plurality of pseudo images using the first pre-trained encoder, each pseudo image corresponding to an image region and comprising a plurality of patch representations obtained by processing a plurality of image patches in the image region, the plurality of pseudo images corresponding to image regions from a second plurality of histopathology images; as well as Using a second encoder, obtaining a second pre-trained encoder by processing the plurality of pseudo images using a second self-supervised learning objective to pre-train the second encoder; The pre-trained portion of the deep learning network includes the first pre-trained encoder and the second pre-trained encoder.
2. The method according to claim 1, wherein The first SSL learning objective comprises DINOv2, and the first encoder comprises a visual transformer (ViT) encoder.
3. The method according to claim 1, wherein The first SSL learning objective comprises DINOv1, and the first encoder comprises a visual transformer (ViT) encoder.
4. The method according to claim 1, wherein The first SSL learning objective includes a comparative learning objective.
5. The method according to claim 4, wherein The comparative learning targets include simCLR.
6. The method according to any one of claims 4 to 5, wherein The first encoder includes a convolutional neural network.
7. The method according to any one of claims 2 to 6, wherein The second SSL learning objective includes DINOv2.
8. The method according to any one of claims 2 to 6, wherein The second SSL learning objective includes DINOv1.
9. The method according to any one of claims 2 to 6, wherein: The second SSL learning objective includes Visual Transformer Mask Auto-Encoding (ViTMAE).
10. The method according to any one of claims 1 to 9, wherein The first plurality of histopathological images is identical to the second plurality of histopathological images.
11. The method according to any one of claims 1 to 9, wherein: The first plurality of histopathological images is different from the second plurality of histopathological images.
12. The method according to any one of claims 1 to 9, wherein The first plurality of histopathology images overlap with the second plurality of histopathology images.
13. The method according to any one of claims 1 to 12, wherein An image region corresponds to a portion of the image that is at least 100 times larger than the image patch.
14. The method according to any one of claims 1 to 12, wherein An image region corresponds to a portion of the image that is at least 200 times larger than the image patch.
15. The method according to any one of claims 1 to 12, wherein The image region corresponds to a portion of the image that is at least 400 times larger than the image patch.
16. A computer-executable deep learning network stored in a non-transitory computer-readable medium and configured to execute on one or more processors of one or more computers to process a histopathology image and analyze tissue corresponding to the histopathology image, the computer-executable deep learning network comprising: The first pre-trained encoder obtained by the method according to any one of claims 1 to 15; The second pre-trained encoder obtained by the method according to any one of claims 1 to 15; and A machine learning analyzer is configured to receive the representation of the histopathology image from the second pre-trained encoder and generate an analysis result of the tissue corresponding to the histopathology image, wherein the analysis result comprises at least one of classification, regression, and segmentation.
17. The computer-executable deep learning network of claim 16, wherein: The machine learning analyzer includes one or more feed-forward neural network layers.
18. The computer-executable deep learning network of claim 16, wherein: The machine learning analyzer includes a visual transformer (ViT) encoder.
19. The computer-executable deep learning network of claim 16, wherein: The machine learning analyzer includes an analyzer based on logistic regression.
20. The computer-executable deep learning network of claim 16, wherein: The machine learning analyzer includes a machine learning classifier, and the analysis results include classifications expressed as class probabilities.
21. A computer-executable deep learning network stored in a non-transitory computer-readable medium and configured to execute on one or more processors of one or more computers to process a whole-slide histopathology image (WSI) and analyze tissue corresponding to the WSI, the computer-executable deep learning network comprising: a first pre-trained encoder, wherein the first pre-trained encoder has been pre-trained using a first self-supervised learning objective to extract patch-level representations from patch-level pixel data corresponding to patches of the WSI; a second pre-trained encoder, wherein the second pre-trained encoder has been pre-trained using a second self-supervised learning objective to extract region-level representations from a set of patch-level representations, the set of patch-level representations corresponding to regions of the WSI; and an attention-based machine learning analyzer configured to receive the region-level representation of the WSI from the second pre-trained encoder, generate a WSI-level representation from the region-level representation of the WSI, and process the WSI-level representation to generate an analysis result of the tissue corresponding to the histopathology image, wherein the analysis result comprises at least one of classification, regression, and segmentation.
22. The computer-executable deep learning network of claim 21 , wherein: The first SSL learning objective comprises DINOv2, and the first encoder comprises a visual transformer (ViT) encoder.
23. The computer-executable deep learning network of claim 21 , wherein: The first SSL learning objective comprises DINOv1, and the first encoder comprises a visual transformer (ViT) encoder.
24. The computer-executable deep learning network of claim 21 , wherein: The first SSL learning objective includes a comparative learning objective.
25. The computer-executable deep learning network of claim 24, wherein: The comparative learning targets include simCLR.
26. The computer-executable deep learning network of any one of claims 24 to 25, wherein: The first encoder includes a convolutional neural network (CNN).
27. The computer-executable deep learning network of any one of claims 22 to 26, wherein: The second SSL learning objective includes DINOv2.
28. The computer-executable deep learning network of any one of claims 22 to 26, wherein: The second SSL learning objective includes DINOv1.
29. The computer-executable deep learning network of any one of claims 22 to 26, wherein: The second SSL learning objective includes Visual Transformer Mask Auto-Encoding (ViTMAE).
30. The computer-executable deep learning network of any one of claims 21 to 29, wherein: The attention-based machine learning analyzer includes a visual transformer (ViT) encoder.
31. The computer-executable deep learning network of any one of claims 21 to 29, wherein: The attention-based machine learning analyzer includes a convolutional neural network (CNN) and a multi-instance attention block configured to assign attention weights to outputs of the CNN to provide an aggregated representation of the WSI.
32. The computer-executable deep learning network of any one of claims 30 to 31, wherein: The attention-based machine learning analyzer includes one or more feed-forward neural network layers.
33. The computer-executable deep learning network of any one of claims 21 to 32, wherein: An image region corresponds to a portion of the image that is at least 100 times larger than the image patch.
34. The computer-executable deep learning network of any one of claims 21 to 32, wherein: An image region corresponds to a portion of the image that is at least 200 times larger than the image patch.
35. The computer-executable deep learning network of any one of claims 21 to 32, wherein: The image region corresponds to a portion of the image that is at least 400 times larger than the image patch.
36. The computer-executable deep learning network of any one of claims 21 to 35, wherein: An area of the WSI having a size of at least 100×100 pixels and not more than 1000×1000 pixels is included in the patch.
37. The computer-executable deep learning network of any one of claims 21 to 36, wherein: The region includes an area of the WSI having a size of at least 1000×1000 pixels and not more than 10000×10000 pixels.
38. The computer-executable deep learning network of any one of claims 21 to 37, wherein: A patch includes an area of the WSI having a size equal to or approximately equal to 224×224 pixels (eg, 224+ / −50×224+ / −50 pixels).
39. The computer-executable deep learning network of any one of claims 21 to 38, wherein: The region includes a region of the WSI having a size equal to or approximately equal to 4480×4480 pixels (eg, 4480+ / −500×4480+ / −500 pixels).