Clustering-based histopathologic phenotypic representation learning through self-supervised multi-class label hierarchical vision TRANSFORMER

A self-supervised learning framework using a multi-class labeled hierarchical Vision Transformer model addresses the challenges of high-cost dataset assembly in digital pathology by leveraging unlabeled data to learn domain-specific features, enhancing model accuracy and robustness in medical imaging analysis.

CN120322798APending Publication Date: 2025-07-15VENTANA MEDICAL SYSTEMS INC +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380084241.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-08
Filing Date
2023-12-07
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

In medical imaging analysis tasks, especially in the field of digital pathology, acquiring large and high-quality labeled data sets is costly and difficult, with great differences between observers, and existing unsupervised machine learning methods have problems of feature differences and extended training processes when processing histopathological images.

Method used

Unsupervised clustering is used for unsupervised clustering, and self-supervised multi-class labeled hierarchical ViT (Cypher ViT) is used as the backbone encoder, combining histopathology-specific proxy tasks to capture semantically meaningful fine-grained regions of interest, and use unlabeled data to learn field-specific background knowledge.

Benefits of technology

It improves the accuracy and robustness of digital pathological image analysis, reduces the dependence on high-cost annotation data, enhances the generalization ability of the model in downstream tasks, and realizes efficient histopathological feature extraction and classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120322798A_ABST
    Figure CN120322798A_ABST
Patent Text Reader

Abstract

The present disclosure relates to a system and method for processing a digital pathological image using a machine learning model that includes a self-supervising hierarchical vision Transformer (ViT) configured to perform unsupervised clustering with a plurality of classification markers. The method includes receiving a digital pathological image depicting a tissue section stained with a histological dye. The digital pathology image may be processed to generate a plurality of predicted classification results including individual tiles of the digital pathology image. The result is generated by a machine learning model using a self-supervised hierarchical vision Transformer (ViT), which may further include a multi-head self-attention module configured to predict, for each individual tile in the digital pathology image, a cross-tile correlation metric using an attention mechanism, and to predict, for each individual tile in the digital pathology image, a cross-tile correlation metric based on the predicted cross-tile correlation metric. Thereby assigning the individual tiles to clusters based on the cross-tile correlation metric.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims the benefit and priority of U.S. Provisional Application No. 63 / 386,617, filed on December 8, 2022, which is incorporated herein by reference in its entirety for all purposes. Background Art

[0003] Access to large - scale and high - quality datasets can prove to be a major driver for machine learning. For example, ImageNet is a dataset that has been used to train computer vision models that perform well in processing natural images. At the same time, for medical image analysis tasks, labeled data can be scarce and expensive because annotations from multiple experts may often be required, and crowdsourcing may not be an option. Additionally, inter - observer variability among medical experts can affect the quality of the dataset. Thus, assembling large and high - quality datasets for medical imaging analysis tasks can often be prohibitively costly and time - consuming, which may limit the progress of research and model development in this field.

[0004] Unsupervised machine learning is a method of using unlabeled data to train models. Unsupervised learning can provide solutions to the above - mentioned challenges and facilitate the development of more accurate artificial intelligence (AI) models. Transfer learning is a technique where a model can be pre - trained (e.g., using the ImageNet dataset) and then fine - tuned using the data type of interest (e.g., medical images). This approach is advantageous because the ImageNet dataset is typically much larger than medical datasets, thus providing a well - grounded model for understanding basic image features. However, challenges arise because the potential differences in features and patterns between natural scene images from ImageNet and medical images hinder model convergence and may prolong the training process.

[0005] Histopathology has widely adopted digitization, providing a unique opportunity to improve the objectivity and accuracy of diagnostic interpretation through machine learning. Among other factors, digital images of tissue specimens may exhibit significant complexity and heterogeneity in preparation, fixation, and staining protocols. This diversity may further exacerbate the availability of large labeled datasets in digital pathology compared to any other medical imaging modality. Additionally, each tissue specimen image is typically a gigapixel file, which may require significantly more labeling effort from experts, resulting in higher inter - observer / intra - observer variability and mislocalization of lesions. These challenges may reinforce the need to utilize unsupervised machine learning methods to leverage large amounts of unlabeled data in the field of digital pathology. Summary of the Invention

[0006] Some embodiments of the present disclosure relate to using a machine learning model to process digital pathology images, the machine learning model including a self-supervised hierarchical vision Transformer (ViT) configured to perform unsupervised clustering. A computer-implemented method includes receiving a digital pathology image depicting a tissue section stained with a histological dye (e.g., without any accompanying pathologist annotations). Processing the digital pathology image to generate a plurality of predicted classification results including individual patches of the images (e.g., in an unsupervised manner). The results are generated by a machine learning model that uses a self-supervised hierarchical ViT as a backbone encoder to capture fine-grained regions of interest that are detailed to the pixel level in terms of semantics. The hierarchical ViT further includes a multi-head self-attention module configured to use an attention mechanism to predict cross-patch correlation metrics for each individual patch in the digital pathology image. Based on the cross-patch correlation metrics, the individual patches can be assigned to clusters.

[0007] The plurality of predicted classification labels can indicate, predict, or correspond to one or more of the following: tissue type characterizing histological features (e.g., cytological features), magnification level, diagnostic category (e.g., non-diagnostic category, malignant negative category, atypical category, tumor: benign category, suspicious category, or malignant positive category), or a prediction regarding whether the digital pathology image depicts a specific histological feature of a malignant tumor or the extent to which the digital pathology image depicts a specific histological feature of a malignant tumor. Specific histological features can include high cell density, cell enlargement, lack of cell adhesion, high nuclear-cytoplasmic ratio, nuclear hyperchromasia, prominent nucleoli, large nucleoli, abnormal nuclear-chromatin distribution, high mitotic activity, abnormal nuclear membrane, cell pleomorphism, nuclear pleomorphism, or tumor diathesis.

[0008] The plurality of predicted classification labels can characterize what is being depicted in a part or all of the digital pathology image. For example, the part can include patches or pixels.

[0009] In some embodiments, a system is provided that includes one or more data processors and a non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of one or more of the methods disclosed herein.

[0010] In some embodiments, a computer program product tangibly embodied in a non-transitory machine-readable storage medium includes instructions configured to cause one or more data processors to perform some or all of one or more of the methods or processes disclosed herein.

[0011] In some embodiments, a system is provided that includes one or more devices for performing some or all of the one or more methods or processes disclosed herein.

[0012] The terminology and expressions employed are used in a descriptive rather than a restrictive sense, and in using such terminology and expressions, no intention is made to exclude any equivalents of the features shown and described or portions thereof, but it should be recognized that various modifications are possible within the scope of the invention as claimed. Accordingly, it should be understood that although the claimed invention has been specifically disclosed by way of embodiments and optional features, those skilled in the art may adopt modifications and variations of the concepts disclosed herein, and such modifications and variations are considered to be within the scope of the invention as defined by the appended claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] This patent or application file contains at least one color drawing. After a request is made and the necessary fee is paid, the Patent Office will provide a copy of this patent or patent application publication with one or more color drawings.

[0014] Figure 1 An exemplary workflow of a self-supervised learning (SSL) framework is shown.

[0015] Figure 2A An illustrative example of the self-supervised contrastive method DINO (Distillation INformation by Non-parametric contrast) is shown.

[0016] Figure 2B Shown from Figure 2A is an illustrative example of a vision transformer.

[0017] Figure 3 Shown from Figure 2A is an example process of chunking and embedding.

[0018] Figure 4 An exemplary implementation of backbone transformer encoder-based clustered histopathological phenotype representation learning via a self-supervised multi-class labeled hierarchical vision transformer (Cypher ViT) according to an embodiment of the present disclosure is depicted.

[0019] Figure 5 Shown from Figure 4 is an example architecture of the Cypher ViT attention module.

[0020] Figure 6 An example operation of an attention mechanism according to an embodiment of the present disclosure is shown.

[0021] Figure 7An example flowchart of a computer-implemented method for processing digital pathology images using a machine learning model according to some embodiments of the present disclosure is shown.

[0022] Figure 8A A two-dimensional (2D) UMAP (Uniform Manifold Approximation and Projection) visualization of feature embeddings extracted from an example implementation of a state-of-the-art self-supervised learning (SSL) framework pre-trained on a dataset is shown.

[0023] Figure 8B A 2D UMAP visualization of feature embeddings extracted from the state-of-the-art SSL framework iBOT (Image BERT Training with an Online Tokenizer) and an example implementation of the present disclosure pre-trained on a dataset is shown.

[0024] Figure 9 Retrieval results of a set of query images according to an example implementation of the present disclosure are shown.

[0025] Figure 10A Attention maps of multiple predicted classification labels extracted from learnable multi-class labels at the final stage of an example implementation of the present disclosure are shown.

[0026] Figure 10B Attention maps of multiple predicted classification labels extracted from learnable multi-class labels at the final stage of an example implementation of the present disclosure are shown. Detailed Description

[0027] Some embodiments of the present disclosure relate to using a machine learning model to process digital pathology images, the machine learning model including a self-supervised hierarchical vision Transformer (ViT) configured to perform unsupervised clustering with multiple classification labels. More specifically, the machine learning model can be configured with multiple levels, each level assembling semantically similar patches into a fixed number of classification labels. An attention mechanism can use the classification labels to determine the degree of similarity of various patches. For example, the classification labels can be used to predict the degree of relevance of a given patch to each of one or more other patches, which can then be used to support the assignment of individual patches to clusters. The classification labels are learned in a self-supervised manner and / or may be related to histological semantics.

[0028] In some embodiments, a framework is provided for performing clustering-based histopathological phenotype representation learning in an SSL pipeline using a self-supervised multi-class token hierarchical ViT (Cypher ViT) as a novel backbone encoder in place of a conventional ViT (Dosovitskiy et al., “An image is worth 16x16 words: Transformers for image recognition at scale” (arXiv preprint arXiv:2010.11929, 2020), which is incorporated herein by reference in its entirety). In some embodiments, the SSL pipeline is structured according to the following:

[0029] · DINO (Information Distillation via Non-Parametric Contrast) (see Caron et al., “Emerging properties in self-supervised vision transformers”. Proceedings of the IEEE / CVF International Conference on Computer Vision, pp. 9650-9660, 2021, which is incorporated herein by reference in its entirety for all purposes);

[0030] · MOCO (see He, Kaiming et al., “Momentum contrast for unsupervised visual representation learning”. Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2020, which is incorporated herein by reference in its entirety for all purposes);

[0031] · SimCLR (see Chen, Ting et al., “A simple framework for contrastive learning of visual representations”. International Conference on Machine Learning. PMLR, 2020, which is incorporated herein by reference in its entirety for all purposes).

[0032] This approach can encourage self-supervised learning (SSL) models to capture semantically meaningful fine-grained regions of interest detailed down to the pixel level. In accordance with an unsupervised clustering scheme, the single-class tokens in a conventional ViT are extended to a set of learnable multi-class tokens, thereby assembling coarse-to-fine-grained features into semantically aware clusters in a hierarchical manner.

[0033] SSL technology has traditionally been used to process natural images and has not been used to process digital pathology images. Adapting existing SSL methods to histopathology data can be challenging because the features important in the digital pathology context (e.g., cell density, cell morphology, co-localization of dyes, etc.) are completely different from the features extracted from natural images.

[0034] To incorporate histopathology-specific knowledge into a self-supervised contrastive learning framework, hybrid methods can be deployed. These methods can combine contrastive learning customized for histopathology patches and domain-specific pretext tasks, which are designed based on the characteristics of histopathology images, such as predicting magnification levels, predicting hematoxylin channels, predicting cross-staining, and color normalization. However, focusing on certain unique histopathology characteristics during SSL pre-training may compromise the generality of the model required as a general feature extractor. For example, the model can selectively focus on color differences in cross-staining prediction or, alternatively, on the association between spatial proximity and semantic proximity in the feature space. However, there is no guarantee for the premise that adjacent patches are also more likely to be adjacent at the feature level than distant patches. Therefore, noisy positive and negative pairs will endanger the network training.

[0035] For example, the self-supervised multi-class labeled hierarchical ViT is a novel backbone that captures both coarse-grained and fine-grained features. (ViT may have been tested on one or more frameworks such as the DINO, MOCO, and / or SimCLR SSL frameworks). Compared with ImageNet pre-training and other state-of-the-art SSL methods, this model has at least two advantages: it produces much higher-quality features compared with other state-of-the-art SSL frameworks and patch retrieval demonstrations; it learns more precise morphological phenotypes down to the pixel level, which is different from the grid-structured attention maps extracted from the multi-head attention heads of a conventional ViT. In addition, when AI algorithms are trained on data from a limited group of subjects, the generalization gap is usually larger, and this limited group of subjects may not be representative of the actual population. The disclosed SSL-based paradigm can help bridge the gap by building more general models that learn from a larger population of subjects, and this is possible because manual labeling is not required in SSL. The robustness and transferability of the model are further verified in exhaustive experiments in downstream tasks such as unsupervised and semi-supervised patch classification of tissue types and fine-grained classification of histological features (e.g., cytological features).

[0036] By extending the class tokens in a conventional ViT in a hierarchical manner until the final stage, the unique design of the backbone encoder has the potential for two possible extensions in addition to serving as a general feature extractor. If trained in the SSL paradigm, by equipping each histopathology image with a list of domain-specific attributes as supervision signals for multiple auxiliary (proxy) tasks (e.g., magnification level, hematoxylin channel), each class token at the final stage can be customized to predict each label that simultaneously guides each auxiliary task; if trained with some supervision signals as in the weakly supervised setting, each class token at the final stage can be customized to uniquely learn the target lesion area.

[0037] Clinical AI models require large amounts of highly curated datasets that are carefully annotated by multiple medical experts, increasing the development time and cost. This disclosure leverages SSL techniques in the context of digital pathology. Self-supervised learning (SSL) is a form of unsupervised learning method that allows AI models to utilize unlabeled data to acquire domain-specific background knowledge to improve performance and generalization regarding various downstream learning tasks. SSL is a form of unsupervised learning that is designed to learn domain-specific salient features from large amounts of unlabeled data. SSL methods can enable AI models to acquire domain-specific background knowledge from large amounts of existing unlabeled data. It learns visual representations based on supervision signals that are entirely derived from the data itself. SSL can enable AI models to discover domain-specific background knowledge about the data without labels from subject matter experts. This means that high-level general knowledge of the domain can be learned from unlabeled data, and task-specific information or skills can only be learned from labeled data in a supervised manner. The knowledge obtained through SSL provides an improved starting point for AI models to converge to a more robust and generalizable solution with a smaller amount of labeled training data.

[0038] Figure 1Illustrates an example workflow 100 of an SSL model for a given dataset 105. SSL relies heavily on unlabeled data 110, such that the machine learning model 120 is not given explicit annotations 112, but is trained to create its own understanding of the data by generating auxiliary tasks also known as surrogate tasks 130, which are inherently related to the data itself. These surrogate tasks 130 are typically designed to encourage the machine learning model 120 to learn meaningful representations 125. For example, predicting missing parts of an image, assembling a jigsaw puzzle, differentiating transformed versions of the same data, predicting the order of sentences in a document or the order of words in a sentence. As such, domain-specific background knowledge is acquired to improve its performance and generalization on various downstream learning tasks 150. SSL implementations focus on developing domain-agnostic / specific surrogate tasks 130 for unlabeled data 110 to derive supervision signals during model training. The principle of developing surrogate tasks 130 is to utilize supervision signals inferred from the unlabeled data 110 itself, without depending on any external guidance. Well-learned representations 125 of the original data are used as an initialization starting point and for performance improvement to further facilitate the training of the desired downstream tasks 150.

[0039] After pre-training the SSL surrogate tasks 130, the pre-trained model 120 can be fine-tuned on a smaller labeled dataset 115 for a specific downstream task 150. This transfer learning 135 process leverages the knowledge acquired during self-supervised pre-training to improve performance on downstream tasks 150 with limited labeled data 115. It is worth noting that the pre-trained machine learning model 120 and the downstream model 140 to be used for the downstream task 150 can be similar or different, depending on the specific implementation and requirements. In some cases, the pre-trained model 120 can be directly used for the downstream task 150. The idea is that the features or representations 125 learned during pre-training can also be useful for performing similar tasks. In other cases, the pre-trained model 120 can be fine-tuned for the downstream task. Fine-tuning can involve updating the parameters of the pre-trained model 120 to adapt to the specific labeled data 115 and downstream task 150. Alternatively, a different model can be trained for the downstream task 150, e.g., using ViT as the surrogate model 120 and using a convolutional neural network (CNN) as the downstream model 140 for image classification. In this example, ViT is used as a feature extractor and CNN is used as a classifier.

[0040] Thus, in pre-training, the model 120 is empowered to extract coarse-grained and / or fine-grained features from an image dataset (e.g., digital pathology images). Once the downstream model 140 is also trained, an image can be tested by first feeding it to the pre-trained model 120 to extract features and then feeding it into the downstream model 140 for downstream tasks (e.g., classification, captioning, segmentation, etc.). Put another way, the key to the success of the SSL model may lie in the wise utilization of the information obtained from the images themselves during pre-training.

[0041] In Figure 2A it, an illustrative example of the self-supervised contrastive method DINO (Information Distillation via Non-Parametric Contrast) is shown. The example architecture of DINO can include a student network 210 and a teacher network 215. Due to different update methods, these two networks, the student 210 and the teacher 215, can have similar architectures but different learnable weights. The network 200 learns in a self-supervised setting through a process called knowledge distillation. Distillation can refer to the process of transferring knowledge from the teacher network 215 to the student network 210. The teacher-student network involves training the teacher network 215 to produce reference representations (e.g., z′2, z2) for the training samples 110. In turn, the student network 210 is trained to mimic the representations, and this process can be called knowledge distillation. The teacher 215 can be a momentum teacher, which means that the weights of the teacher network 215 are the exponentially weighted average of the weights of the student network 210. DINO can use the learning objective 270 to distinguish representations of different augmentations of the same image using a memory bank of features from previous instances in the training data.

[0042] DINO can define proxy tasks that the model needs to learn during training. The proxy tasks can involve augmenting the input unlabeled data 110 and training the model to distinguish different augmentations (e.g., V240 and V′245) of the input image in a self-supervised manner. For example, as depicted in FIG. 2, DINO takes an image x from the unlabeled dataset 110 and applies two different transformations or augmentations 240 and 245 to produce two different views V and V′, which are to be fed into the student 210 and teacher network 215 pipeline.

[0043] Multi-crop augmentation can be applied to extract two sets of images (possibly partially overlapping) from the transformed views V and V'. Small crops can be referred to as local views 220 (e.g., <50% of the image), and large crops (e.g., >50% of the image) can be referred to as global views 225. In other words, the set of global views 225 has a higher dimension than the set of local views 220. All crops pass through the student 210, while only the global views 225 pass through the teacher 215. This promotes a "local-to-global" correspondence, training the student 210 to infer context from small crops. During training, only the student 210 is trained so that the network group can understand that the local and global representations, although significantly different, represent the same subject. It is worth mentioning that multi-crop augmentation and random transformation can be applied in any order. For example, local views 220 and global views 225 can be implemented from the input image x, and then random augmentations (e.g., color jitter, Gaussian blur, solarization, etc.) can be applied to the local views 220 and global views 225 to make the network more robust.

[0044] Before feeding these views into the vision transformers (ViTs) 235 and 237, the views can be passed through the patch and embedding block 230 to obtain enhanced embedding vectors. The patch and embedding block 230 can convert the image into equally sized patch tokens and perform a set of operations to obtain the corresponding embedding vectors for each patch. These enhanced input embedding vectors to the ViTs 235 and 237 can represent the embedding sequence of patch tokens, a learnable multi-class token prefixed to the sequence, and positional information.

[0045] It should be understood that (e.g., MOCO or SimCLR) can be used instead of DINO or as a supplement to DINO.

[0046] Figure 2B An illustrative example of a ViT from Figure 2A is shown. Due to different update methods, the two vision transformers (ViTs) 235 and 237 in the student and teacher networks can have the same architecture but different learnable weights. The vision transformer (ViT) can be a type of neural network based on the transformer architecture. The enhanced embedding vectors from the patch and embedding 230 are fed into the ViT, which can further include a transformer encoder 235a with learnable weights θ in the student network 210 and a multi-layer perceptron (MLP) 235c, and a transformer encoder 237a in the teacher network 215 and MLP 237c. These Transformer encoders can represent a stack of multiple self-attention layers. Self-attention can refer to a mechanism that allows the model to learn long-range dependencies between chunks for tasks such as image classification, as it can allow the model to learn how different parts of an image can contribute to its overall label. The output of the Transformer encoder is a sequence of vectors, e.g., for the student network 210, it is (y1, y′1)235b, which represents a group of intermediate features of the global view 225 and the local view 220, and for the teacher network 215, it is (y2, y′2)237b, which represents a group of intermediate features of the global view 225.

[0047] These intermediate representations can be fed into the MLPs (235c and 237c) of the student-teacher network. The MLPs (235c and 237c) in the teacher-student network can be connected after the MLP head 260. The MLP can act as a pointwise feed-forward neural network consisting of multiple linear transformations and non-linear activations. It can independently apply non-linear transformations to each position of the input sequence to generate a group of projections or embeddings (z1 and z′1)235d and (z2 and z′2)237d for the corresponding student network 210 and teacher network 215. The MLP layer may help increase the expressiveness and representational power of the Transformer encoder.

[0048] The group of projections generated from the MLP is fed into the MLP head 260a for the student network and the MLP head 260b for the teacher network to generate a group of probabilities (q1, q′1) and (q2 and q′2) for the corresponding student network 210 and teacher network 215. In the context of DINO, the MLP head can represent a component of the projection head that is responsible for transforming the input features into a space where the learning objective 270 can be applied. The projection heads (260a and 260b) can be a layer (e.g., an average pooling layer or SoftMax) or a small MLP that takes the embeddings or projections (i.e., 235d and 237d) from the corresponding branches as input and predicts the representations of positive pairs (augmented or transformed views of the same image) and distinguishes them from negative pairs (representations from different images). The loss objective 270 aims to maximize the similarity between the two sets of projections 235d and 237d from the same input while minimizing the similarity with the projections of other images in the same mini-batch. For contrastive metric measurement, DINO adopts the cross-entropy loss.

[0049] In the DINO framework, mode collapse may occur during training. There can be two forms of mode collapse: the model output can be the same along all dimensions regardless of the input (i.e., the same output for any input), or it can be dominated by one dimension. Centering and sharpening can be deployed in the teacher network 215 before the prediction head to prevent these two problems. Sharpening 250 can refer to the process of refining the learned representations to make them more distinct and explicit. In sharpening 250, additional operations (such as feature scaling, gradient clipping, temperature scaling, etc.) can be applied to the projections from the MLP 237c (i.e., 235d and 237d) to enhance the features. The goal of centering 255 is to improve the clustering or concentration of similar instances in the learned feature space. Centering 255 can involve normalization, whitening, spatial attention mechanisms, or other techniques for improving the clustering of similar instances. Both centering 255 and sharpening 250 are performed to improve the discriminative ability and clarity of the features learned during the self-supervised contrastive learning process.

[0050] DINO has an asymmetric architecture for the student 210 and teacher 215 network pipelines, where the teacher encoder 237a 's weights θ are updated via the exponential moving average (EMA) from the student encoder 235a during backpropagation. The update rule is θ t ← λθ t + (1 - θ t )θ s , where λ follows a cosine schedule during training. The output probabilities from the teacher network 215 are considered as the supervision signal guiding the training of the student network 210. The distribution of the student network 210 can be matched to the teacher network 215 for the input image x by minimizing the cross-entropy loss function with respect to the parameters of the student network (i.e., θ s , ψ s ), as given in Equation 1.

[0051]

[0052] As Figure 2A shown, the loss in Equation 1 can be adapted to the self-supervised learning problem that deploys a multi-crop strategy with local views 220 and global views 225 enhanced from the original input image x. In some embodiments, multi-crop augmentation is important, but there is an optimal sweet spot in the number of local views, which is considered an adjustable hyperparameter. The global view can be represented as and several local views of smaller resolutions are represented as where N l can refer to the number of local views, and N gIt can refer to the number of global views. For simplicity, Figure 2A only two global views are shown in Figure 2A . The loss objective function 270 given in Equation 1 can be modified as given in Equation 2, where and represent the output probability distributions of the teacher network 215 and the student network 210, respectively. Gradient propagation is stopped in the teacher network 215, and only the gradient 265 is allowed to pass through the student network 210.

[0053]

[0054] The standard transformer receives an input of a one-dimensional (1D) sequence of token embeddings as it was originally designed for natural language processing (NLP). To process two-dimensional (2D) digital pathology images, the image is reshaped into a sequence of flattened 2D patches. Figure 3 The example process of patch and embedding 230 from Figure 2 is further elaborated. To structure the input image, the image 305 can be passed into the patch and embedding block 230, which can convert the image 305 into a grid 310 of non-overlapping, equal-sized 2D patches, where each patch is treated as a separate entity. Each image patch (also called a token) is flattened into a 1D vector and then passed to a trainable linear projection 320 to transform the 1D high-dimensional patch into a lower-dimensional vector or embedding 340. This transformation can be achieved by applying a linear transformation (e.g., a fully connected layer with fewer output dimensions) on the patch embedding. The purpose of dimensionality reduction is to make the processing computationally more efficient while still capturing the essential information from the input patches. Since the vision transformer does not inherently capture the spatial information of the input, positional information needs to be incorporated. The positional embedding 325 is added to the patch embedding 340 to preserve the spatial arrangement of the patches.

[0055] In addition to the positional and patch embeddings 340, a special token called the [cls] token 330 can be introduced. The semantic image layout can be discovered from the attention maps of the [cls] token. These attention maps may yield promising results in unsupervised segmentation tasks. In some embodiments, different from conventional transformers, multiple [cls] tokens 330 are used. Using a single [cls] token may be challenging for accurately localizing different objects on a single image. Thus, instead of a single [cls] token, multiple [cls] tokens 330 can be used, which will be responsible for learning the representations of different object classes. By doing so, the model can learn to focus on regions of the image belonging to each class and generate class-discriminative object localization maps from the class-to-patch attention. This technique can be useful for weakly supervised semantic segmentation, which is the task of assigning class labels to each pixel in an image using only image-level labels as supervision. The output of the linear projection block 345 (a combination of the patch embedding 340, the positional encoding 325, and the multiple [cls] tokens 330) forms the input to the ViT.

[0056] In some embodiments, an SSL-based framework is provided that utilizes a large amount of unlabeled digital pathology data to improve the extent to which the model is used to generate digital pathology label predictions, where the model is generalizable and robust. The system may include a self-supervised backbone transformer (such as Figure 4 the Cypher ViT 405 shown) in place of a conventional ViT (e.g., 235 and 237). Similar to the ViT architecture 235, the extended [cls] tokens 330 have been included along with additional hierarchical feature aggregation attention modules (e.g., 410, 415, 420).

[0057] As described above, the input image is first split into non-overlapping patches, and then these non-overlapping patches are transformed into a sequence of patch tokens 340 and positional embeddings 325. These [cls] tokens are concatenated with the patch tokens 340 and the embedded position information 325 to form the input tokens 345 for the transformer encoder. For illustrative purposes, three consecutive layers of the Cypher ViT attention module with the same structure (i.e., 410, 415, and 420) are shown in Figure 4 . The Cypher ViT 405 may further include an average pooling layer 425 to compute attention scores. It can also use an MLP for classification prediction.

[0058] Targets with class - specific tokens cannot be achieved by simply increasing the number of class tokens in ViT because these class tokens may still have no specific meaning. To be able to effectively learn high - level discriminative features of each class token for a specific object class, a class - aware training strategy for multiple class tokens 330 can be adopted. More specifically, an average pooling layer 425 can be applied along the embedding dimension to the output class tokens from the last stage of the Cyber ViT attention module 420 to generate class scores, which are directly supervised by the ground - truth class labels. Thus, a one - to - one strong connection can be established between each class token and the corresponding class label. With this design, a significant advantage may be that the learned class - to - patch attentions for different classes can be directly used as class - specific localization maps.

[0059] Conventional ViT models (e.g., 235 and 237) maintain the full - length sequence during the forward pass across multiple consecutive layers of the ViT. This design may suffer from redundancy and lack of multi - level hierarchical representations that may contribute to successful recognition tasks in digital pathology images. One solution can be to gradually reduce the sampling of the sequence length as the model gets deeper. At each stage of CypherViT 405, the number of learnable class tokens 330 can be progressively reduced, which is driven by the intuition that as more abstract features are acquired, the features can be grouped into a smaller number of clusters.

[0060] Figure 5 Shows an example architecture of the Cypher ViT attention module 410 from Figure 4 The Cypher ViT attention module 410 can include multi - head self - attention 510 and semantic clustering 520. In the Cypher ViT attention module 410, the input embedding 505 is passed to the multi - head self - attention block 510. Due to the hierarchical structure of Cypher ViT 405, the input embedding 505 at each stage is different, i.e., the output of the previous stage becomes the input of the next stage. For example, at stage 1 410, the input embedding 505 comes from the output 345 of the patching and embedding block 230, which is the concatenation of the patch embedding 340, the positional encoding 325, and the multi - class tokens 330.

[0061] The multi - head self - attention module 510 can refer to a component of a vision transformer that allows each input token to attend to every other token in a parallel and efficient manner. The number of heads in the multi - head self - attention module 510 is a hyperparameter that can be selected based on the task and the model architecture. Each head represents a different subspace of the input embedding and can learn to attend to different parts of the input sequence. For each head, query (q), key (k), and value (v) vectors of the same size are obtained using linear projections through q = W o M, k = Wk M, v = W v is calculated by M, where M is the input embedding vector 505, and W q 、W k and W v are the learned weighted matrices for each vector. Then, the scaled dot - product attention function can be applied as , where f can represent the scaling factor. These vectors are used to calculate the correlation scores and weighted outputs for each input token. The number of heads can affect the dimensions of the query, key, and value matrices as well as the output of the self - attention module. Typically, the number of heads is a factor of the model dimension to be maintained, e.g., 8, 12, or 16. Then, the outputs of different heads are concatenated and projected to produce the final output of the module.

[0062] The multi - head attention module 510 can allow each token to attend to every other token in the sequence and produce a new feature map. The output embedding of the multi - head attention module 510 can be split into P1 (chunk tokens) and C1 ([cls] token set) as inputs, and the self - attention mechanism is applied twice in the semantic clustering block 520. The semantic clustering module 520 can take these chunks as inputs and perform clustering. The output of the semantic clustering module 520 is a sequence of the clustered tokens, which can be fed into the next Cypher ViT attention module (e.g., 415 or 420) or used for downstream tasks.

[0063] The embedded vector 345 to the multi - head self - attention module 510 can be a high - dimensional feature map that captures the global semantic information of the image. Thus, semantic clustering 520 can aim to reduce the computational complexity of self - attention in the vision transformer. It can work by grouping visual tokens with similar semantic information into clusters and then aggregating the key and value tokens within each cluster. In this way, the number of tokens is reduced, and self - attention can be performed more efficiently. The self - attention of a single head can be reformulated for semantic clustering 520, Clust where the decision value γ can be calculated to locate the density peak of the cluster. Semantic clustering 520 can also preserve the global context and diversity of the original tokens, which is beneficial for visual representation learning. For clustering, the semantic clustering block 520 can apply a clustering algorithm (e.g., K - means, hierarchical clustering, or DBSCAN, etc.) to group the pixels in the feature map into different clusters based on their similarity. Each cluster can represent a potential object category in the image. At the feature level, following a bottom - up approach without being interfered by external supervision signals, each attention module 510 aggregates the chunk tokens with semantically similar visual concepts into a fixed number of clusters, as Figure 6 shown.

[0064] Figure 6Shows an example operation of the attention mechanism according to an embodiment of the present disclosure. As described above, each attention module 410 in CypherViT may mainly include two blocks: a multi-head self-attention block 510 for exploring cross-chunk correlations, followed by a semantic clustering block 520 for assembling similar tokens together. Then, the intermediate results of the inherently aggregated features obtained through the self-attention mechanism can be retained in a set of multi-class tokens, which are defined with learnable weights during backpropagation. Then, the tokens with the merged features can be fed as new inputs to the next stage in the hierarchical clustering pyramid, as Figure 4 shown, until it reaches the last stage. At the last stage, each token may inherently capture a specific visual concept corresponding to histological phenotypes such as cells, stroma, white space, etc.

[0065] To elaborate on the multi-head self-attention block 510, as Figure 6 shown, the input includes two components: the chunk tokens 340 and the class or classification [cls] tokens 330. In Figure 6 , the chunk embeddings 340 also include positional encodings (omitted in the figure for simplicity of presentation). The chunk tokens 340 remain the same as in a conventional ViT and are denoted as where N p refers to the number of chunk tokens. However, the expansion of [cls] to multi-[cls] token groups - denoted as - where N c refers to the number of learnable class tokens, where s refers to the stage, and g refers to the stage number (a non-learnable fixed hyperparameter). At each stage s g , the number of learnable class tokens (N c ) can be progressively reduced by applying clustering. Next, at the starting stage, the concatenation of the chunks with the multi-[cls] tokens is fed as input to the multi-head self-attention block 510 denoted as , where d represents the latent space dimension. Inside the multi-head attention block 510, linear transformations are applied to generate queries, keys, and values at stage 1, as: where M1 can represent the output, and M represents the input at stage 1 of the multi-head self-attention module 510, h represents the number of attention heads, and each head operates on a transformed version of the input. Subscripts (e.g., k1, v1, M1) can represent the stage numbers for the corresponding values in the Cypher ViT hierarchy. In Figure 6 , @ can refer to the scalar product of two vectors.

[0066] As the input to the upcoming semantic clustering block 520, 610 can be split into those for the chunk tokens and those for the [cls] token set and the self-attention mechanism can be used as applied twice. Similarly, the same applies to and as well.

[0067] C1, C2, and C3 are used for equation simplification, while W q , W k and W v are learnable weights in the linear projection to obtain the corresponding queries, keys, and values for each stage. The attention calculation from the above module highlights the similarity between the learnable multi-[cls] tokens and . In Figure 6 , only the attention block in the first stage is shown. However, for further stages starting from the first stage, instead of chunk and multi-[cls] tokens, the input becomes the multi-[cls] tokens from the previous stage and the current stage, i.e., the similarity matrix between the measurement and will be measured, where g > 0, as shown in Figure 4 . To formulate it mathematically, the attention vector for i chunks to j classes at stage g(s g ) can be the SoftMax of the similarity matrix, as given in Equation 3.

[0068]

[0069] The learnable multi-[cls] tokens as the input to the next stage can be obtained through the following Equation 4, where W, W c , W p and W v refer to the learnable weights in the linear projection.

[0070]

[0071] Once the teacher-student network is pre-trained, the distilled student model learned from the teacher's knowledge is then fine-tuned or directly used for downstream tasks. It is worth noting that in the SSL design, the number N l of local views and the number of class tokens may be irrelevant. These two parameters can be considered as two independent tunable hyperparameters.

[0072] Figure 7FIG. 0 shows an example flow diagram of a computer-implemented method for processing digital pathology images using a machine learning model. At block 705, a digital pathology image is received that depicts a tissue section stained with a histological dye, and the image is split into a plurality of equally sized patches. The process at 710 involves using a machine learning model to process the digital pathology image, the machine learning model including a self-supervised hierarchical vision transformer (ViT) configured to perform unsupervised clustering with a plurality of classification tokens. The self-supervised hierarchical ViT further includes a multi-head self-attention module configured to use an attention mechanism to predict cross-patch correlation metrics for each individual patch in the digital pathology image. At block 715, the individual patches are assigned to clusters based on the cross-patch correlation metrics. Finally, at block 720, a plurality of predicted classification results including the digital pathology image are generated by the machine learning model using Cypher ViT.

[0073] Example implementation:

[0074] An example implementation of the framework is provided to perform clustering-based histopathological phenotype representation learning in a self-supervised multi-class token hierarchical ViT (CypherViT) 405 (as Figure 4 shown) as a novel backbone encoder replacing the conventional ViT 235 in the SSL pipeline inherited from DINO. (It should be understood that various other encoders, such as those inherited from MOCO or SimCLR, can be used.) This approach has proven to capture semantically meaningful fine-grained regions of interest down to the pixel level. In accordance with the unsupervised clustering scheme, the single-class token in the conventional ViT is extended to a set of learnable multi-class tokens 330, thereby assembling coarse-to-fine-grained features into semantically aware clusters in a hierarchical manner.

[0075] In the following demonstration, all models were pre-trained without using any tags. The processing was distributed across 4 GPUs connected in parallel. The batch size for each GPU was 100, where AdamW was used as the optimizer for the student network. For the chunk embedding, the chunk size was set to 16. Following the implementation in DINO, the base learning rate of 5e-04 and batch size of 256 were configured to linearly scale up during the first 10 epochs and then decay with a cosine schedule. The weight decay followed the same schedule from 0.04 to 0.4. To provide different views as shown in Figure 2, the data augmentation design used in DINO was adopted. Specifically, the augmentation pipeline for creating global views included random resized cropping, random horizontal flipping, color jittering, Gaussian blur, solarization, and normalization. For local views, a multi-crop strategy was adopted to randomly crop the input image to half of its original size. It was determined experimentally that adding an appropriate number of local views helped to stabilize SSL training and improve performance.

[0076] VGH dataset:

[0077] The network was trained on the VGH (Vancouver General Hospital) dataset, which is a hematoxylin and eosin (H&E) breast cancer dataset constructed from the Netherlands Cancer Institute (NKI) cohort and the VGH cohort. Patches with tissue coverage less than 70% were filtered out. The patches were cropped from the original resolution of 1128×720 pixels to a smaller size of 224×224 pixels, with an overlap rate of up to 50%. The dataset was also augmented by applying transformations 240 involving rotations of 90 degrees and 180 degrees as well as vertical and horizontal flips. The post-processed dataset contained a total of ≈300,000 images.

[0078] BreastPathQ dataset:

[0079] BreastPathQ is a challenging dataset with noisy and fine-grained labels. For training / validation, a set of 2,579 / 187 patches were extracted from 96 H&E slides of residual invasive breast cancer from the TCGA-BRCA cohort at 20x magnification, measuring the tumor cell density, i.e., the occupancy fraction of the part of the image patch where tumor cells are present. Each patch was assigned a tumor cell density score on a continuous scale from 0 to 1. Therefore, the mean squared error (MSE) was reported in Table 1 using linear regression and Kendall-Tau concordance. The Kendall tau correlation coefficient is a non-parametric association measure based on the number of agreements and disagreements in paired observations.

[0080] PanNuke dataset:

[0081] The PanNuke dataset consists of semi-automatically generated nuclear instance images with exhaustive nuclear labels across 19 different tissue types. These exhaustive nuclear labels are sampled from 20,000 whole-slide images at different magnifications from multiple data sources. The dataset includes a total of 205,343 labeled nuclei, each with an instance segmentation mask. However, these labels are for experimental purposes only.

[0082] CRC dataset - downstream task (patch-level tissue phenotype)

[0083] As the training set in downstream task 150, the CRC (colorectal cancer) includes 100,000 hematoxylin and eosin (H&E) stained 224×224 tissue patches of human colorectal cancer (CRC) and normal tissues manually extracted from 86 slides at 20x magnification, with and without Macenko normalization. Each image is annotated with tissue type labels (Adipose (Adi), Background (Back), Debris (Deb), Lymphocyte (Lym), Mucus (Muc), Smooth Muscle (Mus), Normal Colon Mucosa (Norm), Stroma Associated with Cancer (Str), Colorectal Adenocarcinoma Epithelium (Tum)).

[0084] For evaluation, standard protocols are used on the four (above) datasets by using frozen features or fine-tuning the features. The k-nearest neighbor (k-NN) classifier and linear classifier (linear probing) are trained on frozen features extracted from a pre-trained SSL backbone network by sweeping different numbers of nearest neighbors for KNN and different learning rates for linear probing. Additionally, the network is initialized with pre-trained weights to conduct semi-supervised experiments using different percentages of labeled images evenly distributed to each class in the training set, while the test data remains the same as the official split.

[0085] Table 1 reports the results of KNN accuracy and linear detection accuracy for block-level tissue type classification evaluated on CRC datasets with and without normalization and PanNuke datasets. In addition, Table 1 also shows the mean square error (MSE) and Kendall-Tau consistency scores on the BreastPathQ dataset. For the DINO Cypher ViT model, the best results among different hyperparameter settings (as shown in Table 4) are reported. All SSL methods use the vanilla ViT backbone network, except DINO CypherViT, which uses the proposed CypherViT backbone network. The arrows next to the labels in all the following tables refer to the indication of the direction in which the value of the corresponding indicator should be, for example, an upward arrow (↑) indicates that the more the better, and a downward arrow (↓) indicates that the less the better. The results in the tables of the present disclosure are bold for highlighting purposes. In Table 1, the Acc@1 indicator (also known as the 1st place) may indicate the percentage of instances where the correct label is the highest prediction made by the model. Similarly, Acc@3 can indicate whether the correct label is among the top three predictions made by the model. It is a relatively looser measure than Acc@1.

[0086] Table 1. Accuracy of KNN and linear probing for block-level tissue type classification evaluated on the CRC dataset (with and without normalization), PanNuke, and BreastPathQ datasets.

[0087]

[0088] In Table 2, the performance variation of the disclosed Cypher ViT model with the DINO framework is examined, which is trained using various percentages (e.g., 1%, 5%, 10%, 20%, 50%, and 100%) of the labeled CRC dataset with Macenko normalization. For fair comparison, other existing state-of-the-art self-supervised methods trained on the VGH dataset are also implemented using the same architecture (ViT-small) with the default hyperparameter settings in the officially released codebase. All SSL methods use the Vanilla ViT backbone network, except DINO-CypherViT which uses the proposed CypherViT backbone network. Adding 5% labeled data can provide results comparable to the performance of using the entire training dataset, exceeding the performance of using the entire training dataset when increasing the labeled data to 10%. It suggests that promising SSL applications can achieve comparable or even better performance by training on less data. Evaluation results and ablation studies are presented in Figure 8A , Figure 8B , Figure 10A and Figure 10B Shown in.

[0089] Table 2. Semi-supervised learning accuracy of patch-level tissue type classification evaluated using different percentages of labeled CRC datasets (with normalization).

[0090]

[0091] Table 3 shows the top-1 KNN classification accuracy of the disclosed DINO framework based on CypherViT 405 in the case of 9 tissue types for downstream tasks using labeled CRC datasets with normalization. For compact illustration, the full tissue names have been replaced with the following abbreviations: Adipose (Adi), Background (Back), Debris (Deb), Lymphocyte (Lym), Mucus (Muc), Smooth Muscle (Mus), Normal Colonic Mucosa (Norm), Cancer-Associated Stroma (Str), Colorectal Adenocarcinoma Epithelium (Tum).

[0092] Table 3. Top-1 KNN accuracy for 9 tissue types obtained from the CRC dataset with normalization. (Note: Adipose (Adi), Background (Back), Debris (Deb), Lymphocyte (Lym), Mucus (Muc), Smooth Muscle (Mus), Normal Colonic Mucosa (Norm), Cancer-Associated Stroma (Str), Colorectal Adenocarcinoma Epithelium (Tum))

[0093]

[0094] Figure 8A Shows 2D UMAP (Uniform Manifold Approximation and Projection) visualizations of the feature embeddings extracted from an example implementation of a model pre-trained on ImageNet, the existing state-of-the-art self-supervised learning (SSL) framework MoCo (Momentum Contrast), and DINO pre-trained on the VGH dataset. Figure 8B Shows 2D UMAP visualizations of the feature embeddings extracted from an example implementation of the state-of-the-art SSL framework iBOT (Image BERT Training using an Online Tokenizer) and the present disclosure pre-trained on the VGH dataset. UMAP is a dimensionality reduction technique known for preserving global and local structures in the data.

[0095] The existing state-of-the-art SSL models (MoCo, DINO, iBOT) and the models in the embodiments according to the present disclosure are pre-trained on the VGH dataset. The default parameter settings for UMAP plotting are kept as (neighbors = 15, distance = 0.1). To study the results of phase 3, interpolation is used to estimate the extracted features And the attention map is visualized by overlaying it on the original image, as shown in Figure 8. The feature embedding is obtained by adopting an SSL backbone encoder on the normalized CRC dataset for downstream tasks. For loss objective calculation, an average pooling layer is applied to the aggregated output from the semantic clustering block 520 at the last stage. From Figure 8A and 8B it can be observed that, compared with other state-of-the-art SSL methods, the embeddings from DINO using the proposed CypherViT backbone network present clearer and less noisy clustering results. The DINO framework based on CypherViT 405 performs better in terms of concentration, intra-cluster and inter-cluster. Interestingly, the post-clustering features extracted from the learnable multi-[cls] token 330 exhibit interpretable attention concentration regions corresponding to the morphological tissue phenotypes of histopathological patches.

[0096] Figure 9 Figure 8 shows the retrieval results of a set of query images according to an example implementation of the present disclosure. The top five retrieved images of each query image are shown for 9 tissue types from the CRC dataset in each row. A database of embeddings that are distinguishably distributed in the embedding space for each patch is created. When a query image is selected, its embedding is compared with the embeddings in the database based on cosine distance. After this process, the L most similar patches are returned, where L can be customized.

[0097] In Figure 10, attention maps of six tissue types from the normalized CRC dataset are obtained from an example implementation of the disclosed Cypher ViT405 integrated as a backbone network within an SSL-based DINO framework. In Figure 10A and 10B the first column, the best results regarding relatively high attention scores representing the precise recognition of pixel-level morphological phenotypes are highlighted.

[0098] Figure 10A Figure 19 shows the attention maps of three types of predicted classification tokens (i.e., fat, normal colonic mucosa, and colorectal adenocarcinoma epithelium) extracted from the learnable multi-class token 330 at the last stage of an example implementation of the present disclosure. Figure 10B Figure 21 shows the attention maps of three other predicted classification tokens (i.e., lymphocytes, smooth muscle, cancer-associated stroma) extracted from the learnable multi-class token at the last stage of an example implementation of the present disclosure. Figure 10A and 10B depict the precise recognition of pixel-level morphological phenotypes, exceeding the grid-structured attention maps obtained from the multi-head attention 510 in CypherViT 405. It can also be observed that the semantic clustering block can learn semantically interpretable features corresponding to different morphological phenotypes.

[0099] Semantically meaningful attention maps from multi-[cls] tokens indicate promising interpretability. To provide a more intuitive visualization of what CypherViT 405 has learned from SSL training, the original image with the attention map is overlaid after interpolating the attention weights extracted from each [cls] token at the final stage of the semantic clustering block 520. It can be observed in Figure 10A and 10B that the regions of interest highlighted from the learned attention indicate morphological phenotypes, such as the cells in the first column, the white space in the second column, and the stromal tissue in the last two columns in each of the classes shown. Compared with state-of-the-art self-supervised ViT models such as DINO, Cypher ViT has two advantages regarding interpretability. First, compared with previous SSL models, the attention maps present finer-grained and more precise details, which provide attention maps in a grid structure restricted within regular shapes; second, the disclosed design has greater potential for development to adapt to other machine learning settings in future work. Specifically, similar to unsupervised clustering, although what each [cls] token actually learns is unknown, it can be easily adjusted in a weakly supervised experimental environment by slightly modifying the loss of each [cls] token to control what each [cls] token should focus on based on the assigned content.

[0100] Table 4 presents the results of an ablation study investigating two SSL series: generative SSL (including MAE) and contrastive SSL (including SimCLR, MoCo, and Dino) using the same SSL framework. The SSL framework is implemented under two different backbones: Vanilla ViT and the disclosed CypherViT. Table 4 reports the KNN accuracy of patch-level tissue type classification evaluated on the normalized CRC dataset and the PanNuke dataset. In addition, it includes the MSE and Kendall-Tau concordance scores for the BreastPathQ dataset. As observed, it is clear that the proposed CypherViT 405 is a plug-and-play solution that can be seamlessly integrated into different SSL frameworks. It demonstrates excellent performance and is model-agnostic, enabling it to adapt to various scenarios.

[0101] Table 4. Comparative table investigating two SSL series (generative and contrastive) using the same SSL framework under two different backbones: Vanilla ViT and the proposed Cypher ViT.

[0102]

[0103] In addition, the disclosed CypherViT 405 is an SSL-framework agnostic backbone network that is seamlessly integrated into different contrast-based SSL pipelines such as DINO, MoCo, and SimCLR. Comprehensive experiments have been conducted to demonstrate its consistent performance improvement compared to the vanilla ViT backbone network, as shown in Table 4. Its "plug-and-play" feature allows for easy and efficient framework adaptation without significant architectural modifications.

[0104] Depending on the task, the function may be more robust with an appropriate number of local views obtained using self multi-crop augmentation. To study the contributions of the key hyperparameters used in the main architecture of CypherViT 405 and the important techniques for stable training, an ablation study in Table 5 was conducted, testing the performance variations with respect to different combinations of the number of [cls] tokens and the number of local views at the last stage. It can be inferred that for tasks that require more fine-grained levels of the learned features (such as calculating the MSE error between the linear regression results and the annotations from BreastPathQ measuring tumor cell density), more local views are beneficial. While at the coarse-grained level or for global features used in classification tasks, fewer local views are preferred. For the experiments conducted in Table 5, using four [cls] tokens at the last stage and two local views from the multi-crop augmentation seems to be the optimal hyperparameter combination.

[0105] Table 5. Ablation study of the two components of the proposed DINO-CypherViT, namely the number of [cls] tokens for feature clustering used at the last stage, and the number of local views used in the multi-crop augmentation strategy.

[0106]

[0107] Some embodiments of the present disclosure include a system that includes one or more data processors. In some embodiments, the system includes a non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform part or all of one or more methods disclosed herein and / or part or all of one or more processes disclosed herein. Some embodiments of the present disclosure include a computer program product tangibly embodied in a non-transitory machine-readable storage medium, which includes instructions configured to cause one or more data processors to perform part or all of one or more methods disclosed herein and / or part or all of one or more processes disclosed herein.

[0108] The terms and expressions employed are used in a descriptive rather than a restrictive sense, and in using such terms and expressions there is no intention of excluding any equivalents of the features shown and described or portions thereof, it being recognized, however, that various modifications are possible within the scope of the invention as claimed. Accordingly, it should be understood that although the present invention has been specifically disclosed by way of embodiments and optional features, modifications and variations of the concepts herein disclosed may be resorted to by those skilled in the art, and that such modifications and variations are considered to be within the scope of the invention as defined by the appended claims.

[0109] This description merely provides preferred exemplary embodiments and is not intended to limit the scope, applicability, or configuration of the disclosure. On the contrary, this description of the preferred exemplary embodiments will provide those skilled in the art with a viable description for implementing various embodiments. It should be understood that various changes may be made in the function and arrangement of elements without departing from the spirit and scope set forth in the appended claims.

[0110] Specific details are given in this description to provide a thorough understanding of the embodiments. However, it should be understood that the embodiments may be practiced without these specific details. For example, circuits, systems, networks, processes, and other components may be shown in block diagram form as components so as not to obscure the embodiments in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail to avoid obscuring the embodiments.

Claims

1. A computer-implemented method, comprising: Receiving a digital pathology image depicting a tissue section, the tissue section being stained with one or more histological dyes; Generating, by processing the digital pathology image using a machine learning model, a plurality of predicted classification results including at least a portion of the digital pathology image, wherein the machine learning model includes a self-supervised hierarchical vision Transformer (ViT) configured to perform unsupervised clustering with a plurality of classification labels; and Outputting the results.

2. The computer-implemented method according to claim 1, wherein the plurality of predicted classifications includes tissue type.

3. The computer-implemented method according to claim 1 or 2, wherein the plurality of predicted classifications includes magnification level.

4. The computer-implemented method according to claims 1 to 3, wherein the plurality of predicted classifications includes diagnostic categories characterizing histological features.

5. The computer-implemented method according to claim 4, wherein the diagnostic categories include non-diagnostic category, malignant negative category, atypical category, tumor: benign category, suspected category, or malignant positive category.

6. The computer-implemented method according to claims 1 to 5, wherein the plurality of predicted classifications includes categories for predicting whether the digital pathology image depicts specific histological features of a malignant tumor; or the degree to which the digital pathology image depicts the specific histological features of a malignant tumor.

7. The computer-implemented method according to claim 6, wherein the specific histological features of the malignant tumor comprise: High cell density, cell enlargement, lack of cell adhesion, high nuclear-to-cytoplasmic ratio, hyperchromatic nuclei, prominent nucleoli, large nucleoli, abnormal nuclear-chromatin distribution, high mitotic activity, abnormal nuclear membrane, cell pleomorphism, nuclear pleomorphism, or tumor diathesis.

8. The computer-implemented method according to any one of claims 1 to 7, wherein for each part in a set of parts of the digital pathology image, the plurality of predicted classifications includes a classification characterizing what is being depicted within the part.

9. The computer-implemented method according to claim 8, wherein each part in the set of parts is a patch.

10. The computer-implemented method according to claim 8, wherein each part in the set of parts is a pixel.

11. The computer-implemented method according to any one of claims 1 to 10, wherein the self-supervised hierarchical vision Transformer (ViT) includes a multi-head self-attention module configured to: For each pair in a plurality of pairs of a set of patches in the digital pathology image, use an attention mechanism to predict a cross-patch correlation metric; and Based on the cross-patch correlation metric, assign each patch in one or more patches in the set of patches to a cluster.

12. A system, comprising: One or more data processors; And A non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform operations including the following: Receive a digital pathology image depicting a tissue section that is stained with one or more histological dyes; Generate, by processing the digital pathology image using a machine learning model, a plurality of predicted classification results including at least a portion of the digital pathology image, wherein the machine learning model includes a self-supervised hierarchical vision Transformer (ViT) configured to perform unsupervised clustering with a plurality of classification tokens; and Output the results.

13. A computer program product tangibly embodied in a non-transitory machine-readable storage medium, the computer program product including instructions configured to cause one or more data processors to perform operations including the following: Receive a digital pathology image depicting a tissue section that is stained with one or more histological dyes; Generate, by processing the digital pathology image using a machine learning model, a plurality of predicted classification results including at least a portion of the digital pathology image, wherein the machine learning model includes a self-supervised hierarchical vision Transformer (ViT) configured to perform unsupervised clustering with a plurality of classification tokens; and Output the results.