Systems and methods for multiple instance learning for whole slide classification based on diverse global representation

DGR-MIL addresses computational challenges and instance diversity in whole slide image classification by employing learnable global vectors and diversity loss, enhancing diagnostic accuracy and efficiency in pathology.

WO2026035765A1PCT designated stage Publication Date: 2026-02-12THE ARIZONA BOARD OF REGENTS ON BEHALF OF THE UNIV OF ARIZONA +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/040769
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-06
Filing Date
2025-08-05
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

Conventional deep learning methods for whole slide image classification face computational intractability due to the gigapixel resolution of histological images, and existing multiple instance learning (MIL) methods overlook instance diversity, leading to inefficiencies in pathology diagnosis.

Method used

A novel multiple instance learning (MIL) aggregation method, DGR-MIL, models instance diversity through learnable global vectors using a cross-attention mechanism and diversity loss, incorporating positive instance alignment and a determinantal point process (DPP) to enhance diagnostic accuracy.

Benefits of technology

DGR-MIL significantly reduces computational load and improves prediction accuracy for whole slide image classification, outperforming state-of-the-art models on datasets like CAMELYON16 and TCGA-lung cancer.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025040769_12022026_PF_FP_ABST
    Figure US2025040769_12022026_PF_FP_ABST
Patent Text Reader

Abstract

Methods and systems for computerized whole slide image classification are disclosed. Exemplary methods include utilizing multiple instance learning (MIL) including storing, by a processor and in a database, a bag of instances X = {x 1, x 2, …, x n ,} denoting a whole slide image (WSI) with n tiled patches, projecting, by the processor and using a pre-trained feature extractor, each instance into an L-dimensional vector, pooling, by the processor, instance-level embeddings into a bag-level feature, and generating, by the processor and using the bag-level feature as input, a bag-level prediction as output.
Need to check novelty before this filing date? Find Prior Art

Description

60980.15316 TITLE: Systems and Methods for Multiple Instance Learning For Whole Slide Classification Based on Diverse Global Representation INVENTORS: Wenhui Zhu Peijie Qiu Yalin Wang Aristeidis Sotiras Xiwen Chen Abolfazl Razi CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to and the benefit of U.S. Provisional Application No. 63 / 679,813 filed August 6, 2024, entitled “SYSTEMS AND METHODS FOR MULTIPLE INSTANCE LEARNING FOR WHOLE SLIDE CLASSIFICATION BASED ON DIVERSE GLOBAL REPRESENTATION.” The foregoing application is hereby incorporated by reference in its entirety, including but not limited to those portions that specifically appear hereinafter, but except for any subject matter disclaimers or disavowals, and except to the extent that the incorporated material is inconsistent with the express disclosure herein, in which case the language in this disclosure shall control. TECHNICAL FIELD

[0002] The present disclosure relates to imaging, and in particular to techniques for machine learning in connection with pathology diagnosis. BACKGROUND

[0003] Prior image processing systems and methods have suffered from various deficiencies. Histological whole slide images (WSIs) are instrumental in identifying and classifying cancerous tissues of a variety of cancers, e.g., breast cancer, lung cancer, etc. Artificial intelligence (AI)-aided analysis methods have become increasingly important in pathology for several reasons, such as enhanced diagnostic precision and high throughput analysis. However, WSIs are characterized by their gigapixel resolution, which complicates the application of conventional deep learning methods due to computational intractability. Accordingly, improved image processing systems and methods remain desirable. SUMMARY

[0004] Various embodiments of the present disclosure relate to methods of whole slide image classification. While the ways in which various embodiments of the present disclosure60980.15316 address drawbacks of prior methods and systems are discussed in more detail below, in general, exemplary embodiments of the disclosure provide improved methods and systems for whole slide image classification using multiple instance learning. Exemplary methods provide desired reduction in computational intractability and desired increased prediction accuracy.

[0005] In accordance with various embodiments of the disclosure, methods of computerized whole slide image classification utilizing multiple instance learning are provided. Exemplary methods can include storing, by a processor and in a database, a bag of instances X = {x1, x2, …, xn,} denoting a whole slide image (WSI) with n tiled patches, projecting, by the processor and using a pre-trained feature extractor, each instance into an L- dimensional vector, pooling, by the processor, instance-level embeddings into a bag-level feature, and generating, by the processor and using the bag-level feature as input, a bag-level prediction as an output.

[0006] In various embodiments, the bag-level prediction may be either positive or negativefor the presence of a lesion. The pooling may be according to σ^x^୧^ =^^୮^^^ ^^ୟ୬୦^^^^^^^ ^ ^୧^୫^^^^^^^∑^^సభ ^^୮^^^ ^^ୟ୬୦^^^^^^^ ^ ^୧^୫^^^^^^^ where W, V, and U may be learnable parameters. The…x^୬^^൯, x^୧ ∈ R^ with ^x^^, … x^୬^ =, … where fcls can include a bag-level classifier function, and where fproj can include the pre-trained feature extractor. The pooling may utilize a diverse global representation. The diverse global representation of targetpositive instances can include a set of learnable vectors given by G = ^g^ ^^, … g୬^ ∈ R^ൈ^ withg୩ ∈ R ^, where K can be the number of global vectors. A feed-forward network can used toembed the instance vectors and the global vectors.

[0007] In various embodiments, exemplary methods can include modeling the diversity by comparing a similarity between each instance vector and the global vectors. The modeling canbe by a cross-attention mechanism according to head୦൫G, ^ X൯ = Attention^Q୦, K୦, V୦^ andQ୦ = GW୦^ , K୦ = X^W^୦ , V୦ = X^W^୦, where G can serve as queries, where the bag of instancepairs, where ^^ொ^ ,^^^^ , ^^^^ ∈ ℝ can be learnableparameters, and where h can be the number ofAttention^Q୦,K , V ^ = ^୦ ୦ softmax൫Q୦K୦∕^d୩൯ V୦ , where dk can be the dimensionality of the vectors. Theoutput of the cross-attention mechanism (MHCA) can be according to MHCA൫G,^ X൯ = concat^head^; … head୦^W^, where ^^ை ∈ ℝ^×^ can be a trainable parameter.Exemplary methods can utilize a tokenized global vector as a summary of all the global vectors.60980.15316

[0008] In various embodiments, exemplary methods can include computing a yielded ^ ore of each instance according to σ x^^ ^ ^^importance sc ^୧^ = softmax൭^ ౪^ౡ^^ ^్^൫ ^^^ే൯^^ౡ^, wherecorresponds to the bag-level prediction, triplet loss, and diversity loss.

[0009] In various embodiments, triplet loss can be calculated according to L^୰୧ =∑^୩ୀ^ [dା^G୩, x^(୮୭^)ୡ ^ − dି^G୩, x^(୬^^)ୡ ^ + μ] , where µ can be the margin parameter, d can be^^^ = m^^^ ( ) 1^ ^ + 1 − ^^ ^ ^^^หℐ ห ^^^^^^^^where m can be theof positive bags, and ℐnegcan be an index set of negative bags.

[0010] In various embodiments, exemplary methods can include utilizing diversity loss to prevent a solution where all global vectors are identical. Diversity loss can be computedaccording to ℒௗ^௩ = − log det (^^^^் + ^^^^), where I can be the identity matrix and where ^^ canbe a pre-determined value. ^^ can be about 1x10-10. In various embodiments, ℒ^^^^^ = ℒ^^ +⋋௧^^ ℒ௧^^ +⋋ௗ^௩ ℒௗ^௩, where ⋋௧^^ and ⋋ௗ^௩ can be pre-determined balance parameters.

[0011] In accordance with further exemplary embodiments of the disclosure, a system for whole slide image classification is provided. The system may be used for performing a method disclosed herein. The system can include one or more processors and one or more non- transitory memories storing computing instructions configured to communicate with the one or more processors and cause the one or more processors to perform multiple instance learning comprising storing, by a processor and in a database, a bag of instances X = {x1, x2, …, xn,} denoting a whole slide image (WSI) with n tiled patches, projecting, by the processor and using a pre-trained feature extractor, each instance into an L-dimensional vector, pooling, by the processor, instance-level embeddings into a bag-level feature, and generating, by the processor and using the bag-level feature as input, a bag-level prediction as output.

[0012] These and other embodiments will become readily apparent to those skilled in the art from the following detailed description of certain embodiments having reference to the attached figures; the invention not being limited to any particular embodiment(s) disclosed.60980.15316 BRIEF DESCRIPTION OF THE DRAWINGS

[0013] With reference to the following description and accompanying drawings:

[0014] FIG. 1 illustrates a method for whole slide image classification using multiple instance learning, in accordance with exemplary embodiments;

[0015] FIG. 2A illustrates examples of positive instances of with-bag and between-bag diversities measured by rate-distortion theory;

[0016] FIG.2B illustrates an exemplary histogram of the diversity measure within positive bags on the CAMELYON16 dataset;

[0017] FIG. 2C illustrates exemplary between-bag distinction, showing the pair-wise similarity between bags;

[0018] FIG. 3 illustrates an exemplary MIL aggregation model termed DGR-MIL, in accordance with an exemplary embodiment;

[0019] FIG.4A illustrates an exemplary similarity matrix for the global vectors G learned from the CAMELYON16 dataset when G is orthogonal, in accordance with an exemplary embodiment;

[0020] FIG.4B illustrates an exemplary similarity matrix for the global vectors G learned from the CAMELYON16 dataset when G is non-orthogonal, in accordance with an exemplary embodiment;

[0021] FIG.5A illustrates exemplary ablation studies on a number of non-tokenized global vectors on both CAMELYON16 and TCGA-NSCLC datasets, in accordance with an exemplary embodiment;

[0022] FIG. 5B illustrates accuracy, F1 score, and AUC compared to a set balance parameter λtri on CAMELYON16 dataset, in accordance with an exemplary embodiment;

[0023] FIG. 5C illustrates accuracy, F1 score, and AUC compared to a set balance parameter λdiv on CAMELYON16 dataset, in accordance with an exemplary embodiment;

[0024] FIG. 5D illustrates an exemplary comparison of the number of positive instances per bag, in accordance with an exemplary embodiment;

[0025] FIG. 6A illustrates a visualization of an exemplary attention map, showing raw WSI with the ground-truth annotation, in accordance with an exemplary embodiment;

[0026] FIG. 6B illustrates an exemplary attention map computed using the tokenized global vectors, in accordance with an exemplary embodiment; and

[0027] FIG.6C illustrates an exemplary attention map computed using a second of the (K- 1) global vectors when K = 6, in accordance with an exemplary embodiment60980.15316

[0028] FIG. 6D illustrates an exemplary attention map computed using a third of the (K- 1) global vectors when K = 6, in accordance with an exemplary embodiment

[0029] FIG.6E illustrates an exemplary attention map computed using a fourth of the (K- 1) global vectors when K = 6, in accordance with an exemplary embodiment

[0030] FIG.6F illustrates an exemplary attention map computed using a fifth of the (K-1) global vectors when K = 6, in accordance with an exemplary embodiment

[0031] FIG. 6G illustrates an exemplary attention map computed using a sixth of the (K- 1) global vectors when K = 6, in accordance with an exemplary embodiment. DETAILED DESCRIPTION

[0032] The following description is of various exemplary embodiments only, and is not intended to limit the scope, applicability or configuration of the present disclosure in any way. Rather, the following description is intended to provide a convenient illustration for implementing various embodiments including the best mode. As will become apparent, various changes may be made in the function and arrangement of the elements described in these embodiments without departing from principles of the present disclosure.

[0033] For the sake of brevity, conventional techniques and components for mathematical processes, transforms, image manipulation, and / or the like may not be described in detail herein. Furthermore, the connecting lines shown in various figures contained herein are intended to represent exemplary functional relationships and / or communicative couplings between various elements. It should be noted that many alternative or additional functional relationships or communicative connections may be present in exemplary methods and systems for imaging and / or components thereof.

[0034] Exemplary embodiments may be operative on and / or utilize computing resources of sufficient capability. For example, certain embodiments may utilize a personal computing workstation utilizing the Linux or Windows operating system. Other embodiments may utilize different operating systems, GPUs, or the like, or may be operative under virtual machines, across distributed or cloud computing resources, or the like.

[0035] In many embodiments, the techniques described herein can provide a practical application and several technological improvements. In some embodiments, the techniques described herein can provide for reduced computational load and increased prediction accuracy for lesions. These techniques described herein can provide a significant improvement over conventional approaches of whole slide image classification by multiple instance learning, such as those based upon a standard attention based multiple instance learning model, which treat60980.15316 each instance independently and do not take correlations between instances into account. Moreover, these methods are technical improvements over other possible approaches, such as assigning correlation values to instances and modeling the correlation, as such methods are prone to trapping the MIL model by incorrectly aggregating instances when making predictions. In many embodiments, the techniques described herein can beneficially make determinations based on dynamic information such as a present whole slide image and / or conditions that have occurred in previous whole slide images. In this way, the techniques described herein can avoid problems with stale and / or outdated machine learned models by continually updating.

[0036] In many embodiments, the techniques described herein can be used continuously at a scale that cannot be reasonably performed using manual techniques or the human mind. For example, the techniques include analyzing patterns of pixels within gigapixel (i.e., 1,000,000,000 pixel) images, comparing large sections of pixels to each other, learning statistical parameters for evaluating gigapixel images based on training on a large quantity of gigapixel images, performing extensive summations, trigonometrical functions, transposing and / or multiplying large matrices, and performing Laplace transformations.

[0037] In a number of embodiments, the techniques described herein can solve a technical problem that arises only within the realm of computers and computer networks, as multiple instance learning for whole slide image classification does not exist outside the realm of computers and computer networks.

[0038] Multiple instance learning (MIL) stands as a powerful approach in weakly supervised learning, employed in histological whole slide image (WSI) classification for detecting lesions. However, existing MIL mainstream methods focus on modeling correlation between instances while overlooking the inherent diversity among instances. Further, MIL methods aimed at diversity modeling face diversity limits due to computational constraints.

[0039] Disclosed herein is a novel multiple instance learning (MIL) aggregation method based on diverse global representation (DGR-MIL), which is designed for pathology diagnosis based on whole slide images (WSIs). Under the MIL setting, a WSI is treated as a bag, while all patches within an image are treated as instances of the bag. An exemplary novel method models diversity among instances through a set of learnable global vectors, which serve as a summary of diverse instances of interest (i.e., tumor instances in WSIs). First, we turn the instance correlation into the similarity between instance embeddings and the predefined global vectors through a cross-attention mechanism. Second, we present two mechanisms to enforce60980.15316 the diversity among the global vectors to be more descriptive of the entire bag: (i) positive instance alignment and (ii) a novel, efficient, and theoretically guaranteed diversification learning by utilizing a determinantal point process (DPP). The positive instance alignment module encourages the global vectors to align with instances of interest center (e.g., tumor WSI / bag). To further diversify the global representations, we propose a novel diversity loss for inclusion in the DPP. The proposed model outperforms the state-of-the-art MIL aggregation models by a substantial margin on the CAMELYON-16 and the TCGAlung cancer datasets.

[0040] In pathology diagnostics, many computer-aid methods frequently encounter challenges in analyzing whole slide images (WSIs) due to computational difficulties introduced by the inherent high resolution of WSIs. Hence, recent methods based on multiple instance learning (MIL) are commonly used in WSI analysis. However, existing multiple instance learning (MIL) methods primarily characterize correlations between instances within a bag, yet often overlook the critical variance among these instances. An exemplary innovative framework disclosed herein, DGR-MIL, navigates these obstacles by utilizing a novel mechanism to model instance diversity within WSIs, e.g. sub-type tumor detection, lymph node metastasis. This approach enhances the accuracy of diagnostic interpretations by ensuring a comprehensive analysis of tissue samples, thereby supporting pathologists in making informed decisions based on high-quality, detailed examinations.

[0041] It will be appreciated that the commercial potential of these concepts is significant, given the growing reliance on digital pathology and the need for more accurate and efficient diagnostic tools in healthcare. By improving the accuracy of WSI classification, DGR-MIL can lead to better patient outcomes through faster and more reliable diagnoses. Additionally, the ability of exemplary embodiments to handle the inherent diversity within WSIs makes it a valuable tool for research and development in pathology, potentially facilitating new insights into disease mechanisms and the development of targeted therapies. Additionally, principles of the present disclosure hold promise for widespread adoption in pathology labs and research institutions worldwide, offering substantial improvements in the diagnostic process and supporting the advancement of personalized medicine.

[0042] Diversity may be modeled by a set of learnable global vectors. The learned global vectors may serve as a summary of diverse instances of interest (i.e., tumor instances in WSIs). As a result, the diversity between instances can be implicitly modeled by computing the correlation between instance embeddings and the global vectors through a cross-attention mechanism. Tokenized global vectors may be implemented to enhance the ability of the global vectors to capture the most discriminative global context for WSI classification. The60980.15316 importance map for instances can be calculated based on the attention between the tokenized global vector and the embedding of each individual instance.

[0043] Standard transformers discover contextually relevant information by modeling the correlation between elements within a sequence through the self-attention mechanism. However, the traditional self-attention operation has quadratic time and space complexity O(n2), with respect to a sequence containing n elements. In the context of MIL, sequence length typically becomes quite large since one bag often approximately comprises ten thousand instances. This extremely long sequence poses significant computational intractability. In contrast, the cross-attention mechanism, which was originally proposed to relate positions from one sequence to another, allows models to consider cross-sequence information. Diversity between and among instances can be modeled through a cross-attention mechanism between instances and the proposed global vectors, described in more detail below. This approach can dramatically reduce the complexity compared to the self-attention mechanism since the number of global vectors is significantly less than the sequence length. Turning ahead to FIG. 2A, examples of positive instances of with-bag and between-bag diversities measured by rate- distortion theory are provided. With reference to FIG. 2B, an exemplary histogram of the diversity measure within positive bags on the CAMELYON16 dataset is provided. With reference to FIG. 2C, exemplary between-bag distinction, showing the pair-wise similarity between bags is provided.

[0044] The DGR-MIL model disclosed herein can comprise two distinct parts: i) the design of the global representation in MIL pooling, and ii) the strategy of learning diverse global representation, which can further include positive instance alignment and a computational- efficient diversity loss with a theoretical guarantee. With reference to FIG. 3, an exemplary framework for a DGR-MIL model is provided.

[0045] With reference to FIG.1, an exemplary method for MIL classification is provided. For example, in binary MIL classification the objective is to predict the bag-level label Y ∈ {0, 1}, given a bag of instances X = {x1, x2, …·, xn}, denoting a WSI with n tiled patches. The bag of instances X may be stored in a database 110. However, the corresponding instance-levellabels ^Y୧^ nare unknown in most WSI analyses due to the laboriousness of obtaining patch-levelThis turns the WSI classification into a weakly-supervised learning scheme according to the MIL formulation:

[0046] Y = ^0, iff ∑୧ y୧ = 01, otherwise.60980.15316

[0047] Because of the gigapixel resolution of WSIs, MIL typically is not performed in an end-to-end fashion and instead employs a simplified learning scheme. A simplified MIL learning process can comprise three main parts: i) a pre-trained feature extractor fproj(·) that can project each instance into a L-dimensional vector 120, ii) a MIL pooling operator σ(·) that can combine instance-level embeddings into a bag-level feature 130, and iii) a bag-level classifier fcls(·) that can take the bag-level feature as input and can produce the bag-levelprediction as an output 140. Mathematically, this process is given by Ŷ = fୡ୪^൫σ(^x^^, … x^୬^)൯,x^୧ ∈ R^with^x^^, … x^୬^= f୮୰୭୨(^x^^, … x^୬^), where Ŷ denotes the predicted bag-level label. In theattention-based MIL (AB-MIL) framework, the formulation for the MIL pooling operator canbe σ(x^) ^^୮^^^ (^ୟ୬୦(^^^^)) ^ ^୧^୫(^^^^)^୧= ∑^^సభ ^^୮^^^ (^ୟ୬୦(^^^^)) ^ ^୧^୫(^^^^)^, where W, V, and U are learnable parameters.

[0048] To accommodate the variability of the target lesions within and between bags, a diverse global representation may be developed in the MIL pooling stage. The global representation of the target (i.e., positive) instances may be defined as a set of learnable global vectors given by G = [g^^, … g^୬] ∈ R^×^with g୩∈ R^where K is the number of global vectors.A feed-forward network (FFN) may be used to embed the input instance vectors^X =^x^୧^୬୧ୀ^and / or the global vectors G. Within this disclosure, G ∈ R^×^may be used to denote global vectors for notation brevity.

[0049] The standard AB-MIL framework assumes the instances are independent and identically distributed while overlooking the correlation effect between instances. Hence, the self-attention mechanism becomes a natural choice for modeling the inter-instance correlation. However, due to the large number of instances within a bag in MIL, the quadratic time and space complexity O(n2) of standard self-attention poses a significant challenge in computation. Alternatively, the previous transformer-based MIL mitigates this problem by employing Nystrom-Attention, approximating the standard self-attention with linear complexity, which has proved effective of modeling correlation between positive and negative instances. However, self-attention usage only guarantees the general separation of the positive and negative instances in a bag, overlooking the diversity between instances and between bags.

[0050] The diversity between instances may be implicitly modeled by comparing the similarity between each instance vector and the proposed diverse global vectors. The comparison may be achieved via a cross-attention mechanism, where the global vector G may serve as queries, and a bag of instance vectors X̃ may be used as key-value pairs. The h-th head of the proposed cross attention is given by60980.15316head୦൫G, ^ X൯ = Attention(Q୦, K୦, V୦), Q୦ = GW୦^ , K୦ = ^ XW^୦ , V୦ = ^ XW^୦ , where ^^ொ^ ,^^^^ ,^^^ ∈ ℝ^×^⁄ ு^ can be H ofheads. For theas Attention(Q୦, K୦, V୦) = softmax൫Q୦K^୦∕^d୩൯V୦. The output of the yielding multi-headcross attention (MHCA) may be the concatenation of the outputs from all heads through a linearprojection MHCA൫G,^X൯ = concat(head^; … head୦)W^, where W^O∈ R^(L×L) can be atrainable parameter. The proposed cross-attention mechanism may reduce the quadratic timeand space complexity O(n2) in the standard self-attention mechanism to linear O(Kn) where K≪ n. In various embodiments, a Nystrom-Attention model may be applied to the instancevectors and global vectors before performing the cross-attention. Applying a Nystrom- Attention model to input instance vectors can facilitate filtering out the background, and applying self-attention to the global vectors can increase their discrepancies.

[0051] The transformer can include a class token, which can encode the globally discriminative representation associated with certain labels in image classification tasks. The class token can be added to the input token embedding by serving as a summary of the entire image. Further, a tokenized global vector gtokenmay be added as a summary of all the otherglobal vectors. The yielding global vectors can be denoted as G^ = ^g^୭୩^୬, g^, … g^^ ∈R(^ା^)×^. The output of the tokenized global vectors after applying the cross-layermay then be used for bag-level classification. The yielded importance score of each instance ^can be computed as σ(x^^^౪^ౡ^^^^్^൫^^^^^ే൯୧)= softmax൭ ^^ౡ ^.

[0052] In various embodiments, the proposed global vectors are learned in an unsupervised way, which poses a challenge in eliminating information from negative instances in the global vectors. This may be attributed to the similarity between positive instances and their adjacent negative instances, as tumor-adjacent regions often exhibit high-density, quantitative expression in the spatial relationships of cells. Each diverse global vector encapsulates a collection of analogous tissue features. As a result, certain global vectors emphasize certain types of positive instances. Accordingly, adding tokenized global vectors facilitates the model’s ability to capture the most discriminative global representation while suppressing the information from the negative instances.

[0053] Learning a reliable and diverse global representation in MIL may be accomplished by (i) positive instance alignment and (ii) diversity learning via utilizing the linear algebra property of the DPP.60980.15316

[0054] Positive instance alignment may be used to push global representation towards aligning with the instances of interest (i.e., positive instances). For example, global vectors may be pushed towards positive bag centers and / or may be pushed away from negative bag centers. First, a center of the positive bags and a center of the negative bags may be defined by^^^(^^^)^ ∈ ℝ^ and ^^^(^^^)^ ∈ ℝ^ respectively. The positive bag centers and the negative bagvia a momentum mechanism, for example upon each iteration of

[0055] ^^^(^^^)^ = m^^^(^^^)^ + (1 − ^^) ^หℐ^^ೞห ∑^∈ℐ^^ೞ ^^^^.which can be set empirically. For example, m may be set between 0 and 1, between 0.2 and 0.6, or at about 0.4. ℐ^^^and ℐ^^^can represent the index sets of positive bags and negative bags respectively. In this manner, the update of the positive instance center occurs only if a positive bag is fed into the network. The same strategy can be applied to the negative center update (i.e., updated only if a negative bag is encountered). A set of triplet ^G,^ x(୮୭^)ୡ , x^(୬^^)ୡ ^ can be formulated. The triplet loss can then be adopted to enforce the G being close to the positive bag centerwhile away from the negative bag center:

[0057] L (୮୭^) (୬^୰୧ = ∑^୩ୀ^ ^dା ^G୩, ^ xୡ ^ - d-^G୩, ^ x ^^)ୡ ^ + µ^ +,the distance measure. The distance measure can be determined by cosine similarity.

[0059] While positive instance alignment pushes the global representation to be aligned with the positive bag center, if used alone it may result in a trivial solution where all the global vectors are identical. A diverse global representation is desired to capture the variability of positive instances. Diversity loss inspired by DPP for data selection to maximize the diversity among global vectors and hence better summarize the instances, may be used to correct for such results. DPP is a diversification tool and may be used to select diverse subsets. Rather than use DPP for selection, exemplary embodiments utilize DPP as a differential diversity measurement.

[0060] P may be an L-ensemble DPP if the likelihood of an arbitrary subset A ⊆ S drawn from the entire set S satisfies PL(A) ∝ det(LA), where LAdenotes a submatrix of the similarity Gram matrix L indexed by A. In the case of prompting diversity of global vectors G = ^g^^, …60980.15316 g^^൧, the similarity matrix is given as L = GG^∈ R^×^, A = B = [K] may be set, each global vector g୧, i ∈ A may be treated as a data point, and the total number of subsets can be calculated as 2|ୗ|= 2^. The matrix L can be positive semi-definite.

[0061] Additionally, P^(A) ∝ det(L^) = Volଶ(^g୧^୧∈^) provides that a diverse subset ismore likely to span larger volumes. As the similarity between two data points (i.e., Lij:i≠j) increases, they will span fewer areas, hence decreasing the probabilities of sets containing both of them. Accordingly, feature vectors that are more orthogonal to each other may span the largest volumes, resulting in the most diverse subsets. With reference to FIG.4A an exemplary similarity matrix for the global vectors G learned from the CAMELYON16 dataset when G is orthogonal is provided. With reference to FIG. 4B, an exemplary similarity matrix for the global vectors G learned from the CAMELYON16 dataset when G is non-orthogonal is provided.

[0062] Given a set of global vectors G = ^g^^, … g^^൧ with ‖g୧‖ = C, ∀i ∈ [K], maximizingthe DPP-based diversity (i.e. maxdet(GG^)) results in orthogonal global vectors with g୧ ⊥g୨, ∀i ≠ j, i, j ∈ [K]. The determinant det(L) = det(GG⊤) is upper-bounded according toHadamard’s inequality:

[0063] |det(^^)| (^)ୀ det(^^) (^)ஸ ∏^^ୀ^ ^^^^

[0064] Condition (a) the matrix L is positive semi-definite. The equalityof Condition (b) is achieved if all non-diagonal entries of G are zeros, meaning rows of the global vectors are orthogonal. The normalization constraint leads the upper bound to be theinfimum, since ^^^^ = ‖^^‖ଶ^ ≤ ^^ଶ and it can be achieved if the equality of Condition (b) issatisfied.

[0065] According to Theorem 1, we propose a diversity loss ℒdiv to diversify the proposed global vectors by minimizing the negative logarithm of det(GG⊤):

[0066] ℒ^୧^ = − log det(GG^), s. t. ‖^^^‖ = 1 = C

[0067] Accordingly, optimal diversity through minimizing loss is theoretically achievable. Enforcing the constraints ‖^^^‖ = 1 may lead the infimum of ℒdivto reach zero due to log(GG⊤)ii= log(‖^^^‖2) = 0. In contrast, the diversity loss ℒdivcan be arbitrarily small (up to −∞) without the constraint ‖^^^‖ = 1, which can result in an unstable training.

[0068] A predetermined value can be added to prevent the logarithm of the determinant from being negative infinity (i.e. any two global vectors become collinear). The predetermined value can be ^^ = 1 x 10−10The final diversity loss can be given as:

[0069] ℒௗ^௩ = − log det(^^^^்) + ^^^^,60980.15316

[0070] where I can be the identity matrix. The complexity to compute the loss is approximate O(L), which is negligible.

[0071] The MIL model disclosed herein can be trained in an end-to-end fashion, for example by jointly optimizing the weighted combination of cross entropy () loss that may correspond to bag-level classification, triplet loss, and the proposed diversity loss:

[0072] ℒ^^^^^ = ℒ^^ +⋋௧^^ ℒ௧^^ +⋋ௗ^௩ ℒௗ^௩parameters. The balance parameters can bebetween 0 and 1, between 0.01 and 0.5, or about 0.1.

[0074] The disclosed DGR-MIL model was validated on the CAMELYON16 dataset and the TCGA-lung cancer dataset (TCGANSCLC). For the CAMELYON16, the training set is further divided into training and validation sets with a 9:1 ratio. The mean of accuracy, F1 score, and AUC with their corresponding 95% interval on the testing dataset after running five experiments are provided. For the TCGA lung cancer dataset, a 4-fold cross-validation experiments was performed, where the dataset was partitioned into training, validation, and testing sets with a patient ratio of 65:10:25. The mean and standard variation of accuracy, F1 score, and AUC on the testing dataset from 4-fold cross-validation are provided.

[0075] Three sets of instance features were extracted using differing strategies to evaluate the proposed method’s adaptability across various feature embeddings. The first set provided by DTFD-MIL, employing OTSU’s method for patch extraction from WSIs and ResNet-50 for feature extraction, resulting in 1024-dimensional vectors per patch. For validation, two additional sets of features were generated by segmenting each WSI into non-overlapping 224x224 patches using threshold filtering, resulting in 3.4 and 10.3 million patches from CAMELYON16 and TCGA lung cancer datasets, respectively. These patches were processed using ResNet-18 and Vision Transformer, pre-trained on ImageNet, to produce 512 and 768 dimensional feature vectors.

[0076] We compare the DGR-MIL model to eight alternative MIL methods. These models can be roughly divided into two categories: i) AB-MIL and its variants, including CLAM-SB, DS-MIL, and DTFD-MIL; ii) the transformer-based methods including Trans-MIL and ILRA- MIL; and iii) clustering / prototype-based MIL including PMIL. All the models are trained using the same parameter settings.

[0077] The DGR-MIL method disclosed herein outperforms the other alternative MIL models by a large margin in both the CAMELYON16 and TCGA-NSCLC datasets using features extracted by three different means (see Table 1). The DGR-MIL model outperforms the second-best models in terms of accuracy (1.7%; 1.3%), F1 score (3.1%; 1.5%), and AUC60980.15316 (1.1%; 1.7%) when using features extracted from ResNet-50 in CAMELYON16 and TCGA- NSCLC, respectively. A similar performance gain is observed on features extracted from ResNet-18 including accuracy (3.4%; 1.1%), F1 score (3.5%; 1.0%), and AUC (4.4%; 1.1%). An improvement in accuracy (3.4%; 1.1%), F1 score (3.5%; 1.0%), and AUC (4.4%; 1.1%) is shown when using features extracted from the vision transformer.

[0078] Table 1: CAMELYON16 TCGA-NSCLC Accuracy F1 AUC Accuracy F1 AUC ResNet-50 ImageNet Pretrained

[0079] The performance of the three sets of feature embeddings varied: the ViT feature embeddings outperform the ResNet-18 features but show lower performance compared to the ResNet-50 features. This is attributed to the fact that a greater number of positive instances is extracted by ResNet-50. With reference to FIG. 5D, an exemplary comparison of the number of positive instances per bag is provided. A smaller portion of positive instances in the extracted patches may accompany a drop in performance. This phenomenon benefits the pseudo-bag partitions in DTFD-MIL, as more positive instances within a bag are prone to result in less60980.15316 noisy pseudo-bag labels. This accounts for the drop in DTFD-MIL performance when applied to feature embeddings that contain a lower proportion of positive instances.

[0080] Table 2 P D CAMELYON16 TCGA-NSCLC Accuracy F1 AUC Accuracy F1 AUC x ✓ 0.895 0.887 0.922 0.872 0.875 0.928gtokenCAMELYON16 TCGA-NSCLCAccuracy F1 AUC Accuracy F1 AUC x 0.907 0.900 0.935 0.903 0.905 0.957 63ON16 dataset with features extracted by a ResNet-50. Different components of the proposed model were ablated, i.e., the positive instance alignment module and the diversity loss. Incorporating the proposed global vectors, without employing any of the learning strategies, yielded an AUC of 0.922 and 0.928. This AUC exceeds that of most existing MIL models. Subsequently, by including the proposed positive instance alignment module, a performance gain of (2.2%, 2.8%) in accuracy, (2.3%, 2.9%) in F1 score, and (2.2%, 2.8%) in AUC is provided. Up to now, DGR-MIL outperforms the DTFD-MIL in terms of accuracy and F1 score, and achieve a similar AUC (AUC = 0.944,0.956) compare to the DTFD-MIL(AFS) (AUC = 0.946,0.951). Further incorporating the proposed diversity loss into the objective function yields a performance gain of (1.3%,0.7%) in AUC, which outperforms DTFD-MIL (AFS) by (1.1%,1.2%).

[0083] As shown in Table 3, including the tokenized global vector gtokenyields a remarkable performance gain by improving accuracy by (1.0%, 0.5%), F1 score by (1.3%, 0.6%), and AUC by (2.2%, 0.6%). Consistent with the pathological findings that instances are diverse, different global vectors indeed corresponded to different instance representations, which can be depicted by the attention map produced by different global vectors in Fig. 5. However, we also observe that the learned global vectors still include non-tumor related representation, particularly around tumor boundaries, as positive instances around tumor boundaries have a similar appearance to surrounding negative instances. As a result, incorporating tokenized global vectors can mitigate this problem by capturing the most discriminative positive (tumor) regions. With reference to FIG.6A illustrates a visualization of an exemplary attention map is provided. With reference to FIG. 6B, an exemplary attention60980.15316 map computed using the tokenized global vectors is provided. With reference to FIGs.6C-6F, an exemplary attention map computed using the remaining global vectors of the (K-1) global vectors when K = 6 is provided.

[0084] The optimal number of global vectors K in different data sets may vary due to dataset intrinsic properties. The optimal K for the CAMELYON16 and TCGA-NSCLC dataset are K = 5 and K = 3, respectively. An overly large K may decrease performance as it may harden the learning task. However, any suitable and / or desired K may be utilized. With reference to FIG.5A, exemplary ablation studies on a number of non-tokenized global vectors on both CAMELYON16 and TCGA-NSCLC datasets are provided.

[0085] By conducting a grid search, we find that the optimal setting of the balance parameters is ⋋tri= 0.1 and ⋋div= 0.1. An overly small ℒ௧^^and ℒௗ^௩(e.g., 0.01) is likely to enforce inadequate constraints on the learned global representation by deviating it from learning meaningful information of instance of interest. Larger balance parameters (e.g., {0.5, 1.0}) may distract the model from the main classification task, leading to a drop in classification performance. However, any suitable and / or desired balance parameters may be utilized. With reference to FIG. 5B, accuracy, F1 score, and AUC compared to a set balance parameter ⋋tri on CAMELYON16 dataset are provided. With reference to FIG. 5C, accuracy, F1 score, and AUC compared to a set balance parameter ⋋div on CAMELYON16 dataset are provided.

[0086] While the principles of this disclosure have been shown in various embodiments, many modifications of structure, arrangements, proportions, the elements, materials and components, used in practice, which are particularly adapted for a specific environment and operating requirements may be used without departing from the principles and scope of this disclosure. These and other changes or modifications are intended to be included within the scope of the present disclosure.

[0087] The present disclosure has been described with reference to various embodiments. However, one of ordinary skill in the art appreciates that various modifications and changes can be made without departing from the scope of the present disclosure. Accordingly, the specification is to be regarded in an illustrative rather than a restrictive sense, and all such modifications are intended to be included within the scope of the present disclosure. Likewise, benefits, other advantages, and solutions to problems have been described above with regard to various embodiments. However, benefits, advantages, solutions to problems, and any element(s) that may cause any benefit, advantage, or solution to occur or become more pronounced are not to be construed as a critical, required, or essential feature or element.60980.15316

[0088] As used herein, the terms "comprises," "comprising," or any other variation thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but may include other elements not expressly listed or inherent to such process, method, article, or apparatus. Also, as used herein, the terms "coupled," "coupling," or any other variation thereof, are intended to cover a physical connection, an electrical connection, a magnetic connection, an optical connection, a communicative connection, a functional connection, and / or any other connection. When language similar to "at least one of A, B, or C" or "at least one of A, B, and C" is used in the specification or claims, the phrase is intended to mean any of the following: (1) at least one of A; (2) at least one of B; (3) at least one of C; (4) at least one of A and at least one of B; (5) at least one of B and at least one of C; (6) at least one of A and at least one of C; or (7) at least one of A, at least one of B, and at least one of C.

Claims

60980.15316 CLAIMS What is claimed is:

1. A computerized whole slide image classification method utilizing multiple instance learning (MIL), comprising: storing, by a processor and in a database, a bag of instances X = {x1, x2, …, xn,} denoting a whole slide image (WSI) with n tiled patches; projecting, by the processor and using a pre-trained feature extractor, each instance into an L-dimensional vector; pooling, by the processor, instance-level embeddings into a bag-level feature; and generating, by the processor and using the bag-level feature as input, a bag-level prediction as an output.

2. The method of claim 1, wherein the bag-level prediction is either positive or negative for a presence of a lesion.

3. The method of claim 1, wherein the pooling is according toexp^W^ (tanh(Vx^ )) ^ sigm(U )^σ(x^ ) = ୧ x^୧୧ ∑୬ exp^W^ (tanh(Vx^୧)) ^ sigm(Ux^୧)^where W, V, and U4. The method of claim 3, wherein the generating is according toŶ= f σ(^x ^) ^ୡ୪^൫ ^^, … x^୬ ൯, x^୧ ∈ Rwhere Yˆ denotes the bag-level prediction, where fclscomprises a bag-level classifier function, and where fproj is the pre-trained feature extractor.

5. The method of claim 3, wherein the pooling utilizes a diverse global representation.

6. The method of claim 5, wherein the diverse global representation comprises a set of learnable global vectors given by G = [g^^, … g^୬] ∈ R^×^with g୩∈ R^, wherein K is a number of global vectors.

7. The method of claim 6, wherein a feed-forward network is used to embed the instance vectors and the global vectors.60980.15316 8. The method of claim 7, further comprising modeling a diversity by comparing a similarity between each instance vector and the global vectors.

9. The method of claim 8, wherein the modeling is by a cross-attention mechanismaccording tohead୦൫G, X^൯ = Attention(Q୦, K୦, V୦)Q= GW ^ ^ ^ ^୦ ୦^ , K୦ = XW୦ , V୦ = XW୦where G serve as ^^^ vectors and serves as key-valuepairs, where ^^ொ^ ,^^^^ ,and where h is a number ofheads.

10. The method of claim 9, wherein Attention(Q୦, K୦, V୦) = softmax൫Q୦K^୦∕^d୩൯V୦,where dkis a dimensionality of the global vectors.

11. The method of claim 10, wherein the output of the cross-attention mechanism (MHCA)is according toMHCA൫G, ^ X൯ = concat(head^; … head୦)W^where ^^ை ∈12. The method of claim 11, wherein the method utilizes a tokenized global vector as a summary of all the global vectors.

13. The method of claim 12, further comprising computing a yielded importance score of each instance according to ^^σ(x^g W x^ W୧) = softmax^൫ ^୭୩^୬ ୦^ ൯൫ ୧ ୦൯^^d୩where gtoken is the tokenized global vector.

14. The method of claim 13, wherein the MIL is trained in an end-to-end fashion by jointly optimizing a weighted combination of cross-entropy (ce) loss that corresponds to the bag-level prediction, a triplet loss, and a diversity loss.60980.15316 15. The method of claim 14, wherein the triplet loss is calculated according to ^L (୮୭^) (୬^^)^୰୧ = ^[dା^G୩, x^ୡ ^ − dି^G୩, x^ୡ ^ + μ] +where µ is a^^^^ = m^^^^ + (1 − ^^) ^ ^^^หℐ ^^^^ห^^^where m is a bags, and ℐneg an index set of16. The method of claim 15, further comprising utilizing diversity loss to prevent a solution where all global vectors are identical.

17. The method of claim 16, wherein the diversity loss is computed according toℒௗ^௩ = − log det (^^^^் + ^^^^)where I is an identity matrix andvalue.

18. The method of claim 17, where ^^ is about 1x10-10.

19. The method of claim 18, whereinℒ^^^^^ = ℒ^^ +⋋௧^^ ℒ௧^^ +⋋ௗ^௩ ℒௗ^௩where ⋋௧^^and ⋋ௗ^௩are pre-determined balance parameters.

20. A system for whole slide image classification, comprising: one or more processors; and one or more non-transitory memories storing computing instructions configured to communicate with the one or more processors and cause the one or more processors to perform multiple instance learning comprising: storing, by a processor and in a database, a bag of instances X = {x1, x2, …, xn,} denoting a whole slide image (WSI) with n tiled patches;60980.15316 projecting, by the processor and using a pre-trained feature extractor, each instance into an L-dimensional vector; pooling, by the processor, instance-level embeddings into a bag-level feature; and generating, by the processor and using the bag-level feature as input, a bag-level prediction as output.

Citation Information

Patent Citations

  • Critical component detection using deep learning and attention

    US20230245431A1

  • Method of processing an image of tissue and a system for processing an image of tissue

    US20230377155A1

  • Systems and methods for identification of pancreatic ductal adenocarcinoma molecular subtypes

    US20240221159A1