Machine-learning models for tumor grading using rank-aware contextual reasoning on whole slide images
The RACR model addresses the challenges of subjective and variable cSCC tumor grading by using patch-based contextual reasoning and rank-ordering, enhancing localization and classification accuracy for high-risk tumors.
Patent Information
- Application Number
- PCT/US2024/062101
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-29
- Filing Date
- 2024-12-27
- Publication Date
- 2025-07-03
AI Technical Summary
Current methods for grading cutaneous squamous cell carcinoma (cSCC) tumors are subjective, prone to inter-observer variability, and face challenges such as varying tumor regions within the same image, the need for contextual information, and limited data, particularly with an imbalance in low-risk vs. high-risk cases.
A rank-aware contextual reasoning (RACR) model that partitions tumor images into patches, uses self-supervised pre-trained encoders for local features, and incorporates spatial and semantic dependencies through graph convolution, with attention-based aggregation and rank-ordering losses to prioritize severe tumor regions.
The RACR model improves tumor localization and classification accuracy, especially for high-risk grades, by consistently focusing on the most severe tumor regions and leveraging contextual information, reducing inter-observer variability and data limitations.
Smart Images

Figure US2024062101_03072025_PF_FP_ABST
Abstract
Description
Attorney Docket No.07039-2297WO1 Machine-Learning Models for Tumor Grading Using Rank-Aware Contextual Reasoning on Whole Slide Images CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of U.S. Provisional Application Serial No.63 / 616,287, filed December 29, 2023. The disclosure of the prior application is considered part of the disclosure of this application and is incorporated by reference in its entirety into the present application. BACKGROUND
[0002] This specification describes techniques for training and using machine-learningmodels to estimate the grade of cancer of tissue depicted in one or more whole-slide images, including techniques for implementing weakly supervised deep-learning models to grade cutaneous squamous cell carcinoma (cSCC).
[0003] cSCC is the second most prevalent skin cancer in the United States, and itsoccurrence is increasing rapidly. cSCC tumor grade is an important prognostic factor, reflecting the level of cancer aggressiveness, and is strongly linked to outcomes (Thompson et al.2016). The current practice for grading cSCC tumors involves a manual examination of whole slide images (WSI) of skin tissues by pathologists, which is inherently subjective, prone to inter- observer variability, and leads to under-staging of high-risk cSCC tumors (Severson et al.2020). AI-assisted grading has emerged as a promising approach for objective tumor grading.
[0004] Neural networks are machine learning models that employ one or more layers ofnonlinear units to predict an output for a received input. Some neural networks include one or more hidden layers in addition to an output layer. The output of each hidden layer is used as input to the next layer in the network, i.e., the next hidden layer or the output layer. Each layer of the network generates an output from a received input in accordance with current values of a respective set of parameters. SUMMARY
[0005] This specification describes systems, methods, devices, and correspondingtechniques for using a weakly-supervised machine model to predict cSCC grade, where the model is trained on WSI-level grade labels. In some implementations, the model is trained toAttorney Docket No.07039-2297WO1 classify cSCC WSI into one of four grading classes: normal (tumor not present), well- differentiated, moderately differentiated, and poorly-differentiated. The model can be implemented according to a multiple-instance learning (MIL) paradigm by partitioning each WSI into a bag of tiled patches (instances).
[0006] cSCC tumor grading presents three notable challenges: (i) grade difference oftumor regions within the same WSI, (ii) the need for contextual information for determining tumor grade, and (iii) limited data. (i) A given cSCC WSI might comprise multiple tumor regions with varying grades. Pathologists implicitly rank the tumor regions based on their cellular differentiation (from well to poor), providing the grade of the most severe tumor region as the overall label. This specification that proposes that a machine-learning based model to predict cSCC should account for implicit grade order to identify the most severe tumor region and deemphasize the importance of irrelevant tumor regions, even if the less severe tumor captures a larger portion of the WSI. (ii) cSCC grading is context-aware because pathologists consider the local tumor neighborhood (tumor microenvironment) as well as long-range relations between distant tumor regions to determine the WSI-level grade label. The disclosed machine- learning based model likewise leverages local and long-range in a balanced fashion. (iii) Limited number of WSIs and an imbalance in the number of low-risk (well-differentiated) vs. high-risk (moderately and poorly-differentiated) cases further exacerbate the previous two challenges.
[0007] This specification refers to the present machine-learning approach as rank-awarecontextual reasoning (RACR) since it leverages contextual information and maintains the ordinal ranking of tumor grades. The disclosed model, RACR-MIL, predicts WSI tumor grade by dividing the tissue into patches. Local patch features are extracted using a self-supervised pre- trained encoder. A graph is defined on the tissue patches that captures both local (spatial) and non-local (semantic) dependencies between the patches. Multiscale features are derived using self-attention-based graph convolution to incorporate contextual information. Attention-based patch feature aggregation is utilized to derive a WSI-level feature for grade classification. The attention computation can be augmented with a rank ordering mechanism that assigns higher weights to higher-grade tumor regions. Tumor depth can also be used as an auxiliary prediction task for regularizing attention to relevant tumor patches.
[0008] The RACR-MIL model described herein can achieve various technical benefits.First, to emulate the ordinal grading protocol implicitly followed by pathologists, a two-partAttorney Docket No.07039-2297WO1 rank-ordering loss is introduced to train the attention network. It consists of (i) an interclass constraint, which compares patches from different grades and imposes higher attention on more severe tumor patches, and (ii) an intraclass constraint, which imposes higher attention on more likely patches within the same grade. The grade of each patch can be obtained by pseudo- labeling the patches based on their grade class-likelihoods. Rank ordering enables the model to consistently assign higher importance to the most severe tumor region(s), improving tumor localization.
[0009] Second, the effectiveness of the model is improved by combining local and non-local dependencies between tissue regions for grading. A WSI graph can be constructed with patches as nodes and edges defined using a combination of spatial proximity and semantic similarity, e.g., patch feature similarity. This enables long-range message passing extending beyond the immediate neighbors in WSI during graph convolution, allowing us to capture broader tumor structure. Incorporating spatial and semantic context improves the localization and classification of higher-risk tumors (moderately and poorly differentiated), which existing methods find difficult to classify correctly.
[0010] Third, the use of tumor depth as an auxiliary training signal is introduced toenhance grade classification. Tumor grade is significantly associated with tumor depth. Well- differentiated cSCC tumors have lower depth, while poorly-differentiated tumors invade deeper into the tissue. To capture the relationship between depth and grade, a multitask framework is developed to predict depth and grade jointly, sharing the patch features between them.
[0011] In one aspect, this specification describes methods for classifying a grade ofcancer represented in a tumor tissue sample. The methods can include obtaining a whole-slide image (WSI) of the tumor tissue sample and partitioning the WSI into a plurality of patches. For each patch, (i) a contextualized feature representation of the patch is generated based on intrinsic features of the patch, local dependencies between the patch and a subset of local patches of the WSI, and non-local dependencies between the patch and a subset of non-local patches of the WSI; and (ii) an attention weight is determined for the patch. A WSI-level cancer grade for the tumor tissue sample is predicted based on the contextualized feature representations and the attention weights for the plurality of patches.
[0012] The cancer can include cutaneous squamous cell carcinoma (cSCC).Attorney Docket No.07039-2297WO1
[0013] The WSI-level cancer grade for the tumor tissue sample can be selected from agroup comprising normal, well-differentiated, moderately differentiated, and poorly differentiated.
[0014] Generating the contextualized feature representation of the patch can include:generating a non-contextualized feature representation of the patch based on an analysis of the patch to the exclusion of any other of the plurality of patches of the WSI; and deriving the contextualized feature representation of the patch from the non-contextualized feature representation of the patch and information about the subset of local patches and the subset of non-local patches of the WSI.
[0015] The non-contextualized feature representation of the patch can be projected froma higher-dimensional space to a lower-dimensional space before deriving the contextualized feature representation of the patch.
[0016] Generating the non-contextualized feature representation of the patch can includeprocessing the patch with a hierarchical transformer model to extract the intrinsic features of the patch.
[0017] Deriving the contextualized feature representation of the patch can include:generating an adjacency matrix A for the patch based on information about the localdependencies between the patch and the subset of local patches of the WSI and the non-local dependencies between the patch and the subset of non-local patches of the WSI; and processing the non-contextualized feature representation of the patch and the adjacency matrix using graph convolution operations to generate the contextualized feature representation of the patch.
[0018] A graph convolution network can perform the graph convolution operations.
[0019] The graph convolution operations can include graph attention-based convolution.
[0020] An attention network that generates the attention weights for the plurality ofpatches of the WSI can be trained to enforce a first ranking in which greater attention is given to patches representing a more severe grade of cancer in the tumor issue sample than patches representing a less severe grade of cancer in the tumor tissue sample.
[0021] The attention network that generates the attention weights for the plurality ofpatches of the WSI can be trained to enforce a second ranking in which, within a given grade of cancer, greater attention is given to patches that are more confidently predicted to represent theAttorney Docket No.07039-2297WO1 given grade of cancer than patches that are less confidently predicted to represent the given grade of cancer.
[0022] One or more machine-learning models can be used in a process for predicting theWSI-level cancer grade for the tumor tissue sample and are trained based on a tumor depth loss that represents a difference between a predicted tumor depth and a ground-truth tumor depth.
[0023] In another aspect, this specification describes methods for training a cancer-gradeclassifier, comprising: obtaining a plurality of training samples, each training sample comprising a whole-slide image (WSI) of a tumor tissue sample and a first label that indicates a ground-truth cancer grade for the tumor tissue sample; for each training sample of the plurality of training samples: processing the WSI from the training sample to generate a predicted cancer grade for the tumor tissue sample; determining a first loss based on a difference between the ground-truth cancer grade indicated by the first label and the predicted cancer grade for the tumor tissue sample; determining a second loss that enforces an intergrade ranking among patches of the WSI that yields greater attention to patches representing tumor regions with a more severe cancer grade than patches representing tumor regions with a less severe cancer grade; determining a third loss that enforces an intragrade ranking among patches of the WSI that yields greater attention to patches for which a patch-level cancer grade prediction is calculated with greater confidence than patches for which the patch-level cancer grade prediction is calculated with less confidence; and updating trainable parameters of the cancer-grade classifier based on the first loss, the second loss, and the third loss.
[0024] The cancer grade can represent a severity of cutaneous squamous cell carcinoma(cSCC).
[0025] The ground-truth cancer grade and the predicted cancer grade for the tumor tissuesample can be selected from a group comprising normal, well-differentiated, moderately differentiated, and poorly differentiated.
[0026] A first training sample of the plurality of training samples further can include asecond label that indicates a ground-truth tumor depth for the tumor tissue sample, the method further comprising for each training sample: processing the WSI from the training sample to generate a predicted tumor depth for the tumor tissue sample; determining a fourth loss based on a difference between the ground-truth tumor depth indicated by the second label and the predicted tumor depth for the tumor tissue sample; and updating trainable parameters of theAttorney Docket No.07039-2297WO1 cancer-grade classifier further based on the fourth loss, including updating trainable parameters that are used in generating the predicted cancer grade and trainable parameters that are used in generating the predicted tumor depth.
[0027] The trainable parameters can include weights at neurons in one or more neuralnetworks.
[0028] Updating trainable parameters of the cancer-grade classifier can includeperforming gradient descent operations.
[0029] For each training sample, processing the WSI from the training sample togenerate a predicted cancer grade for the tumor tissue sample can include: partitioning the WSI from the training sample into a plurality of patches, each patch corresponding to a different region of the WSI; generating, with a first machine-learning model, non-contextual feature representations of the plurality of patches of the WSI; processing, with a second machine- learning model, the non-contextual feature representations of the plurality of patches and corresponding adjacency matrices to generate contextualized representations of the plurality of patches of the WSI; generating, with a third machine-learning model, attention weights for the plurality of patches of the WSI; and generating, with a fourth machine-learning model, the predicted cancer grade for the tumor tissue sample.
[0030] Each aspect of the disclosure can include a computing system, comprising: one ormore processing devices and one or more computer-readable media storing instructions that, when executed by the one or more processing devices, cause the one or more processing devices to perform any of the methods disclosed herein. Each aspect of the disclosure can additionally or alternatively include one or more non-transitory computer readable media storing instructions that, when executed by one or more processing devices, cause performance of any of the methods disclosed herein. DESCRIPTION OF DRAWINGS
[0031] FIG. 1 is a block diagram of an example computing environment for training andusing a RACR-MIL model to predict WSI-level cSCC and tumor depth.
[0032] FIG. 2 is a flowchart of an example process for predicting a WSI-level cancergrade for a tumor tissue sample using a RACR-MIL model.
[0033] FIG. 3 is a flowchart of an example process for training a RACR-MIL model.Attorney Docket No.07039-2297WO1
[0034] FIG. 4a depicts tiling of a WSI with non-local semantic (blue) and local spatial(black) dependencies between tumor regions with the same grade. FIG.4b depicts between depth and grade: Worse-grade tumor invades deeper into the skin tissue reaching higher depth from the skin surface.
[0035] FIG. 5 depicts a process-flow through an example RACR-MIL model. (a) Tissuetiling and local feature extraction. (b) Derivation of contextual patch features using self- attention-based graph convolution, taking spatial and semantic dependency into account. (c) Joint prediction of grade and depth using multiscale contextual features. (d) Rank-order constraint on the attention network.
[0036] FIG. 6 depicts normalized attention heatmaps highlighting the tumor ROIidentified by different models. The patch level annotations for tumor grade are shown in red in the leftmost panels. In the heatmaps, red indicates high attention region and blue indicates a low attention regions.
[0037] FIG. 7 depicts probability heatmaps highlighting the class likelihood (pn, c) oftumor ROI identified by ABMIL and the RACR-MIL model. The patch level annotations for tumor grade are shown in red in the leftmost panels. ’Red’ indicates a higher grade-class probability (closer to 1) and ‘Blue’ indicates a lower grade-class probability (closer to 0).
[0038] FIG. 8 depicts attention heatmaps highlighting the tumor ROI identified byABMIL and our model. The attention values are normalized between 0 and 1. ‘Red’ indicates a higher attention and ‘Blue’ indicates a lower attention.
[0039] FIG. 9 depicts tSNE visualizations of patch features learnt by ABMIL andvariants of RACR-MIL.
[0040] FIG. 10 is a table that shows comparison of the proposed model with state-of-the-art existing attention-based methods. Bold and underline indicate the highest and second highest performance. ’Mod’ represents moderate-differentiation.
[0041] FIG. 11 is a table that shows impact of ranking ordering constraint and multitasklearning using depth as an auxiliary task. ‘RACR-MIL[Fixed] represents leveraging a graph with fixed edge weights, and RACR-MIL[Learnt] represents leveraging a graph with self- attentionbased dynamic edge weights.
[0042] FIG. 12 is a table that shows impact of WSI graph (spatial edges only vs bothspatial and semantic edges) on grade classification accuracy.Attorney Docket No.07039-2297WO1 DETAILED DESCRIPTION
[0043] FIG. 1 is a block diagram of an example computing environment 100 for trainingand using a RACR-MIL model to predict WSI-level cSCC and tumor depth. Computing environment 100 can host one or more computers in one or more locations to implement the various functional blocks depicted in FIG.1. An initial explanation of the various functional blocks and process flows for training and using the RACR-MIL model follows below with respect to FIGS.1-3. The techniques are then described in further detail below with reference to FIGS.4-12.
[0044] An RACR-MIL system in environment 100 includes WSI partitioner 102.Partitioner 102 receives an digital image file for a WSI 150, which depicts a sample of tissue containing a suspected tumor from a person or other animal of interest. Partitioner 102 segments WSI 150 into a set of non-overlapping patches (tiles) 152 of uniform size, each depicting a different region of the WSI 150. Feature generator 104 processes the set of patches 150 from the WSI 150 to generate, for each patch, an initial set of features 154 for the patch. Feature generator 154 can be a hierarchical transformer pre-trained through self-supervised learning, for example.
[0045] Contextualizer 106 processes the initial feature representations 154 for the patches152 of WSI 150 to generate contextualized feature representations 156, 158 for each patch 152. The contextualized feature representations 156, 158 incorporate local and non-local dependencies, e.g., by using graph convolution that passes information between patch-based feature representations 154 connected in a graph structure based on proximity and morphological similarity. Contextualizer 106 produces two sets of contextualized feature representations 156, 158. The first set 156 is formed by processing the initial feature representations 154 through a single graph convolution layer to integrate context within 1-hop of a given node in the graph structure. The second set 158 is formed by processing the feature representations through an additional graph convolution layer to integrate context within 2-hops of a given node in the graph structure.
[0046] A first attention network 108 processes the first set of contextualized featurerepresentations of the patches 152 to generate a set of attention weights 160. In some implementations, a rank control engine 114 trains first attention network 108 to impose ranking constraints. For example, the rank control can cause first attention network 108 to computeAttorney Docket No.07039-2297WO1 higher (more favorable) attention weights 160 to patches 152 that depict more severe tumor regions and lower (less favorable) attention weights 160 to patches 152 that depict less severe tumor regions. The rank control scheme further cause the first attention network 108 to compute higher (more favorable) attention weights 160 to patches 152 that are more confidently predicted to depict tumor regions within a given cSCC class or grade (e.g., normal (tumor not present), well-differentiated, moderately differentiated, or poorly-differentiated). Patch-level grades can be derived through a pseudo-labeling scheme described further below. The attention weights 160 are provided to cancer grade predictor 112 along with the contextualized feature representations 156, which are processed to predict a WSI-level cancer grade for the tumor shown in WSI 150.
[0047] The system can further be trained to predict the depth of the tumor depicted inWSI 150. To do so, second attention network 110 processes the second set of contextualized feature representations 158 of patches 152 to generate attention weights 166. The attention weights 166 are provided to a separate tumor depth predictor 116 along with the contextualized feature representations 158, which further processes the inputs to generate a tumor depth prediction 116.
[0048] Loss calculator 118 and training engine 120 facilitate training of the RACR-MILmodel. Training engine 120 implements one or more training algorithms to backpropagate losses 168 through the model, which serves to update internal parameters of the model such as the weights or biases of neurons in one or more neural network layers. In some implementations, the overall loss is determined by a weighted sum of four component losses: a cSCC grade prediction loss, a tumor depth prediction loss, an intergrade attention weighting loss, and an intra-grade prediction loss.
[0049] FIG. 2 is a flowchart of an example process 200 for predicting a WSI-level cancergrade for a tumor tissue sample using a RACR-MIL model. The process 200 includes obtaining a WSI (202), partitioning the WSI to generate patches (204), generating contextualized feature representations of the patches (206), generating attention weights with inter- and intra-grade rank control (208), predicting cancer grade (210), and ouputting and applying the cancer grade prediction (212).
[0050] FIG. 3 is a flowchart of an example process 300 for training a RACR-MIL model.The system obtains a set of training samples comprising a WSI and ground truth labels for theAttorney Docket No.07039-2297WO1 cancer grade and tumor depth shown in the WSI (302). For each selected training sample (304), the system then processes the WSI to generate a predicted cancer grade (306), determines the cancer grade prediction loss (308), processes the WSI to generate a tumor depth prediction (310), determines the tumor prediction loss (312), determines the intergrade attention weighting loss (314) and the intragrade attention weighting loss (316), and uses these losses to train the RACR- MIL model (318).
[0051] In more detail, WSI is represented as a bag b of patches Xb = denotesthe number of non-overlapping patches in a WSI, and training samples.level labels b{1,2...B}, where B is the number of training samples. The patch-level labels yn are unknown,and only the bag label Yb {normal, well, moderate, poor} is available.
[0052] Tissue feature extraction includes local feature extraction using self-supervisedlearning followed by contextual feature extraction using a graph convolution network (GCN).
[0053] Self-supervised learning (SSL) can be leveraged to pre-train the patch featureextractor using un-labelled patches extracted from WSIs. To capture fine-grained pathological features (e.g., nuclei details, cell distribution, tumor microenvironment), implementations extract 448x448 sized non-overlapping patches at 20X magnification. Nest-S, a hierarchical transformer, can be used to extract patch features. The feature extractor can be pre-trained using DINO (knowledge distillation-based SSL), which has shown promising performance in MIL-based classification tasks. After pre-training, the transformer network can be used as an offline featureextractor to derive d-dimensional feature fn Rd for each patch xn, leading to a WSIrepresentation of .
[0054] pre-trained features fn^are agnostic to the downstream task, they can beprojected into a lower-dimensional space using a multi-layer perceptron (MLP) with nonlinearactivation to get local patch features h0n. The MLP can be trained along with the graphconvolution network and the rank-aware grade classifier. Thus, the resulting local patch feature h0n^potentially captures information specific to grade and deemphasizes information irrelevant to the downstream task.
[0055] To extract contextual information capturing spatial and semantic dependency, aundirected graph can be derived from each WSI G^= (V,A) with patches as nodes V^and adjacencymatrix A^ RNxN, which represents the connections between the nodes. The edges capture thepathology-related structure and interdependence among tumor regions via two types ofAttorney Docket No.07039-2297WO1contextual dependencies incorporated into the adjacency matrix A: (i) Asem, which representsnon-local dependencies between patches that may be spatially distant, but are similar in terms oftheir tissue structure (grade) in the feature space; and (ii) Asp, which captures local dependenciesbetween a tumor patch and its spatially neighboring patches in the tumor microenvironment. In some implementations, the system may only consider edges to the K-nearest neighbors in both the semantic and spatial space.
[0056] Asem^is defined using pairwise feature similarity between patches:where dsemij^represents the semantic distance between patches I,j^with features fi^and fj. J^ [1,N]denotes the K-nearest neighbors of the ith^patch (K^= 4).
[0057] The derived semantic graph Asem^can be pre-processed using personalizedPageRank kernel-based graph diffusion. This amplifies long-range connections between tumor regions by generating additional edges beyond 1-hop neighbors in the feature space.
[0058] Asp^is computed using inverse distance weighting across spatially K-nearestpatches.where dspij^represents the spatial distance between patches i,j^with spatial coordinates (si,ti) and(sj,tj), respectively. K^= 8 can be used to connect each patch to its immediate neighboringpatches in the tissue. The adjacency matrix A^is the average of the spatial and semantic.
[0059] Aggregation: Graph convolution with residual mapping can beleveraged to derive multiscale contextual patch features (h1n,h2n) from the local patch feature h0n^and adjacency matrix A. Each convolution layer uses message-passing to propagate featureinformation from the 1-hop connected nodes j^of a node i^(Aij^>^0) and updates node features viathe following operation:Attorney Docket No.07039-2297WO1 Hl+1 = Hl + GConv(Hl,A;Wl)(Equation 3) where are the trainable parameters of layer l. In some implementations,can be used and at least two approaches for GConv(·) are contemplated,weights and dynamic-edge weights.
[0060] Fixed-edge weights: This approach uses vanilla graph convolution network(GCN) with weighted message passing using predefined edge weights from adjacency matrix A:(Equation 4) where D˜ = Pj A˜(i,j) is the diagonal matrix of A˜ = A + I. A˜ is used to avoid oversmoothing.
[0061] Dynamic-edge weights: Alternatively, graph attention-based convolution can beleveraged to dynamically define edge weights. This allows us information to be dynamically aggregated from patches connected to a patch based on their pairwise similarity. A single attention head can be used with masking to aggregate the neighboring patch features:where Ql = HlWQl ,Kl = HlWKl and M is the attention mask (WQl ,WKl , WMlijare trainableparameters of layer= 0 if Aij > 0, else Mij = l ). WQl ,WKl are used to learn the edge weightsthrough the dot product QlKl.
[0062] Multi-task learning allows us to use additional information from auxiliary labels(depth) to aid main task prediction (grade). Depth can be selected as an auxiliary label because (i) it is easily available from diagnostic reports, (ii) it is derived from pathology images and reflects tissue structure surrounding the tumor, and (iii) it is well-correlated with grade (Dunn’s pairwise p-values are significant at a 10% level). Separate predictors can be implemented fordepth and grade while the graph network is shared between them to allow feature sharing and to ensure parameter efficiency.
[0063] Attention-based feature aggregation can be used to determine the contribution ofeach patch to overall WSI prediction. Attention-based aggregation outperforms traditional pooling-based methods by up-weighing the most relevant patches.1-hop contextual features canAttorney Docket No.07039-2297WO1 be leveraged to determine the normalized attention weight for each patch using a two-layer network:where a^ Rd^and U^ Rd×d^^are learnable parameters. The overall WSI representation is theattention-weighted average of normalized 1-hop patch features: .
[0064] The WSI class-likelihood is computed using a cosine softmaxclassifier :where c {0,1,2,3} representing {normal, well, moderate, poor}, D represents cosine distance andzc is the prototype (class centroid) of grade class c defined as zc = .
[0065] To counter class imbalance due to a lowercases, class-balanced sampling can be leveraged during training.
[0066] Aspects of the invention can use the ground-truth WSI grade label Yb^to compute aMIL-based cancer grade classification loss:
[0067] Ordinal Ranking of Patches: Two ranking constraints can be applied on theattention network to ensure consistent ranking of tumor regions: (i) an interclass constraint to impose higher attention values for worse patches, and (ii) an intraclass constraint to impose higher attention values for more likely patches within the same class.
[0068] Interclass ranking: Pathologists determine the WSI grade based on the grade ofthe most severe tumor region by implicitly ranking the different tumor sections based on their severity. To model this approach, a ranking can be enforced by using pairwise inequality constraints between patches, such that a more severe patch is ranked higher by the attentionAttorney Docket No.07039-2297WO1network ( . To do so requires the grade of each patch, which isusing a threshold on their class
[0069] In some grade class. A set ofpairs (i,j) of pseudo-labeled patches can be derived belonging to two adjacent classes,c,yjpseudo^= c^+ 1,for some c^ {0,1,2}} and a soft ordinal constraint can beattention weights (wi^<^wj) using the pairwise ranking loss:
[0070] Intraclass ranking: Intraclass ranking can be imposed on the patches to ensure thatthe most confident patches (with higher pn,c) within a particular grade class are weighted higherduring feature aggregation. This ranking constraint directs focus to patches with the highestattention weights or class probability, i.e., S = {n;n^ TopK(wn) TopK(pn,c)} for K^= 50.
[0071] Next, a set of pairs (i,j) can be derived using their class probabilities Z˜ = {(i,j);pi,c^<^pj,c^ ,for i,j^ S,c^ {0,1,2,3}}, where ^= 0.1 is used to limit the number of pairs due tocomputational constraints. A pairwise ranking loss can be imposed on the attention network forpatches in Z˜:11)
[0072] For a pair of patches (i,j), this loss enforces that if pj,c^>^pi,c, then attention weightwj^>^wi.
[0073] Depth Prediction: Depth values are continuous and depend on the global tissuestructure. Since depth is continuous, depth prediction can be formulated as a regression task. Thedepth predictor uses a single-layer MLP regressor Wreg^and a two-layer attention network. 2-hopAttorney Docket No.07039-2297WO1patch features h2n^ can be used for both regression and attention networks since it captures morecontextual information, which is useful for depth prediction.
[0074] The WSI-level feature for depth prediction is Hbdepth^= . where wnd^^is the attention weight computed from h2n^by using an equation 6.
[0075] The depth predictor can be trained using the robust loss (Barron2019) to reduce sensitivity to large errors from outlier depth values (>^10mm):where c^= 2 is the scale
[0076] Overall Loss: The overall loss combines the grade classification loss, depthprediction loss, and the interclass and intraclass attention ranking losses. To balance the losses,the weighing factors 0, 1 and 2 can be used, which are determined using hyperparametertuning: Ltotal^= Lgrade^+ 0Ldepth^+ 1Linter^+ 2Lintra^(Equation 13) Experimental Study
[0077] Datasets: An example RACR-MIL model was evaluated in a real-worldexperimental study. This study utilized a cSCC dataset from a leading US-based hospitalconsisting of 718 hematoxylin and eosin (H&E) stained WSIs scanned at 40X^magnification fortraining and evaluation. The dataset was collected from 2017-2022 and reviewed by a group of 4 expert dermatopathologists. The dataset contains 150 normal, 383 well-differentiated, 108 moderately differentiated, and 77 poorly differentiated cases. A majority of patients were white.
[0078] Pre-processing: To remove background regions and irrelevant tissue sections,each WSI was pre-processed using thresholding and morphological operations. Each tissueregion was downsampled to 20X^magnification and tiled into non-overlapping 448 × 448patches. Patches with minimal texture were removed using image gradient-based entropy.Attorney Docket No.07039-2297WO1
[0079] Training details: Stratified 5-fold cross-validation was performed using a64:16:20 split between training / validation / test sets. The GCN feature extractor and task-specific predictors were jointly trained using the Adam optimizer with a batch size of 16 and a learningrate of 1e 4 for 60 epochs. The evaluation metrics were: cross-validated average of classwiseaccuracy (ACC), macro-averaged F1 score, AUC score, and Matthews Correlation Coefficient (MCC).
[0080] Tumor localization: Fine-grained tumor annotations were obtained to determinethe extent of overlap of tumors with the most-probable tumor regions as predicted by the model. 24 WSIs from the test set were randomly chosen and annotated by two senior pathologists. Each pathologist marked the grades of up to seven most relevant tumor regions in each WSI.
[0081] Baselines: The study compared the present model with state-of-the-art attention-based MIL models, including methods that treat patches independently (ABMIL (Ilse, Tomczak, and Welling 2018), Gated ABMIL (GABMIL) (Ilse, Tomczak, and Welling 2018), CLAM-MB (Lu et al.2021)) and contextual dependency-based methods (PatchGCN (Chen et al.2021), DSMIL (Li, Li, and Eliceiri 2021), TransMIL (Shao et al.2021)). The study also compared variants of RACR-MIL that included only some of the proposed techniques (spatial vs semantic contextual features, fixed vs learned edge weights, attention ranking) to evaluate their contribution to overall performance. To ensure a fair comparison, the study utilized the same pre- trained model as the feature extractor for all approaches. The study adjusted the hyperparameters of existing approaches (embedding size, dropout rate, learning rate) to achieve optimalperformance on our dataset. The study achieved best test accuracy using 0 = 1.0, 1 = 0.5 and 2= 0.25. Results
[0082] Grade Classification: RACR-MIL outperforms state-of-the-art attention modelsby achieving 2-9% improvement in F1-score over existing non-contextual methods (Max / Mean Pooling, CLAM-MB, GABMIL, ABMIL; FIG.10). It achieves a higher classification accuracy for higher-risk tumors (Mod + Poor) compared to the self-attention-based contextual methods TransMIL and DSMIL. The studied RACR-MIL model outperforms TransMIL, which learns pairwise dependencies between all patches, by 12% because it explicitly incorporates spatial and semantic dependency while creating the graph. Moreover, compared to DSMIL and PatchGCN which incorporate semantic dependency and spatial dependency respectively, our approach achieves slightly higher F1-score and is less prone to overfitting on the dominant well-Attorney Docket No.07039-2297WO1 differentiated class. The studied model achieves the best performance in classifying the most challenging class (moderately differentiated) with 19.6% higher accuracy compared to the next best method (DSMIL).
[0083] All three innovations in the framework contribute to improvement in gradeclassification. Graph network-based contextual features combined with rank-ordering loss achieves the highest improvement in F1-score. The higher accuracy is due to the improved feature space that enables better rank ordering and separation of patches that belong to different grade classes (see Appendix). Including depth as an auxiliary task leads to a reduction in F1- score. This might be because depth is a geometric concept, and using 2-hop contextual features might not capture all of the relevant structural information for predicting depth accurately. However, qualitative analysis of tumor localization shows that using depth allows the model to localize the tumor better, capturing more tumor patches corresponding to the grade label.
[0084] Tumor Localization - Qualitative Analysis: The study evaluated the impact of theproposed innovations on tumor localization by considering the normalized attention heatmaps of two representative WSIs (Figure 6). The heatmaps were derived by scaling the attention weightwn for each patch xn^across a . The studied model localizes the tumoraccurately, achieving hightruth fine-grained tumor annotations. Leveraging depth with graph and ranking captures a larger portion of the tumor (cases I and II) while leveraging graph leads to fewer false positives (case II) compared to the baseline approach ABMIL.
[0085] Semantic Dependency: The study finds that leveraging both spatial and semanticdependency leads to higher classification as shown by the table in FIG.12.
[0086] Using spatial dependency allows the model to give consistent importance tonearby tumor regions, and using semantic dependency allows it to aggregate and focus on the relevant spatially distant tumor regions with similar morphology.
[0087] Depth as Auxiliary Task: The study finds that the graph-based contextual featuresand depth information are complementary and incorporating depth further guides the attention weights toward key tumor sections leading to lesser false positives (Figure 6).
[0088] Rank Constraint: Using the rank-ordering constraint allows the model to suitablyweigh the tumor regions by focusing higher attention on the higher-grade tumor within a WSIAttorney Docket No.07039-2297WO1 (Figure 6). The rank-ordering constraint is widely applicable - applying it to the baseline ABMIL framework leads to improved F1-score with higher accuracy in classifying high-risk tumors. Furthermore, imposing this constraint improves the F1-score by 2% for ABMIL and up to 5% for our framework. Additionally, it improves classification accuracy across all grade classes in our framework.
[0089] Fixed vs Learnt Edge Weights: The study found that edge weights using graphattention is more effective than using predefined edge weights for higher-risk cases (FIG.11). It is possible that dynamically learned edge resulted in better performance because the edge weights better reflected the similarities between task-relevant patches.
[0090] In sum, the model achieved improved tumor grading on a real-world datasetcompared to existing methods and also led to improved tumor localization. The proposed innovations are generic and applicable to existing WSI tumor grading methods, as shown in the ablation study. Supplementary Material
[0091] Visualization of Feature Space: As detailed in the discussion of the example studyabove, a random subset consisting of 24 Whole Slide Images (WSI) was selected from the test set for evaluation. This subset encompassed 8 well-differentiated, 9 moderately-differentiated, and 7 poorly-differentiated WSIs. Expert pathologists systematically examined and annotated these WSIs, mainly concentrating on clinically relevant tumor areas. The differentiation levels of these tumor regions were marked by the pathologists. Subsequently, around 18,500 patches, each measuring 448 x 448 pixels, were extracted from the annotated regions within these 24 WSIs. These patches were associated with their corresponding grade labels. The utilization of tSNE plots is illustrated in Figure 6 to represent the acquired patch features visually. The distinct clustering of data points corresponding to patches of varying grades effectively underscores the potency of context-derived features using the graph network. Furthermore, the integration of graph and depth components results in a more distinct separation of high-risk grades (moderately and poorly differentiated) compared to the standard approach (ABMIL).
[0092] Probability Heatmaps: The adoption of an intra-class ranking loss functionenables the RACR-MIL model to concentrate on the most pertinent assortment of patchesAttorney Docket No.07039-2297WO1 aligned with the true grade label. Consequently, the model is adept at identifying regions of interest (ROI) within tumors with a heightened level of confidence, as evidenced by the elevated likelihood of grade-class in all highlighted cases depicted in Figure 4. Additionally, this empowers the model to delineate a broader tumor section in comparison to ABMIL, both for moderately-differentiated (Case c) and poorly differentiated (Case d) tumors.
[0093] Attention Heatmaps: Supplementary heatmaps are presented that encompass allgrade classes to underscore the effectiveness of our novel enhancements in enhancing tumor localization. Figure 5 showcases three instances and their corresponding attention heatmaps. In Case I, the model incorporates the graph and ranking constraints adeptly concentrates on the accurate tumor region (moderately differentiated) with a more pronounced attention level than the well-differentiated region within the WSI. In Case II, the model model demonstrates reduced susceptibility to false positives through the utilization of contextual features and ranking. Case III presents an example where the model, leveraging all innovations (graph, ranking, and depth), allocates slightly greater attention to poorly-differentiated tumor regions as opposed to moderately-differentiated tumor regions.
[0094] Embodiments of the subject matter and the functional operations described in thisspecification can be implemented in digital electronic circuitry, in tangibly embodied computer software or firmware, in computer hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non transitory storage medium for execution by, or to control the operation of, data processing apparatus. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. Alternatively, or in addition, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus.
[0095] The term “data processing apparatus” refers to data processing hardware andencompasses all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. TheAttorney Docket No.07039-2297WO1 apparatus can also be, or further include, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). The apparatus can optionally include, in addition to hardware, code that creates an execution environment for computer programs, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.
[0096] A computer program, which may also be referred to or described as a program,software, a software application, an app, a module, a software module, a script, or code, can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages; and it can be deployed in any form, including as a stand- alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data, e.g., one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple coordinated files, e.g., files that store one or more modules, sub programs, or portions of code. A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a data communication network.
[0097] In this specification, the term “database” is used broadly to refer to any collectionof data: the data does not need to be structured in any particular way, or structured at all, and it can be stored on storage devices in one or more locations. Thus, for example, the index database can include multiple collections of data, each of which may be organized and accessed differently.
[0098] Similarly, in this specification the term “engine” is used broadly to refer to asoftware-based system, subsystem, or process that is programmed to perform one or more specific functions. Generally, an engine will be implemented as one or more software modules or components, installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a particular engine; in other cases, multiple engines can be installed and running on the same computer or computers.
[0099] The processes and logic flows described in this specification can be performed byone or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows canAttorney Docket No.07039-2297WO1 also be performed by special purpose logic circuitry, e.g., an FPGA or an ASIC, or by a combination of special purpose logic circuitry and one or more programmed computers.
[0100] Computers suitable for the execution of a computer program can be based ongeneral or special purpose microprocessors or both, or any other kind of central processing unit. Generally, a central processing unit will receive instructions and data from a read only memory or a random access memory or both. The essential elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, special purpose logic circuitry. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto optical disks, or optical disks. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device, e.g., a universal serial bus (USB) flash drive, to name just a few.
[0101] Computer readable media suitable for storing computer program instructions anddata include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto optical disks; and CD ROM and DVD-ROM disks.
[0102] To provide for interaction with a user, embodiments of the subject matterdescribed in this specification can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user and a keyboard and a pointing device, e.g., a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to a web browser on a user’s device in response to requests received from the web browser. Also, a computer can interact with a user by sendingAttorney Docket No.07039-2297WO1 text messages or other forms of message to a personal device, e.g., a smartphone that is running a messaging application, and receiving responsive messages from the user in return.
[0103] Data processing apparatus for implementing machine learning models can alsoinclude, for example, special-purpose hardware accelerator units for processing common and compute-intensive parts of machine learning training or production, i.e., inference, workloads.
[0104] Machine learning models can be implemented and deployed using a machinelearning framework, e.g., a TensorFlow framework.
[0105] Embodiments of the subject matter described in this specification can beimplemented in a computing system that includes a back end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front end component, e.g., a client computer having a graphical user interface, a web browser, or an app through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (LAN) and a wide area network (WAN), e.g., the Internet.
[0106] The computing system can include clients and servers. A client and server aregenerally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some embodiments, a server transmits data, e.g., an HTML page, to a user device, e.g., for purposes of displaying data to and receiving user input from a user interacting with the device, which acts as a client. Data generated at the user device, e.g., a result of the user interaction, can be received at the server from the device.
[0107] While this specification contains many specific implementation details, theseshould not be construed as limitations on the scope of any invention or on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover,Attorney Docket No.07039-2297WO1 although features may be described above as acting in certain combinations and even initially be claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.
[0108] Similarly, while operations are depicted in the drawings and recited in the claimsin a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0109] Particular embodiments of the subject matter have been described. Otherembodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In some cases, multitasking and parallel processing may be advantageous.
[0110] What is claimed is:
Claims
Attorney Docket No.07039-2297WO1 CLAIMS1. A method for classifying a grade of cancer represented in a tumor tissue sample,comprising: obtaining a whole-slide image (WSI) of the tumor tissue sample; partitioning the WSI into a plurality of patches, each patch corresponding to a different region of the WSI; for each of the plurality of patches of the WSI: (i) generating a contextualized feature representation of the patch based on intrinsic features of the patch, local dependencies between the patch and a subset of local patches of the WSI, and non-local dependencies between the patch and a subset of non-local patches of the WSI; and (ii) determining an attention weight for the patch; and predicting a WSI-level cancer grade for the tumor tissue sample based on the contextualized feature representations and the attention weights for the plurality of patches.
2. The method of claim 1, wherein the cancer comprises cutaneous squamous cellcarcinoma (cSCC).
3. The method of any of claims 1-2, wherein the WSI-level cancer grade for the tumortissue sample is selected from a group comprising normal, well-differentiated, moderately differentiated, and poorly differentiated.
4. The method of any of claims 1-3, wherein generating the contextualized featurerepresentation of the patch comprises: generating a non-contextualized feature representation of the patch based on an analysis of the patch to the exclusion of any other of the plurality of patches of the WSI; and deriving the contextualized feature representation of the patch from the non- contextualized feature representation of the patch and information about the subset of local patches and the subset of non-local patches of the WSI.Attorney Docket No.07039-2297WO15. The method of claim 4, further comprising projecting the non-contextualized featurerepresentation of the patch from a higher-dimensional space to a lower-dimensional space before deriving the contextualized feature representation of the patch.
6. The method of claim 4, wherein generating the non-contextualized feature representationof the patch comprises processing the patch with a hierarchical transformer model to extract the intrinsic features of the patch.
7. The method of claim 4, wherein deriving the contextualized feature representation of thepatch comprises: generating an adjacency matrix A for the patch based on information about the localdependencies between the patch and the subset of local patches of the WSI and the non-local dependencies between the patch and the subset of non-local patches of the WSI; and processing the non-contextualized feature representation of the patch and the adjacency matrix using graph convolution operations to generate the contextualized feature representation of the patch.
8. The method of claim 7, wherein a graph convolution network performs the graphconvolution operations.
9. The method of claim 7, wherein the graph convolution operations include graphattention-based convolution.
10. The method of any of claims 1-9, wherein an attention network that generates the attention weights for the plurality of patches of the WSI is trained to enforce a first ranking in which greater attention is given to patches representing a more severe grade of cancer in the tumor issue sample than patches representing a less severe grade of cancer in the tumor tissue sample.
11. The method of claim 10, wherein the attention network that generates the attention weights for the plurality of patches of the WSI is trained to enforce a second ranking in which,Attorney Docket No.07039-2297WO1 within a given grade of cancer, greater attention is given to patches that are more confidently predicted to represent the given grade of cancer than patches that are less confidently predicted to represent the given grade of cancer.
12. The method of any of claims 1-11, wherein one or more machine-learning models are used in a process for predicting the WSI-level cancer grade for the tumor tissue sample and are trained based on a tumor depth loss that represents a difference between a predicted tumor depth and a ground-truth tumor depth.
13. A computing system, comprising: one or more processing devices; and one or more computer-readable media storing instructions that, when executed by the one or more processing devices, cause the one or more processing devices to perform any of the methods of claims 1-12.
14. One or more non-transitory computer readable media storing instructions that, when executed by one or more processing devices, cause performance of any of the methods of claims 1-12.
15. A method for training a cancer-grade classifier, comprising: obtaining a plurality of training samples, each training sample comprising a whole-slide image (WSI) of a tumor tissue sample and a first label that indicates a ground-truth cancer grade for the tumor tissue sample; for each training sample of the plurality of training samples: processing the WSI from the training sample to generate a predicted cancer grade for the tumor tissue sample; determining a first loss based on a difference between the ground-truth cancer grade indicated by the first label and the predicted cancer grade for the tumor tissue sample; determining a second loss that enforces an intergrade ranking among patches of the WSI that yields greater attention to patches representing tumor regions with a more severe cancer grade than patches representing tumor regions with a less severe cancer grade;Attorney Docket No.07039-2297WO1 determining a third loss that enforces an intragrade ranking among patches of the WSI that yields greater attention to patches for which a patch-level cancer grade prediction is calculated with greater confidence than patches for which the patch-level cancer grade prediction is calculated with less confidence; and updating trainable parameters of the cancer-grade classifier based on the first loss, the second loss, and the third loss.
16. The method of claim 15, wherein the cancer grade represents a severity of cutaneous squamous cell carcinoma (cSCC).
17. The method of any of claims 1-16, wherein the ground-truth cancer grade and the predicted cancer grade for the tumor tissue sample is selected from a group comprising normal, well-differentiated, moderately differentiated, and poorly differentiated.
18. The method of any of claims 1-17, wherein a first training sample of the plurality of training samples further comprises a second label that indicates a ground-truth tumor depth for the tumor tissue sample, the method further comprising for each training sample: processing the WSI from the training sample to generate a predicted tumor depth for the tumor tissue sample; determining a fourth loss based on a difference between the ground-truth tumor depth indicated by the second label and the predicted tumor depth for the tumor tissue sample; and updating trainable parameters of the cancer-grade classifier further based on the fourth loss, including updating trainable parameters that are used in generating the predicted cancer grade and trainable parameters that are used in generating the predicted tumor depth.
19. The method of any of claims 1-18, wherein the trainable parameters comprise weights at neurons in one or more neural networks.
20. The method of any of claims 1-19, wherein updating trainable parameters of the cancer- grade classifier comprises performing gradient descent operations.Attorney Docket No.07039-2297WO1 21. The method of any of claims 1-20, wherein for each training sample, processing the WSI from the training sample to generate a predicted cancer grade for the tumor tissue sample comprises: partitioning the WSI from the training sample into a plurality of patches, each patch corresponding to a different region of the WSI; generating, with a first machine-learning model, non-contextual feature representations of the plurality of patches of the WSI; processing, with a second machine-learning model, the non-contextual feature representations of the plurality of patches and corresponding adjacency matrices to generate contextualized representations of the plurality of patches of the WSI; generating, with a third machine-learning model, attention weights for the plurality of patches of the WSI; and generating, with a fourth machine-learning model, the predicted cancer grade for the tumor tissue sample.
22. A computing system, comprising: one or more processing devices; and one or more computer-readable media storing instructions that, when executed by the one or more processing devices, cause the one or more processing devices to perform any of the methods of claims 15-21.
23. One or more non-transitory computer readable media storing instructions that, when executed by one or more processing devices, cause performance of any of the methods of claims 15-21.
Citation Information
Patent Citations
Methods of diagnosing and treating patients with cutaneous squamous cell carcinoma
US20210238688A1
Systems and methods for processing electronic images for ranking loss and grading
US20230245480A1
Prediction of brcaness / homologous recombination deficiency of breast tumors on digitalized slides
WO2023006843A1
Cited By
Rectum cancer postoperative recurrence risk prediction system and method based on multi-modal time sequence data
CN121617634A
WSI classification method based on dynamic graph modeling and knowledge perception attention mechanism
CN121725286A