Tissue-based classifier of allograft inflammation

EP4802103A1Pending Publication Date: 2026-09-09MAYO FOUNDATION FOR MEDICAL EDUCATION & RESEARCH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
EP2024886985
Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-02
Filing Date
2024-11-01
Publication Date
2026-09-09

AI Technical Summary

Technical Problem

Current methods for diagnosing allograft inflammation in kidney transplant recipients are inadequate, as they rely on qualitative interpretations of biopsies and lack precise identification of inflammation causes.

Method used

A machine-learning-based system that uses imaging mass cytometry (IMC) data to train models for predicting causes of allograft inflammation by analyzing cellular features and generating biosignatures.

Benefits of technology

The system achieves high accuracy in classifying causes of allograft inflammation, enabling early and appropriate therapeutic interventions and improving patient outcomes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024054128_08052025_PF_FP_ABST
    Figure US2024054128_08052025_PF_FP_ABST
Patent Text Reader

Abstract

This specification discloses systems, methods, devices, and other techniques for a tissue based classifier. Systems and methods of predicting causes of allograft inflammation are disclosed. In another aspect, a system comprising an imaging mass cytometry device configured to scan a tissue sample to collect imaging mass cytometry data and a computing system operating an allograft inflammation classifier configured to identify a type of allograft inflammation in the tissue sample by comparing a biosignature with imaging mass cytometry data is disclosed.
Need to check novelty before this filing date? Find Prior Art

Description

TISSUE-BASED CLASSIFIER OF ALLOGRAFT INFLAMMATIONCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of priority to U.S. Patent Application Serial No. 63 / 547,071, filed November 2, 2023. The disclosure of the prior application is considered part of, and is incorporated by reference in its entirety, in the disclosure of the present application.BACKGROUND

[0002] Up to one in three kidney transplant recipients experience some form of transplant rejection, despite the use of immunosuppressants. Molecular phenotyping of transplant cases shows extensive heterogeneity, far exceeding previous notions of discrete transplant rejection categories.

[0003] Current standards-of-care dictate the appropriate method of intervention for when a patient presents with some form of renal transplant rejection. In some examples, a hematoxylin and eosin (H&E)-stained biopsy, followed by qualitative interpretation by a pathologist is performed to determine a tissue’s heterogeneity.

[0004] Imaging mass cytometry (IMC) is used to study complex interaction between different types of cells in a tissue sample. This information can be used to gain insight into various biological processes.SUMMARY

[0005] This document describes systems, methods, devices, and other techniques for training and using machine-learning models to predict causes of allograft inflammation.

[0006] In one aspect, a method of predicting causes of allograft inflammation is disclosed. The method comprising obtaining a training data set that comprises a collection of training samples, each training sample corresponding to a particular cell or a particular cluster of cells in one of a plurality of studied allografts and indicating (i) values for one or more features of the particular cell or the particular cluster of cells derived from an analysis of at least one image of the particular cell or the particular cluster of cells obtained using imaging mass cytometry (IMC) and (ii) a training label that identifies a known cause, from among a plurality of known causes, of inflammation of the studied allograft to which the particular cell or the particular cluster of cellsbelongs, applying a machine-learning algorithm to the training data set to train a machine learning model, thereby learning associations between a plurality of cellular features derived from imaging mass cytometry (IMC) data and different causes of allograft inflammation, extracting, from the machine learning model, a variable importance for each of the plurality of cellular features, generating a biosignature by selecting a subset of the cellular features from the plurality of cellular features based on the extracted variable importance for each cellular feature in the plurality of cellular features, and classifying a cause of allograft inflammation in a tissue sample by comparing the biosignature to corresponding cellular features of the tissue sample derived from IMC data for the tissue sample.

[0007] In some implementations, the studied allografts include a kidney allograft and the tissue sample is a sample of kidney tissue. In some implementations, the studied allografts include a liver allograft and the tissue sample is a sample of liver tissue. In some implementations, the studied allografts includes a lung allograft, and the tissue sample includes a sample of lung tissue. In some implementations, the studied allografts includes a heart allograft, and the tissue sample includes a sample of heart tissue.

[0008] Another aspect is a system, comprising one or more computers and one or more computer-readable storage media encoded with instructions that, when executed by the one or more computers, cause the one or more computers to obtain a training data set that comprises a collection of training samples, each training sample corresponding to a particular cell or a particular cluster of cells in one of a plurality of studied allografts and indicating (i) values for one or more features of the particular cell or the particular cluster of cells derived from an analysis of at least one image of the particular cell or the particular cluster of cells obtained using imaging mass cytometry (IMC) and (ii) a training label that identifies a known cause, from among a plurality of known causes, of inflammation of the studied allograft to which the particular cell or the particular cluster of cells belongs, apply a machine-learning algorithm to the training data set to train a machine learning model, thereby learning associations between a plurality of cellular features derived from imaging mass cytometry (IMC) data and different causes of allograft inflammation, extract, from the machine learning model, a variable importance for each of the plurality of cellular features, generate a biosignature by selecting a subset of the cellular features from the plurality of cellular features based on the extracted variable importance for each cellular feature in the plurality of cellular features, and classify acause of allograft inflammation in a tissue sample by comparing the biosignature to corresponding cellular features of the tissue sample derived from IMC data for the tissue sample.

[0009] Yet another aspect is a system, comprising an imaging mass cytometry device configured to scan a tissue sample to collect imaging mass cytometry data, a computing system operating an allograft inflammation classifier configured to identify a type of allograft inflammation in the tissue sample by comparing a biosignature with imaging mass cytometry data, wherein to generate the biosignature includes to obtain a training data set that comprises a collection of training samples, each training sample corresponding to a particular cell or a particular cluster of cells in one of a plurality of studied allografts and indicating (i) values for one or more features of the particular cell or the particular cluster of cells derived from an analysis of at least one image of the particular cell or the particular cluster of cells obtained using imaging mass cytometry (IMC) and (ii) a training label that identifies a known cause, from among a plurality of known causes, of inflammation of the studied allograft to which the particular cell or the particular cluster of cells belongs, apply a machine-learning algorithm to the training data set to train a machine learning model, thereby learning associations between a plurality of cellular features derived from imaging mass cytometry (IMC) data and different causes of allograft inflammation, extract, from the machine learning model, a variable importance for each of the plurality of cellular features, and generate a biosignature by selecting a subset of the cellular features from the plurality of cellular features based on the extracted variable importance for each cellular feature in the plurality of cellular features.DESCRIPTION OF DRAWINGS

[0010] FIG. 1 shows an example environment in which imaging mass cytometry (IMC) data is collected from a tissue sample and analyzed to classify a cause of allograft inflammation in the tissue sample.

[0011] FIG. 2 depicts a flowchart of an example process for classifying a cause of allograft inflammation in a tissue sample.

[0012] FIG. 3 depicts a flowchart of an example process 300 for generating a training data set for training a classifier to predict a cause of allograft inflammation in a tissue sample.

[0013] FIG. 4A illustrates an Ir channel corresponding to the DNA counterstain, used for segmentation of nuclei.

[0014] FIG. 4B illustrates nucleus segmentation and whole cell simulation using Universal StarDist for QuPath.

[0015] FIG. 4C illustrates multi-channel visualization of most structural and immune-related markers present in panel.

[0016] FIG. 4D illustrates composite classification of all single positive classes combined and applied sequentially.

[0017] FIG. 4E illustrates cells classified as positive for each marker present in the panel, using pathologist-verified thresholds.

[0018] FIG. 4F illustrates percentage of cells positive for each immune marker (x-axis), grouped by clinical phenotype (y-axis).

[0019] FIG. 5A illustrates a pseudocolored image with alpha-SMA, Vimentin, Pan-Cytokeratin, Ecadherin and the Ir DNA counterstain provided to pathologists for annotation.

[0020] FIG. 5B illustrates example of a pathologist’s annotations for glomeruli, arteries, tubules, Interstitium, and empty space used for training the renal structure pixel classifier.

[0021] FIG. 5C illustrates predicted pixel classification from a trained classifier.

[0022] FIG. 5D illustrates a pixel classifier output from trained classifier, with manual pathologist’s annotations.

[0023] FIG. 5E illustrates percent positive scoring for immune markers for glomeruli.

[0024] FIG. 5F illustrates percent positive scoring for immune markers for arteries.

[0025] FIG. 5G illustrates percent positive scoring for immune markers for interstitium.

[0026] FIG. 5H illustrates percent positive scoring for immune markers for tubules.

[0027] FIG. 6A illustrates a pictographic representation of patient classification using modebased data aggregation.

[0028] FIG. 6B illustrate a confusion matrix representing ground truth (y axis) versus predicted (x axis) label.

[0029] FIG. 6C illustrates a sum of all feature importances, grouped by the marker they are derived from.

[0030] FIG. 6D illustrates a heatmap of the top 10 most valuable features (MVF), in one example implementation, after z-score normalization, grouped by clinical phenotype type.

[0031] FIG. 6E illustrates a cumulative model performance when incrementally including all features of a given marker, with the order determined through Fig. 6C.

[0032] FIG. 6F illustrates a visualization of MVF5 (Cluster mean: Gd (158)_158Gd- E Cadherin: Cytoplasm: Median) on an ROI from a patient with ABMR (left) versus ABMR* (right).

[0033] FIG. 7A illustrates a heat map showing the z-scored mean marker expression for each PhenoGraph cluster, colored by cluster identifier.

[0034] FIG. 7B illustrates a map using t-distributed stochastic neighbor embedding (t-SNE) of 900,000 single cells from high-dimensional images of renal allograft inflammation colored by cell-type.

[0035] FIG. 7C illustrates a stacked bar plots depicting the proportion of immune cell clusters identified in each case, distributed in the renal microstructural compartments, and clinical phenotypes.

[0036] FIG. 7D illustrates a panel of pseudo-colored IMC images depicting 3 cell clusters.

[0037] FIG. 8A illustrates a heat map in which squares indicate the Pearson correlation of cell phenotype proportions across all measured tissue regions and circles indicate significant pairwise cell-type interaction or avoidance summarized across the two-sided permutation tests on the individual images.

[0038] FIG. 8B shows an illustration highlighting clusters in Fig.8A that showed significant interactions associated with poor graft outcomes in one example implementation.

[0039] FIG. 8C illustrates a boxplot showing identification of immune clusters that spatially localize in different microcompartments of the kidney, namely tubules, interstitium and glomeruli while correlating with certain Banff lesion scores.

[0040] FIG. 8D illustrates pseudocolored IMC images showing 2 clusters (memory T cells and macrophage / APC clusters) that are seen in increasing proportions with higher scores of Banff tubulitis.

[0041] FIG. 8E: illustrate two pseudocolored IMC images that show two cell clusters (epithelial mesenchymal transformation and memory T cells) that in higher proportion correlated with poor graft outcome.

[0042] FIG. 9 illustrates an example implementation with use of imaging mass cytometry in developing a tissue based classifier of renal allograft inflammation.

[0043] FIG. 10 is a heatmap showing the z-scored mean feature value for each of the 31 clusters identified by PhenoGraph.

[0044] FIG. 11 illustrates the percentage positivity of specific immune subpopulations across the study population standardized within subpopulation using z scores.

[0045] FIG. 12 illustrates an IM based biopsy report.

[0046] FIG. 13A illustrates an analysis flowchart.

[0047] FIG. 13B illustrates heat maps of immune markers.

[0048] FIGs. 13C illustrate example results of the example implementation 3.

[0049] FIGs. 13D illustrate example results of the example implementation 3.

[0050] FIG. 14 shows an example of a computing device and a mobile computing device that can be used to implement the techniques described herein.

[0051] Like reference symbols in the various drawings indicate like elements.DETAILED DESCRIPTION

[0052] This specification describes techniques for training and using a tissue-based classifier that can predict one or more causes of allograft inflammation. The tissue-based classifier processes imaging mass cytometry (IMC) data collected from a tissue sample to predict a cause of the allograft inflammation in a tissue sample. In various implementations, machine learning techniques are used including supervised and / or unsupervised machine learning techniques.

[0053] In some implementations unsupervised analysis of spatially resolved single cell pathology using imaging mass cytometry allows for the identification of single cell subgroups and / or clusters within graft inflammation that associate with poor graft outcomes after kidney transplantation.

[0054] Methods for accurately determining the cause of allograft inflammation are disclosed. A machine learning model is trained to learn the associations between cellular features and different types of allograft inflammation. The machine learning model is configured to compute the variable importance of each feature used in the training, where features with a higher importance are more crucial for the model to achieve its high accuracy. Variable importance is used to evaluate which markers to include as part of the gene signature, based on the cumulative importance of all features derived from a given marker.

[0055] In some examples, the classifier is trained by processing input data representing the processed per-cell statistics containing a label that indicates a cause of transplant rejection (sometimes referred to as “code”) for each cell, and a list of features to include in the training ofthe model (referred to as “features”). The per-cell statistics data is provided without features relating background channels. Features are grouped into variable X, and the code assigned to variable Y. In some cases, a 25 / 75 test-train split is performed, with a fixed seed state to preserve model reproducibility between multiple successive runs. In some examples, z-score normalization is applied to the X train scaled and X test scaled variables and code labels are encoded as discrete values using a label encoder. In some implementations, the machine learning model is a regularized gradient boosting model XGBoost. In some examples, the seed state is fixed, and tree construction is performed using the GPU-accelerated histogram method. Example advantages of the GPU-accelerated histogram method include an increase in model performance and improved accuracy. Variable importances are extracted from the trained model, and each feature has a corresponding importance score. In some examples, accuracy is evaluated by applying the trained model to our test set (X test to infer Y test). To identify the importance of each marker in training, the model is trained on the mean marker intensity measurement of each feature and sorted these features by their variable importance. The model is iteratively retrained, starting with all features relating to the marker with the highest variable importance, and successively appending features with descending variable importance grouped by the marker they were derived from. For example, all features derived from the gene(s) most important in predicting allograft rejection may be used for training first, to teach the model what the most important markers are in the panel.

[0056] Once training of the model has completed and evaluated to yield consistent results across multiple batches, a list of protein markers prognostic for allograft rejection when stained, imaged can be provided in a diagnostic setting, and analyzed through the method as described herein.

[0057] FIG. 1 depicts an example environment 100 in which imaging mass cytometry (IMC) data is collected from a tissue sample and analyzed to classify a cause of allograft inflammation in the tissue sample. The IMC data 110 is collected using an IMC device 102. The IMC device 102 is configured to communicate with the analysis system 104 to send the IMC data. The Analysis system 104 includes a classifier 106.

[0058] In some implementations, the analysis system 104 includes prediction tools for allograft loss, including the classifier 106. In some cases, these predictions can be used as part of a care plan to prevent poor outcomes. In some examples, the prediction tools identify cluster-wise cell- to-cell interactions (e g., colocalization and avoidance) within regions of interest and by spatialcontext analysis to identify interactions of cellular neighborhoods to correlate inflammation subtypes to graft outcomes. In some examples, artificial intelligence algorithms analyze IMC data 110 captured at an IMC device 102. Examples of the IMC data include IMC-based multiplexed immunohistochemistry images of renal allograft inflammation. The artificial intelligence algorithms are used to enhance the accurate phenotypic classification of allograft inflammation permitting early and appropriate therapy.

[0059] The classifier 106 is configured to generate an output indicating a predicted allograft inflammation type (a cause of allograft inflammation) and allows for disease specific therapeutic intervention. In some examples, the classifier 106 is used on a tissue sample of a tumor (e.g., breast cancer), where the applications the classifier 106 are used to reveal multicellular features of the tumor microenvironment and novel subgroups of the tumor that are associated with distinct clinical outcomes, characterizing intratumor phenotypic heterogeneity in a diseaserelevant manner informing patient-specific diagnosis. Example uses of the classifier 106 include disease processes that are harder to diagnose histologically but are treated completely differently once diagnosed: For example, in kidney allograft biopsies, classified conditions can include cellular rejection, antibody-mediated rejection, BK nephropathy and pyelonephritis. The treatment of the first two include boosting immunosuppression while the latter two require reduction of immunosuppression. In liver allograft biopsies, the cases include cellular rejection, antibody-mediated rejection, plasma cell-rich rejection, and idiopathic post-transplant hepatitis. Similar to the case in kidney transplantation, while some of these diagnoses require more immunosuppression, some do not. Hence, in both cases, the IMC-based and Al-assisted diagnostic tools disclosed herein help diagnose the condition earlier, permit selection of the appropriate treatment and permit identification of cell groups predictive of poorer outcomes.

[0060] In some implementations, the classifier 106 includes a machine learning model that is trained on highly multiplexed imaging of renal allograft biopsies with subcellular resolution by IMC. In some implementations, the classifier 106 is further configured to identify single cell phenotypic clusters that correlate with graft loss.

[0061] The IMC device 102 is a device that is configured to scan a tissue (often with the assistance of one or more dyes) to capture IMC data. Other devices in the multiplex spatial imaging modalities are used in some implementations. In some implementations, the IMC device102 is a hyperion imaging system. In some implementations, the IMC device uses a panel of 28 markers. In some implementations, the IMC data 110 includes spatially resolved single cell data.

[0062] Examples of tissues that can be included in the tissue sample are kidney tissue, liver tissue, lung tissue. For example, the studied allografts include a kidney allograft and the tissue sample is a sample of kidney tissue. In another example, the studied allografts include a liver allograft and the tissue sample is a sample of liver tissue. In yet another example, the studied allografts includes a lung allograft and the tissue sample includes a sample of lung tissue. In some examples, the studied allografts includes a heart allograft, and the tissue sample includes a sample of heart tissue.

[0063] FIG. 2 depicts a flowchart of an example process 200 for classifying a cause of allograft inflammation in a tissue sample. The process 200 includes the operations 202, 204, 206, 208 and 210. In some implementations, one or more of the operations 202, 204, 206, 208, and 210 are executed on the analysis system 104 as part of the classifier 106, as illustrated and described in reference to FIG. 1.

[0064] The operation 202 obtains a training data set that comprises a collection of training samples, each training sample corresponding to a particular cell or a particular cluster of cells in one of a plurality of studied allografts. In some implementations, each training sample indicates values for one or more features of the particular cell or the particular cluster of cells derived from an analysis of at least one image of the particular cell or the particular cluster of cells obtained using imaging mass cytometry (IMC) and a training label that identifies a known cause, from among a plurality of known causes, of inflammation of the studied allograft to which the particular cell or the particular cluster of cells belongs.

[0065] The operation 204 applies a machine-learning algorithm to the training data set to train a machine learning model, thereby learning associations between a plurality of cellular features derived from imaging mass cytometry (IMC) data and different causes of allograft inflammation. In some implementations, the machine learning model learns to identify one or more cell clusters of the tissue sample associated with a predicted outcome and uses spatial context analysis to identify interactions with neighboring cells of the one or more cell clusters associated with the predicted outcome.

[0066] In some implementations, the machine learning model is trained on annotated multiplexed immunohistochemistry datasets. In some of these implementations, the annotationsinclude single measurement thresholds based on mean intensity. In some implementations, the machine learning model is an unsupervised machine learning model.

[0067] In some implementations, the machine learning model is trained to learn associations between the cellular features derived from the biomarkers and different types of allograft inflammation by training the machine learning model on a mean marker intensity measurement for each of the cellular features, extracting the variable importance for each biomarker of the biomarkers based on calculated variable importances for the cellular features derived from the biomarker, and retraining the machine learning model iteratively starting with the cellular features derived from the biomarker with a highest importance and successively appending the cellular features derived from the biomarkers with descending variable importance. In some of these implementations, the cellular features include cellular measurements and one or more spatial features.

[0068] The operation 206 extracts, from the machine learning model, a variable importance for each of the plurality of cellular features. In some implementations, the variable importance for each cellular feature is extracted by calculating, for each cellular feature, an importance score for indicating an effect of the feature on an accuracy of predictions made by the machine learning model. In some implementations, the variable importance for each cellular feature of the cellular features is extracted by calculating, for each cellular feature, an importance score for indicating an effect of the feature on an accuracy of predictions made by the machine learning model.

[0069] Referring to operations 204 and 206, in some implementations, the machine learning model is trained to learn the associations between the cellular features derived from the biomarkers and the different types of allograft inflammation by training the machine learning model on a mean marker intensity measurement for each of the cellular features, extracting the variable importance for each biomarker of the biomarkers based on calculated variable importances for each of the cellular features; and retraining the machine learning model iteratively starting with the cellular features with a highest importance and successively appending the cellular features with descending variable importance. In some of these implementations, the cellular features include a cellular measurement and one or more spatial features.

[0070] The operation 208 generates a biosignature by selecting a subset of the cellular features from the plurality of cellular features based on the extracted variable importance for each cellularfeature in the plurality of cellular features. Tn some examples, the biosignature includes a plurality of immunological markers. For example, the biosignature may include twenty-eight immunological markers. In some implementations, the subset of cellular features is based on a cumulative importance of all featured derived from a given biomarker.

[0071] The operation 210 classifies a cause of allograft inflammation in a tissue sample by comparing the biosignature to corresponding cellular features of the tissue sample derived from IMC data for the tissue sample. In some implementations, the IMC data from the tissue sample is captured by scanning the tissue sample with an imaging mass cytometry (IMC) device. For example, the tissue sample may be from a kidney identified as a transplant, where a tissue sample from the kidney is scanned with an IMC device. In other examples, the tissue sample may be from a tumor.

[0072] In some implementations, the operation 210 includes a sub-operation for scoring the tissue sample based, at least in part, on the biosignature and the imaging mass cytometry data, where the score indicates a predicted outcome for a transplant of the tissue sample. In some implementations, the tissue sample is classified is one of a plurality of allograft inflammation types. For example, the classified allograft inflammation type maybe selected from a group of four different allograft inflammation types.

[0073] In some implementations, the process 200 further includes providing an output indicating a relative importance of one or more of the plurality of cellular features for the predicted outcome. In some implementations, the plurality of cellular features and associated variable importance are indicative of the relative importance of one or more markers for a predicted outcome. In some implementations, the relative importance of each of the one or more markers for a predicted outcome are output (e.g., via a user interface on a display device). This output can include measured density of single and multiplexed biomarkers in the tissue sample. In some implementations, the output includes a visualization representing a spatial relation to renal microstructural components of the tissue sample.

[0074] FIG. 3 depicts a flowchart of an example process 300 for generating a training data set for training a classifier to predict a cause of allograft inflammation in a tissue sample. The process includes the operations.

[0075] The operation 302 generates a combined training image. In some implementations, the training images are IMC images. Examples of the biopsies sampled to generate the IMC imagesinclude biopsies of rejection, BK nephropathy, pyelonephritis & normal kidneys. In some implementations, the training images are cropped for a representative field of view, where the cropped images are concatenated into one large image. Thresholds are set for the one large image. In some implementations, each training image is copped to create a 500 by 500 pixel cropped image (capturing a representative area), and these training images are stitched toother to create the large training image. In other implementations, each of the training images are individual annotated.

[0076] In some implementations, the training images are selected by evaluating representative images and channels in consultation with a pathologist. For example, channels with inaccurate staining can be filtered out of the training images dataset.

[0077] The operation 304 performs cell segmentation on the combined training image. In some implementations, cell segmentation is performed on the combined large training image discussed above. In some implementations, cell segmentation is performed using a deep learning technology. For example, Universal StartDist for Qupath can be used to perform cell segmentation. In some of these implementations, segmentation is performed using theIr(193) 193Ir-DNA193 with a 10 pm expansion, 0.5 detection probability threshold, 1stto 99thpercentile normalization, using a pretrained model.

[0078] The operation 306 creates a single-measurement classifier for each marker. In some implementations, the cells are classified based on mean intensity threshold. In some examples, a single measurement classifier is made for each marker using the “cell mean” or “nucleus mean”.

[0079] In some example implementation, each of the IHC markers, the cellular subcompartment (cell or nuclear) to use was identified and a single measurement threshold is set based on mean intensity. In some implementations, the identified cellular sub compartment and / or set single measurement threshold are recorded and verified by a pathologist.

[0080] The operation 308 creates a composite classifier by combining all single measurement classifiers. In some examples, the composite classifier is generated by assigning evry segmented cell to one or more IHC marker classes. In some implementations, the composite classifier is applied to the segmented cells.

[0081] The operation 310 computes spatial features. In some implementations, intensity based descriptive statistics (mean, min, max, std, median) were computed for the cell and cellular subcomponents (nucleus, cytoplasm, membrane) for all IHC markers in the panel. In someimplementations, cell morphology features are included. In some implementations, spatial features include neighborhood averaged features. Neighborhood averaged features take the average of each feature with cells within a defined radius around it (e.g. 50 pm). In some implementations, spatial features include cell-to-cell distances features, where the cell-to-cell distance features provide the distance of each cell to the nearest cell positive for each of the markers included in the panel. In some implementations, the cell-to-cell distance features are determined with a maximum distance cap of 200 pm. In some implementations, the spatial features include cluster features. Cluster features denote descriptive statistics relating to all cells grouped together by a Delaunay triangulation was computed according to a clustering threshold (15 pm apart) and containing the same classification.

[0082] The operation 312 exports the training dataset. In some implementations, the dataset is exported in a “ csv” format, where each row corresponds to an individual cell and each column corresponds to a feature in that cell. In some implementations, patient metadata is included in the dataset. In some implementations, a sperate script is executed to inject the supplementary patient metadata to each cell in the dataset.

[0083] In some implementations, Uniform Manifold Approximation and Projection (UMAP) is used to reduce the dimensionality of the dataset down to two with the number of input dimensions corresponding to the number of IHC markers present in the panel. Z-score normalization is applied prior to fitting into an embedded space for UMAP.

[0084] In some implementations, the allograft inflammation-labelled dataset is used train a regularized gradient boosting model XGBoost. Variable importances for each feature are then computed, and the list of features are sorted to identify histological markers which yielded the highest importance for making accurate predictions. A subset of the top markers of this list will constitute a gene signature for allograft inflammation, after sufficient validation and quality control has been performed. Ultimately, performing the imaging, segmentation, and classification methods described herein can be used to provide prognostically actionable information to a pathologist, by predicting the cause of allograft inflammation.

[0085] In some implementations, prior to starting the process 300 a computing environment used in the process 300 is set up. In some implementations, setting up the environment includes installing QuPath with StartDist, python, and scripts for QuPath and Python. In some implementations, a new QuPath project is created and image from a training batch are imported.

[0086] In some implementations, the dataset is visualized. For example, violin plots of each measurement of interest are generated to show the distribution of IHC marker intensities differed between the different groups of transplant rejection. In some examples, percent positive scoring is performed in a similar manner, quantifying the percentage of single or multiple positives present in each group of transplant rejection cases. For example, a percentage of cells positive for multiple markers is calculated. In some implementations, a script is used to compute the percentage of specific immune cell subtypes defined by more than one histological marker.Example Implementation 1:

[0087] This example delineates a study of an analytical pipeline for imaging mass cytometry that incorporates Universal StarDist for QuPath (USDQP) and “Pathologist-in-the-loop” cell classification in order to characterize the single-cell landscape of renal allograft inflammation. Data from patients consisting of 247 high-dimensional histopathology images containing more than 900,000 cells were used to develop a diagnostic tissue classifier with an accuracy of 84% in a blinded validation set. The machine-learning classifier analyzes spatial features. The interpretation of unsupervised PhenoGraph clustering was performed by exploiting the manual gating thresholds, set by an expert pathologist, to guide cluster nomenclature. Clusters were identified in specific microstructural compartments that associate with poor graft function. Methods to automate the analysis of complex multiplex datasets can be used to facilitate the identification of biologically and prognostically relevant interactions.

[0088] Example implementation 1 introduction.

[0089] Multiplexed in situ imaging platforms allow for in-depth and spatially enriched characterization of tissue architecture. Some techniques permit detailed interrogation of cellular biomarker heterogeneity and location within the tissues. The initial analysis of human kidney tissue by Imaging Mass Cytometry (IMC) helped characterize its cellular composition through development of a machine learning-based analysis pipeline, Kidney-MAPPS (Multiplexed antibody-based profiling with preservation of spatial context). This pipeline used Ilastik for unbiased pixel classification, CellProfiler for nuclear-based cell segmentation and HistoCAT for clustering / neighborhood analysis. The Kidney-MAPPS protocol accurately identified, quantified, and localized -92% of all cells in the human kidney, and was validated with subsequent analysis of preimplantation deceased donor kidneys . One implementation of the example implementationuses a similar approach and incorporates a robust immune marker antibody panel, and demonstrated a dysregulated immune response in both COVID-19 kidneys and immune checkpoint inhibitor-associated kidney injury. IMC -based analysis pipelines utilize the properties of protein marker localization to identify cell types, and can allow unsupervised identification and quantification of microanatomical tissue structures in tissues with pathology. In one implementation of the example implementation, a classifier predicts clinical outcomes through interrogation of cellular interactions with other cells, as well as with the tissue architecture.

[0090] Renal allograft inflammation is sometimes seen after transplantation, and is an important cause of graft failure. This inflammation can be caused by different disease processes (e.g., rejection, viral or bacterial infections). Therefore, accurate phenotypic characterization of the renal allograft inflammation has important prognostic and therapeutic implications. The Banff Classification of Kidney Allograft Pathology provides criteria for the diagnosis of kidney transplant rejection. Example implementations have improved reproducibility and are able to unravel distinct mechanisms of pathogenesis that impact prognosis and therapy. In some examples, molecular diagnostic tools have been incorporated to the Banff classification scheme. These tools identify pathogenesis-based transcript sets that correlate with histologic lesions of rejection. Some implementations tools allow for correlation with the tissue microarchitecture. Imaging platforms with spatial profiling capacity are used in some implementations. Some examples use multiplexed immunofluorescence imaging and whole exome GeoMX Digital Space Profiling platform to evaluate renal allograft rejection.

[0091] The example implementation includes a method to evaluate renal allograft inflammation by combining pathologist-guided cell biomarker classification and a tissue-level deep learning algorithm for segmentation and classification of renal structures to augment the set of spatial localization features available at the single-cell level. Using IMC to simultaneously quantify 28 biomarkers across 247 high-dimensional pathology images of renal allograft inflammation from 32 patients, each with associated graft-survival data, we used the spatial features to develop a novel, highly accurate classifier predicting clinical phenotype. Applying utilized dimensionality reduction and clustering-based IMC methodologies, to identify novel cell clusters in the renal allograft microenvironment that are associated with poor graft outcomes. In some examples, the ability to identify patients who will progress with a high degree of certainty is used to guidefuture patient management. The spatially augmented techniques disclosed herein are translatable to other organ systems and disease groups.

[0092] Example implementation 1 results.

[0093] In some implementations, cell segmentation is performed using universal StarDist for QuPath (USDQP). In some examples, IMC data includes differences in spatial resolution and pixel intensity distribution that is considered for cell segmentation. StarDist is a deep learning segmentation algorithm that uses star-convex shape representation to enable robust and reliable nucleus segmentation, trained from over 37,000 manually annotated nuclei. A transfer-learning style approach to segment IMC datasets using pre-trained Startdist models. The transfer-learning style approach accounts for differences between fluorescence and mass cytometry-based images, we utilized a transfer-learning style approach to segment IMC datasets using pre-trained StarDist models. In some examples, the cell segmentation algorithm is accessible via a published API. USDQP was performed using one or more parameters to segment over 900,000 cells (FIG. 4A and FIG. 4B) within a cohort of 32 cases of renal allograft inflammation, representing 5 clinical phenotypes (Acute T cell-mediated rejection [ACR], clinically acute antibody -mediated rejection [ABMR*], active and chronic active antibody -mediated rejection [ABMR], BK virus nephropathy [BK] and chronic pyelonephritis [C. Pyel]) . These biopsies had been stained with a panel of 28 IMC biomarkers that identify renal structures (e.g., NA-K ATPase, e-cadherin, Vimentin) (FIG. ID), and immune cells (e.g., CD68, CD4, CD8, Vista, PDL1, CD19). The biomarker panel was designed to identify and distinguish different immune cells, their subsets and activation status based on co-expression of the markers.

[0094] In some implementations, cell classification is performed using pathologist-validated thresholds. Following segmentation, cell classification is performed for identifying distinct cell phenotypes. In some implementations, a “pathologist-in-the-loop” approach to incorporate pathologist’s feedback directly during cell phenotype determination. Forty representative areas measuring 250,000 pm 2 were sampled from multiple regions of interest (ROIs) within the dataset and stitched together to form a montaged image. For each marker in the panel, a threshold for marker positivity was set by a pathologist (MPA) by adjusting and visualizing the positively identified cells in real time across the whole montaged image. In some examples, a pathologists’ knowledge of marker expression patterns is used to identify positive cells. Depending on the expected intracellular localization of each marker, thresholding was performedusing the averaged intensity across either the whole cell or specifically over the nucleus. Final thresholds were then combined into one composite classifier (example shown in FIG. 4D), applying each individual threshold sequentially to assign each cell a list of positive markers. A montage of all thresholded biomarkers from one ROI is shown in FIG. 4E. The proportion of cells within each tissue that was positive for immune markers of interest was added up for all regions associated with a particular clinical phenotype, providing a percentage positivity score. In some implementations, general characteristics of the immune influx in these distinct categories are determined (Fig. 4F).

[0095] In the example implementation, an imaging mass cytometry-based tissue classifier was developed. A deep learning-based renal microstructure segmentation algorithm is trained on pathologist annotations of glomeruli, tubules, interstitium and arteries, to incorporate the ability to assess cellular interactions with tissue compartments and spatial localization features at the single-cell level. First, approximately 3000 manually drawn annotations of glomeruli, tubules, interstitium and arteries, utilizing structural markers (alpha smooth muscle actin [SMA], vimentin, e-Cadherin, pan-cytokeratin, and NaK-ATPase) as a visual aide (FIG. 5A and 5B) were obtained. The distribution of annotated structures matched their abundance in the ROIs, with tubules, interstitium and arteries. A series of tissue pixel-based classification algorithms in QuPath were created and trained on 75% of all annotations, with the remaining 25% reserved for independent validation. The tissue classifiers varied in terms of markers used, feature convolutions, resolution subsampling scales and model framework (for example, random trees, artificial neural network, k-nearest neighbors, and linear regression). This created a suite of 13 classifiers to choose from. The model with the highest overall validation set accuracy across all microstructures was selected. The selected tissue classification algorithm was then applied across the entire dataset to generate structural annotations as an additional layer atop of our segmented and classified cells. An example of the optimal tissue classifier annotations is shown in FIG. 5C, and contrasted with the ground truth pathologist annotations in FIG. 5D. Each cell now inherited the property of the renal structure that its centroid fell within. This enabled the scoring immune cell subtypes within each renal structure (FIG. 5E and FIG. 5H) and the tissue-based spatial features were incorporated into the dataset. The tissue microstructural segmentation enabled localization of distinct immune cell populations of interest to various renal sub-compartments.

[0096] Some implementations include a single-cell spatial feature augmentation step. The proximity of certain cells to other cells of a given subtype, as well as the localization of such cells relative to tissue microstructures, is analyzed to identify higher-order patterns of tissue heterogeneity, and can be used to differentiate diseased versus healthy tissue. Two types of spatial features from these data were generated and used: (1) Cell classification-dependent, which includes cell distance analysis; and (2) utilizing the encoded categorical information of the cell classifications to compute the distance to the nearest cell of each base classification (e.g., CD3, 194 CD4, or vimentin single positive cell). In some implementations, Delaunay triangulation was used to allow for the averaging of all features across spatially connected cells. The second type of spatial feature was cell classification-independent, including smoothed features, where absolute intensities were averaged across neighboring cells iteratively over progressively larger radii to interrogate different scales of spatial heterogeneity. Finally, the distance of each cell to the border of the nearest renal microstructure was computed. Cumulatively, these spatial augmentations expand the number of features for each cell from 28 to over 2400, across the -900,000 cells present in this dataset.

[0097] A machine learning-based patient classifier using IMC was developed. The single-cell dataset generated using the techniques disclosed herein to develop a patient classifier capable of predicting the 5 different causes of graft inflammation (ACR, ABMR, ABMR*, BK, C. Pyel). Pathologist-assigned “ground truth” biopsy diagnoses, sometimes referred to herein as “clinical phenotypes”, were assigned. In some examples, the machine learning classifier trained on this data was an “Extreme Gradient Boosting” (XGBoost). XGBoost is a library implementing a regularized gradient boosting framework that is performance optimized, computer operating system-agnostic and can be implementable in several languages. To implement the XGBoost classifier into the pipeline, cells were assigned a categorical variable (encoded as an integer value) denoting the clinical phenotypes. Next, all cells from all patients belonging to the first batch of staining as were selected for a training set (46.5% of the total dataset), and the second batch as a validation set (53.5% of the total dataset). Additional training can also be performed. New batches of data will be generated as more patients are accrued. The training sets are selected to avoid contamination of the validation set. For example, contamination may occur during random partitioning when cells from a given patient are present in both training and validation sets, artificially inflating classifier accuracy. The classifier was trained on the spatial-augmentedfeature set to predict clinical phenotype, and then applied to the cells of the validation set, with accuracy defined as the percentage of correctly predicted. A mode-assigned integer-based label encoder is utilized within Scikit-learn package to determine what the most common predicted clinical phenotype was, first at the ROI and then at the patient level (see Fig. 6A).

[0098] The confusion matrix depicted in Fig. 6B highlights the accuracy of the trained classifier on the validation set at 84%. The algorithm predicted the correct patient diagnosis in 21 of the 25 patients. One patient with ABMR was misclassified as BK, 1 patient with BK was misclassified as ACR and two patients with chronic pyelonephritis were misclassified as BK. The high accuracy of the classifier is ascribed to not only capturing the mean intensity of marker expression across the tissues, but also to capturing intracellular spatial features, which relate to the distribution of marker expression patterns within a cell. Accuracy was also measured on the patient and cell level when the classifier training was limited only to the mean marker expression. In order to understand what features played the biggest role in the development of this patient classifier, the variable importance score was computed, the variable importance defining the relative importance of the feature to the model’s overall performance. In some examples, XGBoost was used to assign variable importance. Features can be classified based on the marker they were derived from (SMA, vimentin, CD3 etc.) or their spatial characteristics. Fig. 6C highlights the relative importance of all the markers used in the patient classifier, in which variable importance of all features (spatial and non-spatial) associated with a given marker was summed, to obtain the cumulative importance of each marker in the panel. The heatmap in Fig. 6D illustrates the distribution of MVF values across clinical phenotypes, revealing distinct differences used by the classifier when assigning the patient classification. To further interrogate the patient classifier on a per-marker basis, and to identify the minimum number of markers necessary to reach a particular accuracy level, a series of models were iteratively trained, starting with all features from the marker with the highest cumulative importance and working our way through subsequently lower importance markers. This generated a figure (FIG. 6E) indicating the stepwise relative contribution of each subsequent biomarker, and the contribution of non-spatial versus all features, towards this classifier accuracy. A representative spatial feature obtained in the MVF list is shown for two clinical phenotypes in Fig. 6F, alongside the single-channel intensity of the marker from which the spatial feature was obtained. A classifier utilizing only thenon-spatial features for that same marker was also trained to illustrate the relative importance of spatial features relative to marker number.

[0099] In some examples, identification of spatially resolved cell clusters by PhenoGraph was improved with manual gating of cells. Unsupervised clustering of single-cell data from the entire study sample was completed using PhenoGraph. A heat map showing the z-scored mean feature profile for each PhenoGraph cluster, colored by cluster identifier, is depicted in Fig. 7A. Singlecell feature data were visualized using the Barnes-Hut t-SNE data dimensionality reduction algorithm on the full feature set described above, reduced to two dimensions (Fig. 7B). 31 total clusters were identified. Five clusters had insufficient cells and were excluded from further analyses. Six clusters contained cells that had intermediate-to-low levels of expression for all markers studied. In the absence of a clear abundance of markers, these clusters could not be definitively named. The remaining were categorized into 11 immune and 9 non-immune (structural) clusters. The immune clusters included proliferation-dominant, lymphocyte-rich (T cell, B cell, 269 and mixed clusters), regulatory (FoxP3, Vista and PD-L1 expressing) and macrophage clusters. The pathologist-guided cell phenotyping allowed for a more specific nomenclature of the clusters. The proportional distribution of immune clusters across renal subcompartments in each biopsy sample of the various clinical phenotypes was analyzed. The stacked bar plots in Fig. 7C highlights the heterogeneous distribution of immune cellular clusters across the clinical phenotypes of allograft inflammation. Representative IMC images enriched with specific clusters are shown in Fig. 7D.

[0100] In some implementations, cellular neighborhoods and interaction networks of allograft inflammation were analyzed. Neighborhood analysis based on permutation tests to quantify celltype colocalization and identify statistically significant interaction or avoidance between pairs of cell clusters (Fig. 8 A). Structural (non-immune) clusters co-occurred frequently across many images, but positive interactions were few. Immune clusters showed variable co-occurrences across all images, with a few significant positive interactions. Significant positive interactions between clusters are depicted in Fig. 8A, as numbered boxes, and the interactions are illustrated in Fig. 8B. Immune cluster interactions (Fig. 8A: boxes 1-3), include those between memory T cells and lymphocyte (B cell)-rich clusters, proliferation-dominant clusters and macrophage / antigenpresenting cell clusters. The enriched interactions between CD45+ SMA+ (myofibroblast) clusters and cytokeratin+ vimentin+ clusters (epithelial-mesenchymaltransformation) (Box 5) was analyzed to determine that the latter is a key component of development of renal fibrosis and progression of kidney disease. The positive interaction between immune checkpoint regulator (VISTA) and epithelial-mesenchymal transformation clusters was found to indicate that VISTA plays a crucial role in preventing renal fibrosis.

[0101] In the example implementation, correlations between immune clusters’ with Banff histological lesions and graft outcomes were calculated. The Banff lesion scores assess histopathological changes in different microcompartments of the kidney. For example, glomeruli, tubules, interstitium and blood vessels. The scores focus on immune cell infiltration in these compartments. For example, tubular inflammation (tubulitis) is scored based on the number of infiltrating immune cells present in a tubule. Given the more detailed information derived in a spatially relevant format obtained through the disclosed methods, the association between the Banff scores and the IMC data is correlated. The Banff scores that pertain to the tubulointerstitial compartments (tubular inflammation [t], interstitial inflammation in non-scarred tissues [i], total inflammation in renal parenchyma [ti], chronic interstitial fibrosis [ci] and chronic tubular atrophy [ct]) correlated with the IMC data to demonstrate a proportional increase in immune clusters. One example advantage of the method is the ability to identify immune clusters that are localized to compartments besides the specific lesional compartment studied. For example, in tubulitis, a significant increase in proportion of macrophage / antigen-presenting cell (APC) clusters and memory T cell clusters were found, not only in 311 tubules, but in interstitium and glomeruli as well (Fig. 8C and FIG. 8D). Furthermore, the IMC data was integrated with clinic information to identify that compared to those with preserved graft function and similar Banff scores, grafts that later failed had a significantly higher proportion of epithelial-mesenchymal transformation, and lower proportions of ATPase expressing tubules, PD-L1 expressing tubules and glomeruli expressing larger amounts of vimentin. Grafts enriched in memory T cells had poorer outcomes. Fig. 8E shows representative images of the two clusters which are associated with poor graft outcome.

[0102] Example implementation 1 discussion.

[0103] The example implementation includes an integrated image analysis pipeline that takes IMC regions of interest from several patient biopsies, and adapts a sequence of analytical tools to localize and identify cells, extract useful spatial features, and use this information to predict clinical phenotypes. The StarDist cellular segmentation algorithm is optimized using USDQP,which provides considerable robustness for modality-agnostic segmentation of circular objects across a broad variety of imaging modalities. USDQP implements several pre- and postprocessing tools to enable models trained on DAPI, hematoxylin, or other DNA counterstained images, to be applied to IMC data directly: First, USDQP uses scale-agnostic resampling to resize individual image tiles from IMC images prior to being passed to the model, such that the resampled IMC resolution can mimic what the model had been trained at. A user-adjustable linear normalization of values is used prior to segmentation to correct for IMC data that has an intensity-based dynamic range that differs greatly from fluorescent DAPI images. Cell simulation can be performed through cell neighbor-constrained dilation of segmented nuclei. Third, in addition to extending StarDist to IMC data, USDQP is configured to allow stain separation of brightfield images, allowing DAPI-based models to be used in segmenting H&E and HD AB images. Finally, USDQP can also utilize GPU-acceleration on versions of QuPath built with Tensorflow or OpenCV libraries, thus presenting a robust method for modalityagnostic segmentation of circular objects across a broad variety of imaging modalities.

[0104] Manual gating methods are used on a montaged image, with real-time gating feedback for cell type identification and classification. This incorporated pathologist expertise in biomarker localization on a per-marker basis, as is routinely done in conventional IHC staining to separate true signal from background. Each individual manual gate identifies cells that are positive or negative for that marker. Individual marker thresholds combine into a multi-marker classification that allows for more complex immune phenotyping. This pathologist-guided cell phenotyping can be utilized both for the generation of spatial features, as well as for improving the specificity of cluster nomenclature during the process of dimensionality reduction and phenograph clustering. Additional training of a tissue-based classifier permitted the breakdown of the tissue into its microstructural components, enabling the development of a highly accurate IMC-derived patient classifier of renal allograft inflammation. The XGBoost patient classifier developed is scalable and adaptable across varied organ-systems and diverse disease categories.

[0105] Advantages of the disclosed tools include in-depth spatial characterization. In some examples, spatially resolved subcellular biomarker distribution improved the accuracy of the machine learning algorithm in predicting patient-level classification. By sequentially sorting markers by feature importance, and then training a set of classifiers on each marker individually, additionally discovered spatial features significantly boosted the cell-level accuracy of theresulting patient classifier. In fact, these spatial features permit the achievement of high classifier accuracy with a limited panel of markers (10 markers with spatial features reach equivalent accuracy of 27 markers with non-spatial features). The top ten features identified by the ML model to play a role in accurate patient-level classification are biologically meaningful. For instance, dendritic cells (CDl lc+), that infiltrate transplanted organs may sustain the alloimmune responses after T-cell activation has already occurred. Similarly, CD8+ 367 T cells play an may play a role in rejection. The analysis in the example implementation identified the heterogeneity within clinical phenotypes. Analysis in the example implementation identified specific immune clusters correlate with unique Banff histological lesions, and that certain cell clusters correlate with poor graft outcome. The disruptive methodologies applied, even on limited needle biopsies, permit biologically insightful and hypothesis-generating data creation.

[0106] In some examples, the panel of biomarkers was identified ahead of time, meaning that novel proteins that might not be thought to play a role in this process would not be identified. USDQP lends itself as a powerful tool for segmentation, but can be further configured to include a StarDist trained on IMC-annotated cells. In some examples, variations in staining conditions across multiple IMC batches can lead to poorer segmentations if the images are dissimilar to what the model was trained on. In some examples, pathologist-guided classification were used within a context where staining was consistent between batches of images. However, additionally processing may adjust to account for training examples that do not have uniform staining due to differences in pre-analytical steps (panel design, differences in fixation and embedding, etc.). The use of IMC in evaluating renal allograft inflammation using the systems and methods disclosed herein allow for a systematic, multidimensional interrogation of renal allograft histology and has generated a detailed spatial map of cell clusters and their relationships with clinical phenotypes. Applying the single-cell-derived data to histological and clinical databases, to be able to arrive at biologically and prognostically relevant information. The example implementation explored the single-cell landscape of renal allografts from a biological, therapeutic, and prognostic perspective. Some of the methods adopted can be configured to operate on various tissue types, including tissues with limited biopsy material, with widespread applicability in clinical, experimental and pharmaceutical fields.

[0107] FIGs. 4A-4F related to cell segmentation using Universal StarDist for QuPath (USDQP). FIG. 4A illustrates an Ir channel corresponding to the DNA counterstain, used for segmentationof nuclei. FIG. 4B illustrates nucleus segmentation and whole cell simulation using Universal StarDist for QuPath. FIG. 4C illustrates multi-channel visualization of most structural and immune-related markers present in panel. FIG. 4D illustrates composite classification of all single positive classes combined and applied sequentially. FIG. 4E illustrates cells classified as positive for each marker present in the panel, using pathologist-verified thresholds. From top left to bottom right, the markers are as follows: alpha-SMA, CD19, Vimentin, CD14, CD16, Pan- Cytokeratin, CDl lb, PD-L1, CD45, CDl lc, FoxP3, CD4, E-Cadherin, CD68, Vista, CD20, CD8a, CD45RA, Granzyme-B, Ki67, Collagen-I, CD3, Histone-H3, CD45RO, HLA-DR, Beta- 2M, NaK-ATPase, and DNA intercalator counterstain. FIG. 4F illustrates percentage of cells positive for each immune marker (x-axis), grouped by clinical phenotype (y-axis). White scale bar corresponds to 100 pm, in both small and large views of image.

[0108] FIGs. 5A-5H relate to cell classification by pathologist-validated binary thresholds. FIG. 5A illustrates pseudocolored image with alpha-SMA (red), Vimentin (blue), Pan-Cytokeratin (yellow), Ecadherin (cyan) and the Ir DNA counterstain (white) provided to pathologists for annotation. FIG. 5B illustrates example of a pathologist’s annotations for glomeruli (blue), arteries (red), tubules (green), Interstitium (magenta), and empty space (grey) used for training the renal structure pixel classifier. FIG. 5C illustrates predicted pixel classification from trained classifier. FIG. 5D illustrates a pixel classifier output (translucent) from trained classifier, with manual pathologist’s annotations (outlined). FIGs. 5E-H illustrate percent positive scoring for various immune markers, grouped by disease subtype. Renal structure shown with yellow outline (top) and profiled for various immune markers (bottom). FIGs, 5E-H show glomeruli, arteries, interstitium, and tubules, respectively.

[0109] FIGs. 6A-F relate to development of an Imaging Mass Cytometry-based tissue classifier. FIG. 6A illustrates a pictographic representation of patient classification using mode-based data aggregation. FIG. 6B illustrate a confusion matrix representing ground truth (y axis) versus predicted (x axis) label. Values correspond to the number of patients in each ground truth- predicted pair. Accuracy is measured as the sum of correctly predicted labels (diagonal) divided by the total number of patients. FIG. 6C illustrates a sum of all feature importances, grouped by the marker they are derived from. Markers were sorted from highest to lowest cumulative importance. FIG. 6D illustrates a heatmap of the top 10 most valuable features (MVF) after z- score normalization, grouped by clinical phenotype type. Importance is the relative importanceof the feature, as extracted from the model’s parameters. FIG. 6E illustrates a cumulative model performance when incrementally including all features of a given marker, with the order determined through Fig. 6C. Each point includes all features associated with that marker, as well as all features from markers to the left of the point. The blue line indicates model performance when only including non-spatially dependent features, whereas the orange includes spatially dependent features. FIG. 6F illustrates a visualization of MVF5 (Cluster mean: Gd (158)_158Gd- E Cadherin: Cytoplasm: Median) on an ROI from a patient with ABMR (left) versus ABMR*(right). The left side of each subfigure denotes the original image (blue=DAPI, yellow=E-Cadherin), and the right side contains an overlaid measurement map for MVF5.

[0110] FIGs. 7A-D relates to single cell phenotypes in high dimensional histopathology of renal allograft inflammation. FIG. 7A illustrates a heat map showing the z-scored mean marker expression for each PhenoGraph cluster, colored by cluster identifier. The absolute cell counts of each PhenoGraph cluster are displayed as a bar plot (left). FIG. 7B illustrates a map using t- distributed stochastic neighbor embedding (t-SNE) of 900,000 single cells from highdimensional images of renal allograft inflammation colored by cell-type. FIG. 7C illustrates a stacked bar plots depicting the proportion of immune cell clusters identified in each case, distributed in the renal microstructural compartments, and clinical phenotypes. FIG. 7D illustrates a panel of pseudo-colored IMC images depicting 3 cell clusters.[0U1] FIG.8A-E relate to a global map of the cellular neighborhoods and interaction networks in renal allograft inflammation. FIG. 8A illustrates a heat map in which squares indicate the Pearson correlation of cell phenotype proportions across all measured tissue regions and circles indicate significant pairwise cell-type interaction or avoidance summarized across the two-sided permutation tests on the individual images. Circle color indicates the percentage of images and size represents the number of images with a significant cell-cell interaction or avoidance (P < 0.01). Significant interactions are depicted in numbered boxes 1-10. FIG. 8B shows an illustration highlighting clusters in Fig.8A that showed significant interactions associated with poor graft outcomes, for example, clusters depicting epithelial mesenchymal transformation and memory T cells. The illustration suggests cluster interactions that depict biological progression (i: early to ii: late). FIG. 8C illustrates a boxplot showing identification of immune clusters that spatially localize in different microcompartments of the kidney, namely tubules, interstitium and glomeruli while correlating with certain Banff lesion scores. FIG. 8C shows how increasingscores of tubulitis significantly correlate with higher proportions of macrophage / antigen presenting cell (APC) clusters which are spatially distributed in tubules, interstitium and glomeruli. Data in boxplots are presented by minimum, 25th percentile, median, 75th percentile and maximum. Values outside of 1.5 times 14 interquartile range are classified as outliers and are denoted as fliers. P < 0.05, two-sided Mann Whitney U-test, Benjamini -Hochberg adjusted. FIG. 8D illustrates pseudocolored IMC images showing 2 clusters (memory T cells and macrophage / APC clusters) that are seen in increasing proportions with higher scores of Banff tubulitis. The annotations allow us to see the localization of the clusters to the interstitial compartment. FIG. 8E: illustrate two pseudocolored IMC images that show two cell clusters (epithelial mesenchymal transformation and memory T cells) that in higher proportion correlated with poor graft outcome.Example Implementation 2:

[0112] FIGs. 9-12 refer to example implementation 2. The example implementation 2 incorporates machine learning and artificial intelligence algorithms into imaging mass cytometry of renal allograft inflammation to develop novel diagnostic and prognostic tools.

[0113] Example implementation 2 introduction.

[0114] Among kidney transplant patients, inflammation of the allograft is a cause of grafts failure. Kidney allograft inflammation can be secondary to both alloimmune and infectious stimuli. For example, in a large cohort (n=5752) of kidney transplant recipients, we have demonstrated that inflammation due to alloimmune injury accounted for 39% of graft failure and BK nephropathy, another cause of graft inflammation, for 3%. Allograft inflammation in areas of allograft fibrosis is strongly associated with death-censored graft failure when compared to recipients whose biopsies had no inflammation, even after adjusting for the presence of interstitial fibrosis. Furthermore, total allograft inflammation, is a robust predictor of death- censored allograft survival. Imaging mass cytometry (IMC) allows for multiplexed immunohistochemistry staining of tissue. Upwards of 40 markers can be simultaneously stained, acquired and visualized, enabling a variety of distinct cell types to be analyzed concurrently in their native microenvironment. IMC provides spatial data for many parameters at subcellular resolution. In this implementation, IMC data interrogating allograft inflammation with machine learning models is used develop a tissue classifier. Computational and artificial intelligencemodels disclosed herein are used to identify novel cell clusters predictive of poor clinical outcomes.

[0115] Example implementation hypothesis.

[0116] Application of unsupervised machine learning and artificial intelligence algorithms to imaging mass cytometry data of renal allograft inflammation can provide spatially resolved, single cell data that improves current tissue classification, identifies novel biomarkers, and predicts prognostic outcomes.

[0117] Example Implementation results

[0118] A tissue-based classifier is of allograft inflammation to improve diagnostic accuracy in biopsies with inflammation of unclear etiology (i-IFTA) is develop and validate. In this example, application of imaging mass cytometry (EMC) allows for accurate tissue diagnosis of conventional phenotypes of allograft inflammation and can improve diagnostic accuracy in challenging biopsies of “non-specific” inflammation in areas of interstitial fibrosis and tubular atrophy. Forty different biopsies across clinically defined categories of allograft inflammation were analyzed, using 27 metal-labeled antibodies. This data was used with machine learning methods to develop a tissue classifier with an accuracy of 80%. Further refinement, includes a) increase the training and validation biopsy data set to 300 biopsies and expand the panel of metal-labeled antibodies (50 antibodies) to improve on the current diagnostic accuracy of the classifier, b) test the classifier across different populations and centers, and c) apply the developed classifier to phenotype challenging cases of “non-specific” inflammation in areas of interstitial fibrosis and tubular atrophy.

[0119] A prognostic tool for allograft outcome integrating clinical, serological and morphological data with spatially resolved single cell data is developed._In this example, Singlecell pathology phenotype clusters improve the ability to predict overall survival of renal allografts, as compared to current conventional morphological categories of allograft inflammation. The current single cell pathology database is expanded by increasing the number of samples using clinically and pathologically well-characterized biopsies from a large transplant biopsy archive, and expand the panel of antibodies tested to a) identify single cell phenotypic clusters that correlate with graft loss, and b) develop and validate a prognostic model for graft loss, using machine learning models that incorporate clinical, serological and novel single cell pathology data.

[0120] In this example, an unsupervised machine learning models to discover novel immune and stromal biomarkers in renal allograft inflammation that improve characterization and prognostication of renal allograft inflammation is used. Here, using unsupervised clustering of single cell pathology data to identify subgroups that better inform on graft outcome, as compared to currently defined clinicopathological classifications. IMC data is used on 40 biopsies and performed unsupervised analysis by computational machine learning models to identify single cell pathology clusters that correlate with adverse graft outcomes. By enriching the biopsy cohort allows for systems and methods to a) analyze the immune, epithelial and stromal composition of single cell pathology clusters that correlate with adverse graft loss, b) compare the immune, epithelial and stromal composition of these phenotypic clusters to those of currently defined clinical subtypes of allograft inflammation to detect novel immune and stromal biomarkers that correlate with poor allograft outcomes, and c) validate these biomarkers on tissue sections using immunohistochemistry across different biopsy populations.

[0121] In this example, multiplexed immunohistochemistry data sets are used to train an artificial intelligence algorithm to improve on and generate the requirements of the Banff classification scheme. In one example, a pathologist-supervised end-to-end workflow for renal allograft reporting using multiplexed immunohistochemistry data that improves on the Banff classification is developed. Using a large data base of pathologist-annotated multiplexed immunohistochemistry images to generate biopsy reports that a) quantify glomeruli, tubules and blood vessels, b) compute total inflammation, inflammation in glomeruli and blood vessels, and c) develop novel predictive tools such as peritubular capillary density, podocyte density. These reports can be applied for diagnostic and research purposes in clinical trial and biopharma.

[0122] Example implementation 2 approach and methods.

[0123] A tissue-based classifier of allograft inflammation to improve diagnostic accuracy in biopsies with inflammation of unclear etiology (i-IFTA) is developed and validated. First it was established whether machine learning model using IMC-derived features can accurately classify well-phenotyped causes of allograft inflammation, for example, acute T cell-mediated rejection (ACR) or acute antibody mediated rejection (AB MR) with enough accuracy to be useful to predict causes of inflammation in areas of interstitial fibrosis and tubular atrophy (i-IFTA). In example implementation 2, a classifier to predict clinical categories of allograft inflammation is developed (Figure 9). For this, an artificial intelligence model is trained to learn the associationsbetween cellular features and different types of allograft inflammation. The machine learning model computes the variable importance of each feature used in the training. Those with a higher importance are more crucial for the model to achieve its high accuracy. Variable importance allowed us to evaluate which markers to include as part of the gene signature, based on the cumulative importance of all features derived from a given marker. A 25 / 75 test-train split was performed, with a fixed seed state to preserve model reproducibility between multiple successive runs, z-score normalization is applied to the X train scaled and X test scaled variables. The model to train was regularized gradient boosting model XGBoost. A fixed seed state is used, and tree construction was performed using the GPU-accelerated histogram method for a considerable increase in model performance, while also yielding a slight benefit in accuracy. Variable importances were obtained by extracting them from the trained model, and each feature has a corresponding importance score. Accuracy was evaluated by applying the trained model to a test set. The accuracy of our trained classifier on a validation set was 80%. To improve the diagnostic accuracy of the classifier, the training and validation biopsy data set was increased to 300 biopsies and expand the panel of metal-labeled antibodies (50 antibodies). The developed classifier was applied to phenotype challenging cases of “non-specific” inflammation in areas of interstitial fibrosis and tubular atrophy (i-IFTA).

[0124] In one exampl e,_based on a 50:50 train / test split, characterized a power by precision in the estimate of overall accuracy in a test set size of 50 samples. For a 95% confidence level, the half-width would be <0.10 when the true accuracy is >0.90.

[0125] A prognostic tool for allograft outcome integrating clinical, serological and morphological data with spatially resolved single cell data was developed. First, it was analyzed whether unsupervised analysis of spatially resolved single cell pathology data permits identification of single cell pathology subgroups / clusters that associate with poor graft outcome, with the intention of integrating the spatial data to clinical and serological data to develop a prediction tool for allograft loss. An unsupervised cell clustering of single-cell IMC intensity data and morphology characteristics using the network-based clustering method PhenoGraph to evaluate the IMC data of allograft inflammation. This was successful in identifying 31 latent cell types, which were enumerated by region of interest (ROI, e.g., glomeruli, tubules) as a subjectlevel feature set. Figure 10 is a heatmap showing the z-scored mean feature value for each of the 31 clusters identified by PhenoGraph, colored and numbered by cluster identifier. Features areon the x-axis and labeled by feature grouping, while clusters are on the y-axis. The adjacent plots to the left indicate the relative proportions of each cell-type cluster by sample type, as well as the total number of cells (far left). To assess the collective prognostic ability of clinical, morphological, and single-cell IMC features on graft loss, we will fit flexible tree-base ensemble machine learning models via random forest (RF). RF is a highly flexible and robust tree-based classification algorithm that can readily accommodate high-dimensional data with non-linear relationships and interaction effects related to case status. Hyperparameter tuning of the RF model (e.g., tree depth) will be conducted using repeated k-fold cross-validation which can readily handle mixed variable types and non-linear relationships with risk. Performance will be assessed using area under the receiver operating characteristic curve (AUC), and cross-validated AUC’s and corresponding 95% Cl’s will be estimated using the cvAUC R package. Multivariable models will be fit with and without IMC features to further examine potential improvement in discrimination when combining these results, and bootstrapping methods will be used to quantify uncertainty in the difference in performance. Additionally characterizations of diagnostic performance metrics such as sensitivity, specificity, PPV, and NVP at clinically relevant thresholds are determined.

[0126] In one example, a target sample size of N = 100 and assuming a graft loss rate of -25%, to frame statistical power in the context of a precision of the AUC for a given prognostic model. Under these conditions, the expected lower bound of the 95% CI width will -0.80 when the true AUC is 0.90 under a simple 50:50 split. Use of nested cross-validation to derive cross-vali dated AUCs further improve precision of these estimates.

[0127] In this example implementation, unsupervised analysis of IMC data to spatial cell interactions is extended to discover novel biomarkers in renal allograft inflammation that improve characterization of the subtypes of renal allograft inflammation. First, biomarkers of immune, epithelial and stromal lineage that associate with clinically defined and novel subtypes of allograft inflammation, with allograft outcomes and are detectable with IHC on kidney biopsy are found. During preliminary analysis, a supervised data analytical approach of a clinically defined cohort. For each of the IHC markers, the cellular sub compartment (cell or nuclear) to use was identified and each threshold was set with a renal pathologist (MPA) on a per-batch basis. A composite classifier was generated, effectively assigning every segmented cell to one or more IHC marker classes. Cellular and annotation-level measurements were exported in csvformat. To quantify the percentage of cells positive for multiple markers, percent multiple _positive.py was developed to query the dataset to identify cells positive for one or more specific markers. FIG. 11 illustrates the percentage positivity of specific immune subpopulations across the study population standardized within subpopulation using z scores. Color scale shows z-score per row. (A) in ABMR. (B) in the glomerulus in the different categories of inflammation and (C) in all regions of interest for the various categories of conventionally classified inflammation.

[0128] The unsupervised analyses is extended to further extract higher order features and test associations with allograft inflammation subtypes. In addition to ROI-wise cell-cluster measures previously mentioned, are used to further exploit spatial information of cell locations and cluster assignments. First, evidence of cluster-wise cell-to-cell interaction (i.e., colocalization and avoidance) within ROI using histoCat is identified, with permutations to identify significant interactions within sample agnostic to subtype. Interactions with sufficient heterogeneity (>5% or <95%) will be considered as candidate feature for association with subtype. Similarly, spatial context analysis is performed to identify interactions of cellular neighborhoods, which can capture more complex cellular phenomena. Formal statistical association analysis of interaction features which were tested using chi-square testing, while enumeration of cell proportions belonging to particular spatial context motifs will be tested using nonparametric Kruskal-Wallis testing. Similarly, cell-cluster proportions and densities per ROI were tested with subtype using Kruskal-Wallis. Significant associations with be declared at a false discovery rate of 0.05.Candidate markers that correlate with inflammation subtype and graft outcome are then validated by standard immunohistochemistry.

[0129] In this example, multiplexed immunohistochemistry data sets are used to train an artificial intelligence algorithm to improve on and generate the requirements of the Banff classification scheme. To use data generated above to develop a pathology reporting tool that improves on the Banff classification. Pathologist annotated multiplexed immunohistochemistry data sets are used to develop artificial intelligence algorithms that can generate glomerular and artery counts per regions of interest (Figure 12). Additionally, scores on total immune composition per region of interest are generated, as well immune composition per renal microstructural compartments. Using the tissue classifier developed in reference to FIG. 9, and the prognostic tool referenced in FIG. 10, we will be able to add novel measures to improve onthe current Banff - allograft report (FIG. 12). With an expanded panel of immune and renal specific metal tagged antibodies, to generate prognostic scores including capillary density, podocyte density, etc.

[0130] In this example, artificial intelligence algorithms in the analysis of IMC -based multiplexed immunohistochemistry images of renal allograft inflammation allows accurate phenotypic classification of allograft inflammation permitting early and appropriate therapy. The prerequisite for successful treatment is an accurate diagnosis, a concept coined precision medicine. The use of imaging mass cytometry. The IMC classifier, developed herein, provides the information that allow disease specific therapeutic intervention. This classifier can be utilized in the evaluation of complex transplant biopsies and as part of a system to serve as a central biopsy review center for clinical trials focused on renal allograft inflammation.

[0131] In some examples, leveraging machine learning models in the unsupervised analysis of spatially resolved single cell data allows for the identification of novel single cell pathology subgroups that better inform on prognosis of allograft inflammation. The applications of IMC in breast cancer have revealed multicellular features of the tumor microenvironment and novel subgroups of breast cancer that are associated with distinct clinical outcomes, characterizing intratumor phenotypic heterogeneity in a disease-relevant manner informing patient-specific diagnosis. The disclosed applications of IMC to the analysis of renal allograft inflammation permits discovery of single cell pathology subgroups that are associated with clinical outcomes, in such serving as novel predictive biomarkers. These biomarkers can be used to developed prognostic signature sets.

[0132] Training of artificial intelligence algorithms on pathology annotated multiplexed immunohistochemistry data sets allows for the development pathology reporting tools that replicate and improve on the current Banff classification. The use of automated reporting tools allows for enhanced standardization of biopsy interpretation, and reproducibility of indices such as glomerular and arterial numbers. Additionally, the use of a multiplexed panel of antibodies which incorporate a broad panel of immune and renal specific markers allows for the development and automation novel prognostic indices such as immune subpopulation specific compositional scores, peritubular capillary density, IFTA scores and podocyte density scores.Example Implementation 3:

[0133] FIGs. 13A-D refer to example implementation 3. Example implementation 3 includes development of a tissue-based classifier of allograft inflammation using imaging mass cytometry. FIG. 13A illustrates an analysis flowchart. FIG. 13B illustrates heat maps of immune markers. FIGs. 13C and FIG. 13D illustrate example results of the example implementation 3.

[0134] Background of example implementation 3.

[0135] Molecular phenotyping of allograft inflammation is used to improve both diagnostic accuracy and understanding of the heterogeneity of rejection. Current molecular techniques lack histological correlation & spatial dimensionality. Imaging Mass Cytometry (IMC) is used to develop a tool to accurately predict the cause of allograft inflammation.

[0136] Example implementation 3 methods.

[0137] A cohort included biopsies of rejection, BK nephropathy, pyelonephritis & normal kidneys. Using a panel of 28 markers, IMC images were processed by the Hyperion imaging system. Details of analysis are in Fig 13A. Cell segmentation was performed using Universal StarDist for Qupath. Cell classification was performed based on mean intensity threshold.

[0138] Example implementation 3 results.

[0139] 139 regions of interest (ROI) were processed. Violin plots ensured there were measurable differences in known markers associated with each allograft inflammation category. Distribution of percent positive scoring of immune cells are seen in the heatmap (FIG. 13B) [e.g. : cellular & mixed rejection cases enriched in CD45+, HLA-DR+ cells and CD4+ memory T cells]. The trained regularized gradient boosting classifier model XGBoost was used to predict the allograft inflammation category for all cells, ROIs and each original histological diagnosis. (Fig 13C & FIG. 13D). The trained model accurately predicted the allograft inflammation category for each cell with an accuracy of 64.3%. When using the mean intensity parameter of each cell, the classifier accuracy improved to 87.8% in predicting the type of renal allograft inflammation, independent of ROI. The accuracy improved to 90.9% when dimension of intracellular spatial features (proximity metrics) were added to the algorithm. Granzyme, CD68 and Vista were the three most important markers in achieving this high accuracy.

[0140] Example implementation conclusion.

[0141] Highly multiplexed imaging of renal allograft biopsies with subcellular resolution by IMC were used to develop a classifier of allograft inflammation, which demonstrates high diagnostic accuracy.Example Computing Devices:

[0142] FIG. 14 shows an example of a computing device 900 and a mobile computing device that can be used to implement computer-based aspects of the techniques described herein. The computing device 900 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The mobile computing device is intended to represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smart-phones, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be exemplary only, and are not meant to limit implementations of the inventions described and / or claimed in this document.

[0143] The computing device 900 includes a processor 902, a memory 904, a storage device 906, a high-speed interface 908 connecting to the memory 904 and multiple high-speed expansion ports 910, and a low-speed interface 912 connecting to a low-speed expansion port 914 and the storage device 906. Each of the processor 902, the memory 904, the storage device 906, the high-speed interface 908, the high-speed expansion ports 910, and the low-speed interface 912, are interconnected using various busses, and may be mounted on a common motherboard or in other manners as appropriate. The processor 902 can process instructions for execution within the computing device 900, including instructions stored in the memory 904 or on the storage device 906 to display graphical information for a GUI on an external input / output device, such as a display 916 coupled to the high-speed interface 908. In other implementations, multiple processors and / or multiple buses may be used, as appropriate, along with multiple memories and types of memory. Also, multiple computing devices may be connected, with each device providing portions of the necessary operations (for example, as a server bank, a group of blade servers, or a multi-processor system).

[0144] The memory 904 stores information within the computing device 900. In some implementations, the memory 904 is a volatile memory unit or units. In some implementations, the memory 904 is a non-volatile memory unit or units. The memory 904 may also be another form of computer-readable medium, such as a magnetic or optical disk.

[0145] The storage device 906 is capable of providing mass storage for the computing device 900. In some implementations, the storage device 906 may be or contain a computer-readablemedium, such as a floppy disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid state memory device, or an array of devices, including devices in a storage area network or other configurations. The computer program product may also contain instructions that, when executed, perform one or more methods, such as those described above. The computer program product can also be tangibly embodied in a computer- or machine-readable medium, such as the memory 904, the storage device 906, or memory on the processor 902.

[0146] The high-speed interface 908 manages bandwidth-intensive operations for the computing device 900, while the low-speed interface 912 manages lower bandwidth-intensive operations. Such allocation of functions is exemplary only. In some implementations, the high-speed interface 908 is coupled to the memory 904, the display 916 (for example, through a graphics processor or accelerator), and to the high-speed expansion ports 910, which may accept various expansion cards (not shown). In the implementation, the low-speed interface 912 is coupled to the storage device 906 and the low-speed expansion port 914. The low-speed expansion port 914, which may include various communication ports (for example, USB, Bluetooth, Ethernet, wireless Ethernet) may be coupled to one or more input / output devices, such as a keyboard, a pointing device, a scanner, or a networking device such as a switch or router, for example, through a network adapter.

[0147] The computing device 900 may be implemented in a number of different forms, as shown in the figure. For example, it may be implemented as a standard server 920, or multiple times in a group of such servers. In addition, it may be implemented in a personal computer such as a laptop computer 922. It may also be implemented as part of a rack server system 924. Alternatively, components from the computing device 900 may be combined with other components in a mobile device (not shown), such as a mobile computing device 950. Each of such devices may contain one or more of the computing device 900 and the mobile computing device 950, and an entire system may be made up of multiple computing devices communicating with each other.

[0148] The mobile computing device 950 includes a processor 952, a memory 964, an input / output device such as a display 954, a communication interface 966, and a transceiver 968, among other components. The mobile computing device 950 may also be provided with a storage device, such as a micro-drive or other device, to provide additional storage. Each of theprocessor 952, the memory 964, the display 954, the communication interface 966, and the transceiver 968, are interconnected using various buses, and several of the components may be mounted on a common motherboard or in other manners as appropriate.

[0149] The processor 952 can execute instructions within the mobile computing device 950, including instructions stored in the memory 964. The processor 952 may be implemented as a chipset of chips that include separate and multiple analog and digital processors. The processor 952 may provide, for example, for coordination of the other components of the mobile computing device 950, such as control of user interfaces, applications run by the mobile computing device 950, and wireless communication by the mobile computing device 950.

[0150] The processor 952 may communicate with a user through a control interface 958 and a display interface 956 coupled to the display 954. The display 954 may be, for example, a TFT (Thin-Film-Transistor Liquid Crystal Display) display or an OLED (Organic Light Emitting Diode) display, or other appropriate display technology. The display interface 956 may comprise appropriate circuitry for driving the display 954 to present graphical and other information to a user. The control interface 958 may receive commands from a user and convert them for submission to the processor 952. In addition, an external interface 962 may provide communication with the processor 952, so as to enable near area communication of the mobile computing device 950 with other devices. The external interface 962 may provide, for example, for wired communication in some implementations, or for wireless communication in other implementations, and multiple interfaces may also be used.

[0151] The memory 964 stores information within the mobile computing device 950. The memory 964 can be implemented as one or more of a computer-readable medium or media, a volatile memory unit or units, or a non-volatile memory unit or units. An expansion memory 974 may also be provided and connected to the mobile computing device 950 through an expansion interface 972, which may include, for example, a SIMM (Single In Line Memory Module) card interface. The expansion memory 974 may provide extra storage space for the mobile computing device 950, or may also store applications or other information for the mobile computing device 950. Specifically, the expansion memory 974 may include instructions to carry out or supplement the processes described above, and may include secure information also. Thus, for example, the expansion memory 974 may be provide as a security module for the mobile computing device 950, and may be programmed with instructions that permit secure use of themobile computing device 950. In addition, secure applications may be provided via the SIMM cards, along with additional information, such as placing identifying information on the SIMM card in a non-hackable manner.

[0152] The memory may include, for example, flash memory and / or NVRAM memory (nonvolatile random access memory), as discussed below. The computer program product contains instructions that, when executed, perform one or more methods, such as those described above. The computer program product can be a computer- or machine-readable medium, such as the memory 964, the expansion memory 974, or memory on the processor 952. In some implementations, the computer program product can be received in a propagated signal, for example, over the transceiver 968 or the external interface 962.

[0153] The mobile computing device 950 may communicate wirelessly through the communication interface 966, which may include digital signal processing circuitry where necessary. The communication interface 966 may provide for communications under various modes or protocols, such as GSM voice calls (Global System for Mobile communications), SMS (Short Message Service), EMS (Enhanced Messaging Service), or MMS messaging (Multimedia Messaging Service), CDMA (code division multiple access), TDMA (time division multiple access), PDC (Personal Digital Cellular), WCDMA (Wideband Code Division Multiple Access), CDMA2000, or GPRS (General Packet Radio Service), among others. Such communication may occur, for example, through the transceiver 968 using a radio-frequency. In addition, short-range communication may occur, such as using a Bluetooth, WiFi, or other such transceiver (not shown). In addition, a GPS (Global Positioning System) receiver module 970 may provide additional navigation- and location-related wireless data to the mobile computing device 950, which may be used as appropriate by applications running on the mobile computing device 950.

[0154] The mobile computing device 950 may also communicate audibly using an audio codec 960, which may receive spoken information from a user and convert it to usable digital information. The audio codec 960 may likewise generate audible sound for a user, such as through a speaker, for example, in a handset of the mobile computing device 950. Such sound may include sound from voice telephone calls, may include recorded sound (e.g., voice messages, music files, etc.) and may also include sound generated by applications operating on the mobile computing device 950.

[0155] The mobile computing device 950 may be implemented in a number of different forms, as shown in the figure. For example, it may be implemented as a cellular telephone 980. It may also be implemented as part of a smart-phone 982, personal digital assistant, or other similar mobile device.

[0156] Various implementations of the systems and techniques described here can be realized in digital electronic circuitry, integrated circuitry, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0157] These computer programs (also known as programs, software, software applications or code) include machine instructions for a programmable processor, and can be implemented in a high-level procedural and / or object-oriented programming language, and / or in assembly / machine language. As used herein, the terms machine-readable medium and computer-readable medium refer to any computer program product, apparatus and / or device (e.g., magnetic discs, optical disks, memory, Programmable Logic Devices (PLDs)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term machine-readable signal refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0158] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0159] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middlewarecomponent (e.g., an application server), or that includes a front end component (e.g., a client computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0160] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.

[0161] In situations in which the systems, methods, devices, and other techniques here collect personal information (e.g., context data) about users, or may make use of personal information, the users may be provided with an opportunity to control whether programs or features collect user information (e.g., information about a user’s social network, social actions or activities, profession, a user’s preferences, or a user’s current location), or to control whether and / or how to receive content from the content server that may be more relevant to the user. In addition, certain data may be treated in one or more ways before it is stored or used, so that personally identifiable information is removed. For example, a user’s identity may be treated so that no personally identifiable information can be determined for the user, or a user’s geographic location may be generalized where location information is obtained (such as to a city, ZIP code, or state level), so that a particular location of a user cannot be determined. Thus, the user may have control over how information is collected about the user and used by a content server.

[0162] Although various implementations have been described in detail above, other modifications are possible. In addition, the logic flows depicted in the figures do not require the particular order shown, or sequential order, to achieve desirable results. In addition, other steps may be provided, or steps may be eliminated, from the described flows, and other components may be added to, or removed from, the described systems. Accordingly, other implementations are within the scope of the following claims.

Claims

CLAIMSWhat is claimed is:

1. A method of predicting causes of allograft inflammation, the method comprising: obtaining a training data set that comprises a collection of training samples, each training sample corresponding to a particular cell or a particular cluster of cells in one of a plurality of studied allografts and indicating:(i) values for one or more features of the particular cell or the particular cluster of cells derived from an analysis of at least one image of the particular cell or the particular cluster of cells obtained using imaging mass cytometry (IMC); and(ii) a training label that identifies a known cause, from among a plurality of known causes, of inflammation of the studied allograft to which the particular cell or the particular cluster of cells belongs; applying a machine-learning algorithm to the training data set to train a machine learning model, thereby learning associations between a plurality of cellular features derived from imaging mass cytometry (IMC) data and different causes of allograft inflammation; extracting, from the machine learning model, a variable importance for each of the plurality of cellular features; generating a biosignature by selecting a subset of the cellular features from the plurality of cellular features based on the extracted variable importance for each cellular feature in the plurality of cellular features; and classifying a cause of allograft inflammation in a tissue sample by comparing the biosignature to corresponding cellular features of the tissue sample derived from IMC data for the tissue sample.

2. The method of claim 1, wherein the biosignature includes a plurality of immunological markers.

3. The method of any one of claims 1-2, the method further comprising: scoring the tissue sample based, at least in part, on the biosignature and the IMC data, wherein the score indicates a predicted outcome for a transplant of the tissue sample.

4. The method of claim 3, the method further comprising: identifying one or more cell clusters of the tissue sample associated with the predicted outcome; and using spatial context analysis to identify interactions with neighboring cells of the one or more cell clusters associated with the predicted outcome.

5. The method of any one of claims 1-4, the method further comprising: providing an output indicating a relative importance of one or more of the plurality of cellular features for a predicted outcome.

6. The method of claim 5, wherein the output includes density of single and multiplexed biomarkers in the tissue sample.

7. The method of claim 5, wherein the output includes a visualization representing a spatial relationships to renal microstructural compartments of the tissue sample.

8. The method of any one of claims 1-7, wherein classifying the cause of allograft inflammation includes classifying the tissue sample as one of five allograft inflammation types.

9. The method of any one of claims 1-8, further comprising: scanning the tissue sample with an imaging mass cytometry device to capture the IMC data.

10. The method of any one of claims 1-9, wherein the tissue sample is from a kidney identified as a transplant candidate.

11. The method of any one of claims 1-9, wherein the tissue sample is from a tumor.

12. The method of any one of claims 1-11, wherein the machine learning model is trained on annotated multiplexed immunohistochemistry datasets.

13. The method of claim 12, wherein annotations in the annotated multiplexed immunohistochemistry datasets include single measurement thresholds based on mean intensity.

14. The method of any one of claims 1-13, wherein training the machine learning model to learn the associations between the cellular features and different types of allograft inflammation comprises: training the machine learning model on a mean marker intensity measurement for each of the cellular features; extracting the variable importance for each biomarker of the biomarkers based on calculated variable importances for each of the cellular features; and retraining the machine learning model iteratively starting with the cellular features with a highest importance and successively appending the cellular features with descending variable importance.

15. The method of claim 14, wherein the cellular features include a cellular measurement and one or more spatial features.

16. The method of any one of claims 1-15, wherein the variable importance for each cellular feature of the cellular features is extracted by: calculating, for each cellular feature, an importance score for indicating an effect of the cellular feature on an accuracy of predictions made by the machine learning model.

17. The method of any one of claims 1-16, wherein the machine learning model is an unsupervised machine learning model.

18. The method of any one of claims 1-17, wherein selecting the subset of the cellular features is based on a cumulative importance of all features derived from a given biomarker.

19. The method of any one of claims 1 -18, obtaining a training data set that comprises a collection of training samples further comprises:generating a combined training image by analyzing a plurality of training images, cropping a region of interest in one or more of the plurality of training images, and stitching together the one or more of the plurality of training images; performing cell segmentation on the combined training image; creating a single-measurement classifier for each marker, wherein the singlemeasurement classifier includes a threshold for a mean marker expression; creating a composite classifier by combining the single-measurement classifiers; computing spatial features for cells and cellular subcomponents for each marker used in capturing the training images; and and exporting the training data set.

20. The method of any one of claims 1-19, wherein the studied allografts includes a kidney allograft, and wherein the tissue sample includes a sample of kidney tissue.

21. The method of any one of claims 1-19, wherein the studied allografts includes a liver allograft, and wherein the tissue sample includes a sample of liver tissue.

22. The method of any one of claims 1-19, wherein the studied allografts includes a lung allograft, and the tissue sample includes a sample of lung tissue.

23. The method of any one of claims 1-19, wherein the studied allografts includes a heart allograft, and the tissue sample includes a sample of heart tissue.

24. A system, comprising: one or more computers and one or more computer-readable storage media encoded with instructions that, when executed by the one or more computers, cause the one or more computers to: obtain a training data set that comprises a collection of training samples, each training sample corresponding to a particular cell or a particular cluster of cells in one of a plurality of studied allografts and indicating:(i) values for one or more features of the particular cell or the particular cluster of cells derived from an analysis of at least one image of the particular cell or the particular cluster of cells obtained using imaging mass cytometry (IMC); and(ii) a training label that identifies a known cause, from among a plurality of known causes, of inflammation of the studied allograft to which the particular cell or the particular cluster of cells belongs; apply a machine-learning algorithm to the training data set to train a machine learning model, thereby learning associations between a plurality of cellular features derived from imaging mass cytometry (IMC) data and different causes of allograft inflammation; extract, from the machine learning model, a variable importance for each of the plurality of cellular features; generate a biosignature by selecting a subset of the cellular features from the plurality of cellular features based on the extracted variable importance for each cellular feature in the plurality of cellular features; and classify a cause of allograft inflammation in a tissue sample by comparing the biosignature to corresponding cellular features of the tissue sample derived from IMC data for the tissue sample.

25. A system, comprising: an imaging mass cytometry device configured to scan a tissue sample to collect imaging mass cytometry data; and a computing system operating an allograft inflammation classifier configured to identify a type of allograft inflammation in the tissue sample by comparing a biosignature with imaging mass cytometry data; wherein to generate the biosignature includes to: obtain a training data set that comprises a collection of training samples, each training sample corresponding to a particular cell or a particular cluster of cells in one of a plurality of studied allografts and indicating:(i) values for one or more features of the particular cell or the particular cluster of cells derived from an analysis of at least one image of the particular cellor the particular cluster of cells obtained using imaging mass cytometry (IMC); and(ii) a training label that identifies a known cause, from among a plurality of known causes, of inflammation of the studied allograft to which the particular cell or the particular cluster of cells belongs; apply a machine-learning algorithm to the training data set to train a machine learning model, thereby learning associations between a plurality of cellular features derived from imaging mass cytometry (IMC) data and different causes of allograft inflammation; extract, from the machine learning model, a variable importance for each of the plurality of cellular features; and generate a biosignature by selecting a subset of the cellular features from the plurality of cellular features based on the extracted variable importance for each cellular feature in the plurality of cellular features.

26. The system claim 25, wherein the imaging mass cytometry data, includes spatially resolved single cell data.

27. The system of any one of claims 25-26, wherein the allograft inflammation classifier is further configured to identify single cell phenotypic clusters that correlate with graft loss.

28. A method of predicting causes of liver allograft inflammation, the method comprising: obtaining a training data set that comprises a collection of training samples, each training sample corresponding to a particular cell or a particular cluster of cells in one of a plurality of studied liver allografts and indicating:(i) values for one or more features of the particular cell or the particular cluster of cells derived from an analysis of at least one image of the particular cell or the particular cluster of cells obtained using imaging mass cytometry (IMC); and(ii) a training label that identifies a known cause, from among a plurality of known causes, of inflammation of the studied liver allograft to which the particular cell or the particular cluster of cells belongs;applying a machine-learning algorithm to the training data set to train a machine learning model, thereby learning associations between a plurality of cellular features derived from imaging mass cytometry (IMC) data and different causes of liver allograft inflammation; extracting, from the machine learning model, a variable importance for each of the plurality of cellular features; generating a biosignature by selecting a subset of the cellular features from the plurality of cellular features based on the extracted variable importance for each cellular feature in the plurality of cellular features; and classifying a cause of liver allograft inflammation in a liver tissue sample by comparing the biosignature to corresponding cellular features of the tissue sample derived from IMC data for the liver tissue sample.