Machine learning classifiers for diagnosing and / or predicting kidney allograft t cell-mediated rejection
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2026-02-06
- Publication Date
- 2026-08-13
Smart Images

Figure IMGF000032_0001 
Figure IMGF000032_0002 
Figure IMGF000032_0003
Abstract
Description
[0001] MACHINE LEARNING CLASSIFIERS FOR DIAGNOSING AND / OR PREDICTING KIDNEY ALLOGRAFT T CELL-MEDIATED REJECTION
[0002] FIELD OF THE INVENTION
[0003] The present invention is in the field of medicine, in particular organ transplantation.
[0004] BACKGROUND OF THE INVENTION
[0005] During the last three decades, the international Banff classification has significantly shaped and standardized kidney allograft rejection diagnosis, yet it remains complex due to the limitations of histological analysis, including intra- and inter-observer variability1,2. These challenges necessitate the development of new diagnostic approaches which might offer greater accuracy and objectivity.
[0006] One such approach could be machine learning classifiers based on the Banff Human Organ Transplant (B-HOT) gene panel, a targeted panel designed to measure a wide range of genes associated with solid-organ transplant rejection3. Machine learning classifiers are models that analyze gene expression patterns to identify biological states, such as the presence of diseases or specific immune responses. In the context of kidney transplantation, these classifiers could help to detect early signs of antibody-mediated rejection (AMR) and T cell-mediated rejection (TCMR) by recognizing molecular signatures associated with immune-mediated damage4,5. Their use in clinical practice can improve the precision of diagnostics, allowing for earlier and more personalized interventions for transplant rejection. However, creating precise machine learning classifiers is hampered by the need for large and multicentric cohorts of deeply phenotyped kidney transplant patients.
[0007] One possibility to address the issue of a limited number of samples is the use of synthetic data to augment cohort size and models’ performances. For instance, Generative Adversarial Networks (GANs) are becoming increasingly prevalent in clinical studies to create synthetic data, particularly in image-related research6–8. They use a unique structure that includes two neural network architectures: the ‘generator’ and the ‘discriminator’9. They show adversarial behavior to sample and model from the training dataset. Ultimately, the modeled GAN generates a synthetic dataset similar to the original dataset.
[0008] Several studies have focused on developing a new GAN structure to optimize for tabular data. Notably, based on the GAN algorithm, TGAN was introduced, followed by conditional tabular GAN (CTGAN)10,11. CTGAN aims to deal with the imbalanced discrete variables and generates synthetic data more like the real data than TGAN by a recent benchmark study10,11. However, the evidence supporting the use of synthetic data toenhance the performance of transcriptomic-based prediction classifiers has not yet been assessed.
[0009] SUMMARY OF THE INVENTION
[0010] The present invention is defined by the claims. In particular, the present invention relates to machine learning classifiers for predicting kidney allograft T cell-mediated rejection.
[0011] The present invention relates to a method of predicting (“prediction method”) whether a kidney transplant recipient is at risk of T cell-mediated rejection (TCMR) comprising: a) quantifying the expression levels of a plurality of genes selected from the group consisting of CD96, CD45R0, SELL, MAPK3, CD27, IL16, CTLA4, CD84, LAIR1, TM4SF1, CD3E, CD247, HLADMB, MS4A6A, CD5, IKZF1, ZAP70, INPP5D, HLADMA, LEF1, ALOX5, STAT5A, PLAT, KDR, GZMK, ADAM8, ADAMDEC1, AGR2, AIM2, ANKRD22, AOAH, BATF, BCL2A1, BIRC3, BK large T Ag, BK VP1, BLK, BTK, BTLA, C1QA, C1QB, CALHM6, CAV1, CCL18, CCL19, CCL3 / L1, CCL4, CCL5, CCR2, CCR4, CCR5, CCR6, CCR7, CD163, CD19, CD1D, CD2, CD22, CD28, CD38, CD3D, CD3G, CD4, CD40LG, CD45RA, CD45RB, CD48, CD6, CD69, CD7, CD72, CD74, CD80, CD86, CD8A, CD8B, CIITA, CLEC4C, CSF2RB, CTSS, CTSW, CXCL13, CXCL9, CXCR3, CXCR4, CXCR5, CXCR6, DUSP2, EOMES, EZH2, FAM30A, FSLG, FCER1G, FCGRA1, FCGR3A / B, FCRL2, FGD2, FLT3, FOS, FOXP3, FPR1, GBP1, GBP2, GBP5, GIMAP5, GZMA, GZMB, GZMH, HLADPA1, HLADPB1, HLADQB1, HLADRA, HLADRB1, HLADRB3, HLAF, HLAG, IGOS, IDO1, IFI30, IFNG, IGHG1, IL10RA, IL21R, IL27RA, IL2RA, IL2RB, IL2RG, IL7R, IRF1, IRF4, IRF8, ISG20, ITGA4, ITGAX, ITGB2, JAK3, KLRB1, KLRC1, KLRG1, KLRK1, LAG3, LCK, LCP2, LILRB1, LILRB2, LILRB4, LST1, LTA, LTB, MIR155HG, MME, MMP9, MS4A1, MS4A4A, NFAM1, NFATC2, NKG7, NLRC5, NLRP3, NOD2, NPHS2, OR2 / 1P, PAX5, PDCD1, PDCD1LG2, PDPN, PHEX, PIK3CD, PIK3CG, PLAAT4, POU2AF1, PRF1, PSMB9, PSTPIP1, PTPN22, PTPN6, PTPN7, PTPRC, PTPRO, SAMHD1, SELPLG, SH2D1A, SIRPG, SLA, SLAMF6, SLAMF7, SLAM8, SLC4A1, SP140, SPIB, ST8SIA4, STAT1, STAT4, TAP1, TBX21, TCF7, TCL1A, TFF3, THEMIS, TIGIT, TLR2, TLR7, TLR8, TNFRSF18, TNFRSF1B, TNFRSF4, TNFRSF9, TNFRSF14, TNFSF8, TRAT1, TRDC and XCL1 / 2, in a sample obtained from the recipient;
[0012] b) implementing an algorithm on data comprising the quantified plurality of gene expression levels so as to obtain an algorithm output; and
[0013] c) determining the probability of TCMR from the algorithm output of step b).The present invention also relates to a method for selecting a therapeutic regimen for a kidney transplant recipient identified as having a high risk of TCMR comprising:
[0014] - performing the prediction method according to the invention, and
[0015] - if it is concluded that the kidney transplant recipient has a high risk of TCMR, then selecting a therapeutic regimen chosen from immunosuppressive treatment, corticosteroids and T cell-depleting treatments, preferably chosen from anti-thymocyte globulin, dexamethasone, prednisone and methylprednisone.
[0016] The present invention also relates to a method for monitoring the efficacy of a therapeutic regimen for a kidney transplant recipient at risk of TCMR comprising:
[0017] i) performing the prediction method according to the invention with a first sample obtained from the recipient as a first testing,
[0018] ii) if it is concluded from step i) that the kidney transplant recipient has a high risk of TCMR, then selecting a therapeutic regimen at a first dose, wherein said therapeutic regimen is chosen from immunosuppressive treatments, corticosteroids and T celldepleting treatments, preferably chosen from anti-thymocyte globulin, dexamethasone, prednisone and methylprednisone, and
[0019] iii) performing the prediction method according to the invention with a second sample obtained from the recipient as a second testing, after a time period after steps i) and ii), wherein if it is concluded from step iii) that the kidney transplant recipient has a low risk of TCMR, then a maintenance dose of the therapeutic regimen ii) which is lower than the first dose, is selected.
[0020] The present invention also relates to a method for discriminating a responder kidney transplant recipient from a non-responder kidney transplant recipient at risk of TCMR comprising:
[0021] i) performing the prediction method according to the invention with a first sample obtained from the recipient before any treatment, as a first testing which results in a first probability of TCMR,
[0022] ii) performing the prediction method according to the invention with a second sample obtained from the recipient after a treatment, as a second testing which results in a second probability of TCMR,
[0023] iii) comparing the first and second probabilities of TCMR,
[0024] wherein if it is concluded from step iii) that the recipient has a low risk of TCMR, then concluding that said recipient is a responder,and wherein if it is concluded from step iii) that the recipient has a high risk of TCMR, then concluding that said recipient is a non-responder.
[0025] The present invention also relates to a method for identifying a biomarker for TCMR, the biomarker being a diagnosis biomarker of TCMR, a susceptibility biomarker of TCMR, a prognostic biomarker of TCMR or a predictive biomarker in response to the treatment of TCMR, the method comprising at least the steps of:
[0026] - carrying out the steps of the prediction method according to the invention for a first subject, to obtain a first predicted biological state, wherein the first subject is suffering from TCMR, - carrying out the steps of the prediction method according to the invention for a second subject, to obtain a second predicted biological state, wherein the second subject is not suffering from TCMR, and
[0027] - selecting a biomarker target based on the comparison of the first and second predicted biological states.
[0028] The present invention further relates to an electronic device for obtaining at least one molecular classifying function (FMC) intended to be used to predict whether a kidney transplant recipient is at risk of TCMR, wherein the molecular classifying function (FMC) is adapted to take as input representative parameters of the expression levels of genes of the recipient and to provide as output the risk of TCMR of the recipient, the electronic device comprising:
[0029] - an obtaining module configured to obtain an initial database (IDB), the initial database (IDB) comprising several elements, each element corresponding to a respective recipient and providing, for said respective recipient, several gene expression level parameters (GELP) and other parameters relative to the recipient,
[0030] - a training module configured to train a conditional tabular generative adversarial network on the initial database (IDB) to output elements fulfilling at least one similarity condition with the elements of the initial database (IDB), to obtain a trained conditional tabular generative adversarial network,
[0031] - a using module configured to use the trained conditional tabular generative adversarial network, to obtain synthetic elements, the set of the initial database (IDB) and the synthetic elements forming an augmented database (ADB),
[0032] - a determining module configured to determine gene expression level parameters (GELP) of the augmented database (ADB) impacting a prediction of the risk of TCMR of the recipient by a molecular classifying function, to obtain determined gene expression level parameters (GELPDET),- a selecting module configured to select gene expression level parameters (GELP) among the determined gene expression level parameters (GELPDET), to obtain a set of selected gene expression level parameters (GELPSEL), the selecting module selecting said gene expression level parameters (GELPSEL) iteratively for each determined gene expression level parameter (GELPDET) by:
[0033] - analyzing the impact of adding said determined gene expression level parameters (GELPDET) in a learning process of a molecular classifying function adapted to take as input at least one already selected gene expression level parameter (GELPSEL) and said determined gene expression level parameters (GELPDET) and to provide as output the risk of TCMR of the recipient, and
[0034] - adding the determined gene expression level parameter (GELPDET) in case a selection criterion is fulfilled or discarding said determined gene expression level parameter (GELPDET) in case said selection criterion is not fulfilled,
[0035] - a generating module configured to generate molecular classifying functions (FMCCAND) taking as input the set of selected gene expression level parameters (GELPSEL) and providing as output the risk of TCMR of the recipient, to obtain molecular classifying function candidates (FMCCAND), and
[0036] - a choosing module configured to choose at least one molecular classifying function (FMC) among the molecular classifying function candidates (FMCCAND) based on a performance criterion, to obtain at least one molecular classifying function (FMC) intended to be used to predict a risk of TCMR of the recipient.
[0037] BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1. Development of the T cell-mediated rejection classifiers on augmented transcriptomic data using generative model in kidney transplant patients. Schematic diagram illustrating the development of generative model, generating synthetic data, augmenting the datasets, developing machine learning classifiers, and assessing them in the external validation cohort processes to observe the impact of synthetic data augmentation and to construct the molecular classifiers in kidney transplant patients.
[0039] Figure 2. Flow of the importance of the top 30 genes from the development set to synthetically augmented GanAug200 set. Using the generative adversarial networks (GAN), the inventors generated eight synthetic cohorts resembling the original development set. They were merged to the development set to augment. The inventors performed Boruta algorithm, Monte Carlo method which uses multiple random forests and shadow features.Genes were included as features to predict TCMR. The inventors performed 2000 iterations of random forests. Among the selected features considered predictive by the Boruta algorithm, they calculated median importance to explore the feature importance changes by datasets. The slope graph depicts the flow of the top 30 gene features’ importance. The median importances of genes from the development set to incrementally augmented sets were represented from the left to the right side. (Dark) lines represent genes considered important from the development set. (Light) lines represent genes appeared to be important in the augmented sets.
[0040] Figure 3. Learning curves of the machine learning models. Using the generative adversarial networks (GAN), the inventors generated eight synthetic cohorts resembling the original development set. They were merged to the development set to augment. After the Boruta feature importance analysis, the inventors incrementally increased the number of selected features to observe the performance change for all datasets. The performance was measured in precision recall area under the curve (PRAUC) and receiver operating characteristic area under the curve (ROCAUC) in the split test set.
[0041] Figure 4. Calibration plots on the external validation cohort. This figure details the performance of final molecular classifiers assessed on the external validation cohort. The calibration performance is demonstrated as plots in this figure.
[0042] DETAILED DESCRIPTION OF THE INVENTION
[0043] The inventors aimed to i) develop machine learning classifiers for TCMR of kidney allografts using original data and generative synthetically augmented transcriptomic data, ii) compare the performances of the original data-based and synthetically augmented data-based machine learning classifiers, and iii) assess if synthetically augmented data-based machine learning classifiers capture additional biologically meaningful predictive genes of TCMR. The invention thus aims to harness the potential of these classifiers to create more accurate and reliable diagnostic tools for kidney transplant rejection.
[0044] Main definitions:
[0045] As used herein, the term “kidney transplantation” refers to the process of taking a kidney (i.e. “graft”) from one individual and placing it or them into a different individual. The individual who provides for the transplant is called the “donor” and the individual who receives the transplant is called the “recipient”. As used herein, the term “allograft” refersto a tissue or organ that is transplanted from one individual to another within the same species.
[0046] As used herein, the term " T cell-mediated rejection” or “TCMR” has its general meaning in the art and refers to the pathological process characterized by infiltration of the interstitium by T cells and macrophages, intense IFNγ and TGFβ effects, and epithelial deterioration. The three primary sites of acute TCMR in transplanted kidneys are indeed the tubular epithelial cells, interstitium, and the vascular endothelial cells. TCMR can be acute TCMR or chronic acitve TCMR. Banff classification recognizes three gradings: tubulitis (the t score), interstitial inflammation (the i score) and endarteritis (the v score).
[0047] As used herein, the term “risk” in the context of the present invention, relates to the probability that an event will occur over a specific time period, as in the conversion to allograft loss, and can mean a subject's “absolute” risk or “relative” risk. Absolute risk can be measured with reference to either actual observation post-measurement for the relevant time cohort, or with reference to index values developed from statistically valid historical cohorts that have been followed for the relevant time period. Relative risk refers to the ratio of absolute risks of a subject compared either to the absolute risks of low risk cohorts or an average population risk, which can vary by how clinical risk factors are assessed. Odds ratios, the proportion of positive events to negative events for a given test result, are also commonly used (odds are according to the formula p / (l— p) where p is the probability of event and (l-p) is the probability of no event) to no conversion. Accordingly, the expression “predicting whether a kidney transplant recipient is at risk of TCMR” in the context of the present invention encompasses making a prediction of the probability, odds, or likelihood that allograft loss mediated by TCMR may occur. The methods of the present invention may be used to make continuous or categorical measurements of the risk of conversion to TCMR. According to the present invention, the prediction is a dynamic prediction. As used herein, the term "dynamic prediction" refers to providing an assessment of probability or likelihood for a subject to develop TCMR, which may change over time. The method of the present invention is particularly suitable to predict the risk of TCMR at 3, 5 and / or 7 years from the date of prediction. As used herein, a “high risk of TCMR” means that allograft loss mediated by TCMR may occur more probably than not. As used herein, a “low risk of TCMR” means that it is more probable that allograft loss mediated by TCMR may not occur.
[0048] As used herein, the term “algorithm” is any mathematical equation, algorithmic, analytical or programmed process, or statistical technique that takes one or more continuous parameters and calculates an output value, sometimes referred to as an “index” or “index value”.As used herein, the term "machine learning classifier" refers to a computational model that analyzes gene expression data to predict biological states, such as disease presence or immune responses. These classifiers leverage complex algorithms and statistical methods to identify patterns within the data that correlate with specific biological outcomes. The use of machine learning classifiers facilitates the development of personalized medicine approaches, enabling more accurate predictions of disease progression and responses to therapy.
[0049] As used herein, the term “random forest” refers to an ensemble learning method used for classification and regression tasks. It operates by constructing a multitude of decision trees during training time and outputting the mode of the classes (classification) or mean prediction (regression) of the individual trees. Random forests correct for decision trees' habit of overfitting to their training set by introducing randomness into the tree-building process. Each tree is trained on a random subset of the training data, and features are randomly selected at each split, ensuring that the resulting model is both robust and generalizable.
[0050] As used herein, the term “gene” has its general meaning in the art and refers to a sequence of nucleotides in DNA or RNA that encodes the synthesis of a gene product, either RNA or protein, which performs a specific function in the organism. Genes are the fundamental units of heredity and are transferred from parent to offspring during reproduction. They are responsible for the inherited characteristics and functions of living organisms, and their expression levels can be quantitatively assessed to provide insights into various biological processes and disease states. In the present specification, the name of each of the genes of interest refers to the internationally recognised name of the corresponding gene, as found in internationally recognised gene sequences and protein sequences databases, in particular in the database from the HUGO Gene Nomenclature Committee, that is available notably at the following Internet address http: / / www.gene.ucl.ac.uk / nomenclature / index.html. In the present specification, the name of each of the various biological markers of interest may also refer to the internationally recognised name of the corresponding gene, as found in the internationally recognised gene sequences and protein sequences databases ENTRE ID, Genbank, TrEMBL or ENSEMBL. Through these internationally recognised sequence databases, the nucleic acid sequences corresponding to each of the gene of interest described herein may be retrieved by the one skilled in the art.
[0051] As used herein, the term “expression level” refers to the quantity of mRNA produced from a gene, which serves as an indicator of gene activity within a sample.As used herein, the term “sample” refers to any biological sample obtained for evaluation in vitro. In the context of transplantation, this typically includes tissue samples. The term “tissue sample” encompasses sections of tissues such as biopsy or autopsy samples and frozen sections taken for histological analysis. According to the present invention the sample is obtained from the transplanted organ, preferably the transplanted kidney. Thus, in some embodiments, the tissue sample may result from a biopsy performed on the transplanted organ to assess its health and function.
[0052] As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” include plural references unless the context clearly dictates otherwise. Thus, for example, “a ribonucleotide” is understood to represent one or more ribonucleotides. As such, the terms “a,” “an,” “one or more,” and “at least one” can be used interchangeably herein. Throughout this specification and embodiments, the words “have” and “comprise,” or variations such as “has,” “having,” “comprises,” or “comprising” will be understood to imply the inclusion of a stated integer or group of integers but not the exclusion of any other integer or group of integers. It is further understood that wherever embodiments are described herein with the language “comprising” or “having,” or grammatical equivalents thereof, otherwise analogous embodiments described in terms of “consisting of” and / or “consisting essentially of’ are also provided.
[0053] Methods of the present invention:
[0054] The present invention relates to a method of predicting (“prediction method”) whether a kidney transplant recipient is at risk of T cell-mediated rejection (TCMR) comprising: a) quantifying the expression levels of a plurality of genes selected from the group consisting of CD96, CD45R0, SELL, MAPK3, CD27, IL16, CTLA4, CD84, LAIR1, TM4SF1, CD3E, CD247, HLADMB, MS4A6A, CD5, IKZF1, ZAP70, INPP5D, HLADMA, LEF1, ALOX5, STAT5A, PLAT, KDR, GZMK, ADAM8, ADAMDEC1, AGR2, AIM2, ANKRD22, AOAH, BATF, BCL2A1, BIRC3, BK large T Ag, BK VP1, BLK, BTK, BTLA, C1QA, C1QB, CALHM6, CAV1, CCL18, CCL19, CCL3 / L1, CCL4, CCL5, CCR2, CCR4, CCR5, CCR6, CCR7, CD163, CD19, CD1D, CD2, CD22, CD28, CD38, CD3D, CD3G, CD4, CD40LG, CD45RA, CD45RB, CD48, CD6, CD69, CD7, CD72, CD74, CD80, CD86, CD8A, CD8B, CIITA, CLEC4C, CSF2RB, CTSS, CTSW, CXCL13, CXCL9, CXCR3, CXCR4, CXCR5, CXCR6, DUSP2, EOMES, EZH2, FAM30A, FSLG, FCER1G, FCGRA1, FCGR3A / B, FCRL2, FGD2, FLT3, FOS, FOXP3, FPR1, GBP1, GBP2, GBP5, GIMAP5, GZMA, GZMB, GZMH, HLADPA1, HLADPB1, HLADQB1, HLADRA, HLADRB1, HLADRB3, HLAF, HLAG, IGOS, IDO1, IFI30, IFNG, IGHG1, IL10RA, IL21R, IL27RA, IL2RA, IL2RB, IL2RG, IL7R, IRF1, IRF4, IRF8, ISG20, ITGA4, ITGAX, ITGB2, JAK3, KLRB1, KLRC1, KLRG1, KLRK1,LAG3, LCK, LCP2, LILRB1, LILRB2, LILRB4, LST1, LTA, LTB, MIR155HG, MME, MMP9, MS4A1, MS4A4A, NFAM1, NFATC2, NKG7, NLRC5, NLRP3, NOD2, NPHS2, OR2 / 1P, PAX5, PDCD1, PDCD1LG2, PDPN, PHEX, PIK3CD, PIK3CG, PLAAT4, POU2AF1, PRF1, PSMB9, PSTPIP1, PTPN22, PTPN6, PTPN7, PTPRC, PTPRO, SAMHD1, SELPLG, SH2D1A, SIRPG, SLA, SLAMF6, SLAMF7, SLAM8, SLC4A1, SP140, SPIB, ST8SIA4, STAT1, STAT4, TAP1, TBX21, TCF7, TCL1A, TFF3, THEMIS, TIGIT, TLR2, TLR7, TLR8, TNFRSF18, TNFRSF1B, TNFRSF4, TNFRSF9, TNFRSF14, TNFSF8, TRAT1, TRDC and XCL1 / 2, in a sample obtained from the recipient;
[0055] b) implementing an algorithm on data comprising the quantified plurality of gene expression levels so as to obtain an algorithm output; and
[0056] c) determining the probability of TCMR from the algorithm output of step b).
[0057] In some embodiments, the expression levels of two or more genes selected from the group consisting of CD96, CD45R0, SELL, MAPK3, CD84, LAIR1, TM4SF1, CD3E, CD27, CD247, HLADMB, IL16, MS4A6A, CD5, IKZF1, CTLA4, ZAP70, INPP5D, HLADMA, LEF1, ALOX5, STAT5A, PLAT, KDR, GZMK, ADAM8, ADAMDEC1, AGR2, AIM2, ANKRD22, AOAH, BATF, BCL2A1, BIRC3, BK large T Ag, BK VP1, BLK, BTK, BTLA, C1QA, C1QB, CALHM6, CAV1, CCL18, CCL19, CCL3 / L1, CCL4, CCL5, CCR2, CCR4, CCR5, CCR6, CCR7, CD163, CD19, CD1D, CD2, CD22, CD28, CD38, CD3D, CD3G, CD4, CD40LG, CD45RA, CD45RB, CD48, CD6, CD69, CD7, CD72, CD74, CD80, CD86, CD8A, CD8B, CIITA, CLEC4C, CSF2RB, CTSS, CTSW, CXCL13, CXCL9, CXCR3, CXCR4, CXCR5, CXCR6, DUSP2, EOMES, EZH2, FAM30A, FSLG, FCER1G, FCGRA1, FCGR3A / B, FCRL2, FGD2, FLT3, FOS, FOXP3, FPR1, GBP1, GBP2, GBP5, GIMAP5, GZMA, GZMB, GZMH, HLADPA1, HLADPB1, HLADQB1, HLADRA, HLADRB1, HLADRB3, HLAF, HLAG, IGOS, IDO1, IFI30, IFNG, IGHG1, IL10RA, IL21R, IL27RA, IL2RA, IL2RB, IL2RG, IL7R, IRF1, IRF4, IRF8, ISG20, ITGA4, ITGAX, ITGB2, JAK3, KLRB1, KLRC1, KLRG1, KLRK1, LAG3, LCK, LCP2, LILRB1, LILRB2, LILRB4, LST1, LTA, LTB, MIR155HG, MME, MMP9, MS4A1, MS4A4A, NFAM1, NFATC2, NKG7, NLRC5, NLRP3, NOD2, NPHS2, OR2 / 1P, PAX5, PDCD1, PDCD1LG2, PDPN, PHEX, PIK3CD, PIK3CG, PLAAT4, POU2AF1, PRF1, PSMB9, PSTPIP1, PTPN22, PTPN6, PTPN7, PTPRC, PTPRO, SAMHD1, SELPLG, SH2D1A, SIRPG, SLA, SLAMF6, SLAMF7, SLAM8, SLC4A1, SP140, SPIB, ST8SIA4, STAT1, STAT4, TAP1, TBX21, TCF7, TCL1A, TFF3, THEMIS, TIGIT, TLR2, TLR7, TLR8, TNFRSF18, TNFRSF1B, TNFRSF4, TNFRSF9, TNFRSF14, TNFSF8, TRAT1, TRDC and XCL1 / 2, are quantified in step a).Preferably, the present invention relates to a method of predicting whether a kidney transplant recipient is at risk of TCMR comprising:
[0058] a) quantifying the expression levels of a plurality of genes selected from the group consisting of CD96, CD45R0, SELL, MAPK3, CD84, LAIR1, TM4SF1, CD3E, CD27, CD247, HLADMB, IL16, MS4A6A, CD5, IKZF1, CTLA4, ZAP70, INPP5D, HLADMA, LEF1, ALOX5, STAT5A, PLAT, KDR and GZMK in a sample obtained from the recipient;
[0059] b) implementing an algorithm on data comprising the quantified plurality of gene expression levels so as to obtain an algorithm output; and
[0060] c) determining the probability of TCMR from the algorithm output of step b).
[0061] In some embodiments, the expression levels of two or more genes selected from the group consisting of CD96, CD45R0, SELL, MAPK3, CD27, IL16, CTLA4, CD84, LAIR1, TM4SF1, CD3E, CD247, HLADMB, MS4A6A, CD5, IKZF1, ZAP70, INPP5D, HLADMA, LEF1, ALOX5, STAT5A, PLAT, KDR and GZMK, are quantified in step a).
[0062] In some embodiments, the expression levels of 2; 3; 4; 5; 6; 7; 8; 9; 10; 11; 12; 13; 14; 15; 16; 17; 18; 19; 20; 21; 22; 23; 24; 25; 50; 55; 60; 70; 80; 90; 100; 105; 110; 120; 130; 140; 150; 160; 170; 180; 190; 200 or 207 genes are determined in step a).
[0063] In some embodiments, the expression levels of CD96 and one or more genes selected from the group consisting of CD45R0, SELL, MAPK3, CD27, IL16, CTLA4, CD84, LAIR1, TM4SF1, CD3E, CD247, HLADMB, MS4A6A, CD5, IKZF1, ZAP70, INPP5D, HLADMA, LEF1, ALOX5, STAT5A, PLAT, KDR and GZMK, are quantified in step a).
[0064] In some embodiments, the expression levels of CD96, CD45R0, and one or more genes selected from the group consisting of SELL, MAPK3, CD27, IL16, CTLA4, CD84, LAIR1, TM4SF1, CD3E, CD247, HLADMB, MS4A6A, CD5, IKZF1, ZAP70, INPP5D, HLADMA, LEF1, ALOX5, STAT5A, PLAT, KDR and GZMK, are quantified in step a).
[0065] In some embodiments, the expression levels of CD96, CD45R0, SELL and one or more genes selected from the group consisting of MAPK3, CD27, IL16, CTLA4, CD84, LAIR1, TM4SF1, CD3E, CD247, HLADMB, MS4A6A, CD5, IKZF1, ZAP70, INPP5D, HLADMA, LEF1, ALOX5, STAT5A, PLAT, KDR and GZMK, are quantified in step a).
[0066] In some embodiments, the expression levels of CD96, CD45R0, SELL, MAPK3 and one or more genes selected from the group consisting of CD27, IL16, CTLA4, CD84, LAIR1,TM4SF1, CD3E, CD247, HLADMB, MS4A6A, CD5, IKZF1, ZAP70, INPP5D, HLADMA, LEF1, ALOX5, STAT5A, PLAT, KDR and GZMK, are quantified in step a).
[0067] In some embodiments, preferably, the expression levels of at least CD96, CD45R0, SELL and MAPK3 are determined. More preferably, the expression levels of at least CD96, CD45R0, SELL, MAPK3, CD27, IL16 and CTLA4 are determined.
[0068] In some embodiments, the expression levels of one or more genes selected from the group consisting of genes related to immune cell recruitment and chemotaxis (e.g., CCR2), immune cell interactions (e.g., CD40L), inflammatory and immune response (e.g., IL1B, IL7), and cell adhesion and migration (e.g., ITGA4, ITGB2), is / are further determined.
[0069] Typically the prediction method of the invention is performed after the recipient has been transplanted.
[0070] Quantifying the expression levels:
[0071] Methods for determining (quantifying) the expression level of a gene product such as nucleic acid (e.g. RNA) are also well known in the art. Conventional methods typically involve polymerase chain reaction (PCR). For instance, U. S. Pat. Nos. 4,683,202, 4,683,195, 4,800,159, and 4,965,188 disclose conventional PCR techniques. PCR typically employs two oligonucleotide primers that bind to a selected target nucleic acid sequence. Primers useful in the present invention include oligonucleotides capable of acting as a point of initiation of nucleic acid synthesis within the target nucleic acid sequence. A primer can be purified from a restriction digest by conventional methods, or it can be produced synthetically. PCR involves use of a thermostable polymerase. The term “thermostable polymerase” refers to a polymerase enzyme that is heat stable, i.e., the enzyme catalyzes the formation of primer extension products complementary to a template and does not irreversibly denature when subjected to the elevated temperatures for the time necessary to effect denaturation of double-stranded template nucleic acids. Thermostable polymerases have been isolated from Thermus fiavus, T. ruber, T. thermophilus, T. aquaticus, T. lacteus, T. rubens, Bacillus stearothermophilus, and Methanothermus fervidus. Nonetheless, polymerases that are not thermostable also can be employed in PCR assays provided the enzyme is replenished. Typically, the polymerase is a Taq polymerase (i.e. Thermus aquaticus polymerase). Quantitative PCR is typically carried out in a thermal cycler with the capacity to illuminate each sample with a beam of light of a specified wavelength and detect the fluorescence emitted by the excited fluorophore. The thermalcycler is also able to rapidly heat and chill samples, thereby taking advantage of the physicochemical properties of the nucleic acids and thermal polymerase. In order to detect and measure the amount of amplicon (i.e. amplified target nucleic acid sequence) in the sample, a measurable signal has to be generated, which is proportional to the amount of amplified product. All current detection systems use fluorescent technologies. Some of them are non-specific techniques, and consequently only allow the detection of one target at a time. Alternatively, specific detection chemistries can distinguish between non- specific amplification and target amplification. These specific techniques can be used to multiplex the assay, i.e. detecting several different targets in the same assay. For example, SYBR® Green I probes, High Resolution Melting probes, TaqMan® probes, LNA® probes and Molecular Beacon probes can be suitable. TaqMan® probes are the most widely used type of probes. They were developed by Roche (Basel, Switzerland) and ABI (Foster City, USA) from an assay that originally used a radio-labelled probe (Holland et al. 1991), which consisted of a single-stranded probe sequence that was complementary to one of the strands of the amplicon. A fluorophore is attached to the 5’ end of the probe and a quencher to the 3’ end. The fluorophore is excited by the machine and passes its energy, via FRET (Fluorescence Resonance Energy Transfer) to the quencher. Traditionally, the FRET pair has been conjugated to FAM as the fluorophore and TAMRA as the quencher. In a well-designed probe, FAM does not fluoresce as it passes its energy onto TAMRA. As TAMRA fluorescence is detected at a different wavelength to FAM, the background level of FAM is low. The probe binds to the amplicon during each annealing step of the PCR. When the Taq polymerase extends from the primer which is bound to the amplicon, it displaces the 5’ end of the probe, which is then degraded by the 5’-3’ exonuclease activity of the Taq polymerase. Cleavage continues until the remaining probe melts off the amplicon. This process releases the fluorophore and quencher into solution, spatially separating them (compared to when they were held together by the probe). This leads to an irreversible increase in fluorescence from the FAM and a decrease in the TAMRA.
[0072] In some embodiments, the expression level is determined by RNA-seq. As used, the term " RNA-Seq" or "transcriptome sequencing" refers to sequencing performed on RNA (or cDNA) instead of DNA, where typically, the primary goal is to measure expression levels, detect fusion transcripts, alternative splicing, and other genomic alterations that can be better assessed from RNA. RNA-Seq typically includes whole transcriptome sequencing. As used herein, the term “whole transcriptome sequencing” refers to the use of high throughput sequencing technologies to sequence the entire transcriptome in order to get information about a sample's RNA content. Whole transcriptome sequencing can be donewith a variety of platforms for example, the Genome Analyzer (Illumina, Inc., San Diego, Calif.) and the SOLiD™ Sequencing System (Life Technologies, Carlsbad, Calif.), However, any platform useful for whole transcriptome sequencing may be used. Typically, the RNA is extracted, and ribosomal RNA may be deleted as described in U. S. Pub, No.
[0073] 2011 / 0111409. cDNA sequencing libraries may be prepared that are directional and single or paired-end using commercially available kits such as the ScriptSeq™ M mRNA-Seq Library Preparation Kit (Epicenter Biotechnologies, Madison, Wis.). The libraries may also be barcoded for multiplex sequencing using commercially available barcode primers such as the RNA-Seq Barcode Primers from Epicenter Biotechnologies (Madison, Wis.). PCR is then carried out to generate the second strand of cDNA to incorporate the barcodes and to amplify the libraries. After the libraries are quantified, the sequencing libraries may be sequenced. Nucleic acid sequencing technologies are suitable methods for expression analysis. The principle underlying these methods is that the number of times a (DNA sequence is detected in a sample is directly related to the relative RNA levels corresponding to that sequence. These methods are sometimes referred to by the term Digital Gene Expression (DGE) to reflect the discrete numeric property of the resulting data. Early methods applying this principle were Serial Analysis of Gene Expression (SAGE) and Massively Parallel Signature Sequencing (MPSS). See, e.g., S. Brenner, et al., Nature Biotechnology 18(6):630-634 (2000). Typically RNA-seq uses Next Generation Sequencing or NGS. As used herein, the term " Next Generation Sequencing" (NGS) refers to a relatively new sequencing technique as compared to the traditional Sanger sequencing technique. For review, see Shendure et al., Nature Biotech., 26(10): 1135-45 (2008), which is hereby incorporated by reference into this disclosure. For purpose of this disclosure, NGS may include cyclic array sequencing, microelectrophoretic sequencing, sequencing by hybridization, among others. By way of example, in a typical NGS using cyclic-array methods, genomic DNA or cDNA library is first prepared, and common adaptors may then be ligated to the fragmented genomic DNA or cDNA. Different protocols may be used to generate jumping libraries of mate-paired tags with controllable distance distribution. An array of millions of spatially immobilized PCR colonies or "polonies" is generated with each polonies consisting of many copies of a single shotgun library fragment. Because the polonies are tethered to a planar array, a single microliter- scale reagent volume can be applied to manipulate the array features in parallel, for example, for primer hybridization or for enzymatic extension reactions. Imaging-based detection of fluorescent labels incorporated with each extension may be used to acquire sequencing data on all features in parallel. Successive iterations of enzymatic interrogation and imaging may also be used to build up a contiguous sequencing read for each array feature.In some embodiments, the nCounter® Analysis system is used to detect intrinsic gene expression. The basis of the nCounter® Analysis system is the unique code assigned to each nucleic acid target to be assayed (International Patent Application Publication No. WO 08 / 124847, U. S. Patent No. 8,415,102 and Geiss et al. Nature Biotechnology. 2008. 26(3): 317-325; the contents of which are each incorporated herein by reference in their entireties). The code is composed of an ordered series of colored fluorescent spots which create a unique barcode for each target to be assayed. A pair of probes is designed for each DNA or RNA target, a biotinylated capture probe and a reporter probe carrying the fluorescent barcode. This system is also referred to, herein, as the nanoreporter code system. Specific reporter and capture probes are synthesized for each target. The reporter probe can comprise at a least a first label attachment region to which are attached one or more label monomers that emit light constituting a first signal; at least a second label attachment region, which is non-over-lapping with the first label attachment region, to which are attached one or more label monomers that emit light constituting a second signal; and a first target- specific sequence. Preferably, each sequence specific reporter probe comprises a target specific sequence capable of hybridizing to no more than one gene and optionally comprises at least three, or at least four label attachment regions, said attachment regions comprising one or more label monomers that emit light, constituting at least a third signal, or at least a fourth signal, respectively. The capture probe can comprise a second targetspecific sequence; and a first affinity tag. In some embodiments, the capture probe can also comprise one or more label attachment regions. Preferably, the first target- specific sequence of the reporter probe and the second target- specific sequence of the capture probe hybridize to different regions of the same gene to be detected. Reporter and capture probes are all pooled into a single hybridization mixture, the "probe library". The relative abundance of each target is measured in a single multiplexed hybridization reaction. The method comprises contacting the tissue sample with a probe library, such that the presence of the target in the sample creates a probe pair - target complex. The complex is then purified. More specifically, the sample is combined with the probe library, and hybridization occurs in solution. After hybridization, the tripartite hybridized complexes (probe pairs and target) are purified in a two-step procedure using magnetic beads linked to oligonucleotides complementary to universal sequences present on the capture and reporter probes. This dual purification process allows the hybridization reaction to be driven to completion with a large excess of target-specific probes, as they are ultimately removed, and, thus, do not interfere with binding and imaging of the sample. All post hybridization steps are handled robotically on a custom liquid-handling robot (Prep Station, NanoString Technologies). Purified reactions are typically deposited by the Prep Station into individual flow cells of asample cartridge, bound to a streptavidin-coated surface via the capture probe, electrophoresed to elongate the reporter probes, and immobilized. After processing, the sample cartridge is transferred to a fully automated imaging and data collection device (Digital Analyzer, NanoString Technologies). The level of a target is measured by imaging each sample and counting the number of times the code for that target is detected. For each sample, typically 600 fields-of-view (FOV) are imaged (1376 X 1024 pixels) representing approximately 10 mm² of the binding surface. Typical imaging density is 100- 1200 counted reporters per field of view depending on the degree of multiplexing, the amount of sample input, and overall target abundance. Data is output in simple spreadsheet format listing the number of counts per target, per sample. This system can be used along with nanoreporters. Additional disclosure regarding nanoreporters can be found in International Publication No. WO 07 / 076129 and W007 / 076132, and US Patent Publication No.
[0074] 2010 / 0015607 and 2010 / 0261026, the contents of which are incorporated herein in their entireties. Further, the term nucleic acid probes and nanoreporters can include the rationally designed (e.g. synthetic sequences) described in International Publication No. WO 2010 / 019826 and US Patent Publication No.2010 / 0047924, incorporated herein by reference in its entirety.
[0075] Typically, the expression level of a gene may be expressed as absolute level or normalized level. For instance, levels are normalized by correcting the absolute level of a gene by comparing its expression to the expression of a gene that is not a relevant for determining the risk. This normalization allows the comparison of the level in one sample, e.g., a subject sample, to another sample, or between samples from different sources.
[0076] Machine learning classifier:
[0077] The algorithm described in the present invention involves applying the determined expression levels of genes to a machine learning classifier designed to predict the risk of T cell-mediated rejection (TCMR).
[0078] Typically, the algorithm produces a score. The term “score” refers to a numerical value derived by combining one or more parameters using a mathematical algorithm or formula. The combination of parameters can be achieved, for example, by multiplying each expression level with a defined coefficient and summing the resulting products to yield a score.
[0079] According to the present invention, the machine learning classifier is a random forest algorithm.
[0080] The generation of the machine learning classifier intended for predicting the risk of TCMR in kidney allograft recipients encompasses a series of meticulously outlined steps.Initially, a comprehensive data collection is assembled, which includes a derivation cohort composed of over 600 kidney allograft biopsy samples sourced from multiple medical centers. Furthermore, an external validation cohort is also incorporated to ensure the robustness and reliability of the classifier.
[0081] To measure the gene expression levels in biopsy samples from the derivation cohort, the Banff Human Organ Transplant panel on the NanoString nCounter® platform is employed. This panel includes 770 genes across 37 pathways to identify biomarkers for rejection. This provides a precise quantification of gene expression, which is critical for the subsequent stages of classifier development. The feature selection and importance analysis are conducted using the Boruta algorithm. This algorithm utilizes multiple random forests and shadow features to identify the most significant genes that contribute to the prediction of TCMR.
[0082] To enhance the performance of the classifier, synthetic data is generated using a Conditional Tabular GAN (CTGAN) model. This model augments the original dataset, ensuring that the synthetic data maintains statistical similarity to the actual data, thereby upholding the quality and reliability of the dataset. To validate the quality of the synthetic data, rigorous statistical assessments are performed.
[0083] Learning curves are generated using a random forest algorithm to optimize the number of features for the final classifier. This process involves forward stepwise feature selection, which systematically adds features to the model to determine the optimal combination that maximizes performance. The development of the machine learning classifiers involves the use of random forest algorithms, with hyperparameters being finetuned through a grid search on a fixed one-fold cross-validation.
[0084] The classifiers' efficacy is rigorously validated using the external validation cohort. The performance of the classifiers is assessed using a range of metrics, including the Receiver Operating Characteristic Area Under the Curve (ROCAUC), Precision-Recall Area Under the Curve (PRAUC), accuracy, balanced accuracy, F1 score, negative predictive value (NPV), positive predictive value (PPV), sensitivity, specificity, and Brier Score. These metrics provide a comprehensive evaluation of the classifiers' ability to accurately predict the risk of TCMR in kidney allograft recipients.
[0085] The algorithm described in this invention can be executed by one or more programmable processors running computer programs that perform functions by processing input data and generating output. Additionally, the algorithm can be implemented using special-purpose logic circuitry, such as FPGA (field-programmable gate arrays) or ASIC (application-specific integrated circuits). Processors suitable for executing a computer program include both general-purpose and special-purpose microprocessors,capable of operating as part of any digital computer. Generally, a processor obtains instructions and data from read-only memory or random-access memory or both. Essential components of a computer include a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer is also equipped with or connected to mass storage devices for data storage, such as magnetic disks, magneto-optical disks, or optical disks. However, it is not mandatory for a computer to possess these devices, and it can be embedded within another device. Suitable computer-readable media for storing program instructions and data encompass all forms of nonvolatile memory, including semiconductor memory devices (e.g., EPROM, EEPROM, flash memory), magnetic disks (e.g., hard disks, removable disks), magneto-optical disks, and CD-ROM and DVD-ROM disks. The processor and memory may be supplemented by or incorporated into special-purpose logic circuitry.
[0086] To facilitate user interaction, embodiments of this invention can be implemented on a computer equipped with a display device, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for presenting information to the user, along with input devices like a keyboard and a pointing device (e.g., mouse, trackball). Interaction with the user can also involve other sensory feedback mechanisms, including visual, auditory, or tactile feedback, and inputs can be received through various means, including acoustic, speech, or tactile input.
[0087] Furthermore, the algorithm can be deployed in a computing system comprising a back-end component (e.g., a data server), middleware component (e.g., an application server), or front-end component (e.g., a client computer with a graphical user interface or Web browser) to enable user interaction. These components can interconnect via any form of digital communication network, such as a LAN (local area network) or WAN (wide area network), including the Internet. Such a computing system may consist of clients and servers, typically remote from each other, interacting through a communication network. The client-server relationship is defined by the computer programs running on the respective computers, establishing a client-server relationship.
[0088] The present invention further relates to an electronic device for obtaining at least one molecular classifying function (FMC) intended to be used to predict whether a kidney transplant recipient is at risk of TCMR, wherein the molecular classifying function (FMC) is adapted to take as input representative parameters of the expression levels of genes of the recipient and to provide as output the risk of TCMR of the recipient, the electronic device comprising:- an obtaining module configured to obtain an initial database (IDB), the initial database (IDB) comprising several elements, each element corresponding to a respective recipient and providing, for said respective recipient, several gene expression level parameters (GELP) and other parameters relative to the recipient,
[0089] - a training module configured to train a conditional tabular generative adversarial network on the initial database (IDB) to output elements fulfilling at least one similarity condition with the elements of the initial database (IDB), to obtain a trained conditional tabular generative adversarial network. More specifically, a conditional tabular generative adversarial network is a specific generative Adversarial Network (more often named GAN). Such network uses a unique structure that includes two neural network architectures: the ‘generator’ and the ‘discriminator’. More precisely, the generator creates synthetic data samples, while the discriminator evaluates them against real data samples to distinguish between the two. In the context of tabular data, the generator is trained to produce synthetic data that mimics the statistical properties and distribution of the real tabular dataset. It learns to generate rows of data with realistic feature values. The discriminator's role is to differentiate between real and synthetic data samples. It provides feedback to the generator, helping it improve the quality of the generated data over time,
[0090] - a using module configured to use the trained conditional tabular generative adversarial network, to obtain synthetic elements, the set of the initial database (IDB) and the synthetic elements forming an augmented database (ADB),
[0091] - a determining module configured to determine gene expression level parameters (GELP) of the augmented database (ADB) impacting a prediction of the risk of TCMR of the recipient by a molecular classifying function, to obtain determined gene expression level parameters (GELPDET),
[0092] - a selecting module configured to select gene expression level parameters (GELP) among the determined gene expression level parameters (GELPDET), to obtain a set of selected gene expression level parameters (GELPSEL), the selecting module selecting said gene expression level parameters (GELPSEL) iteratively for each determined gene expression level parameter (GELPDET) by:
[0093] - analyzing the impact of adding said determined gene expression level parameters (GELPDET) in a learning process of a molecular classifying function adapted to take as input at least one already selected gene expression level parameter (GELPSEL) and said determined gene expression level parameters (GELPDET) and to provide as output the risk of TCMR of the recipient, and- adding the determined gene expression level parameter (GELPDET) in case a selection criterion is fulfilled or discarding said determined gene expression level parameter (GELPDET) in case said selection criterion is not fulfilled,
[0094] - a generating module configured to generate molecular classifying functions (FMCCAND) taking as input the set of selected gene expression level parameters (GELPSEL) and providing as output the risk of TCMR of the recipient, to obtain molecular classifying function candidates (FMCCAND), and
[0095] - a choosing module configured to choose at least one molecular classifying function (FMC) among the molecular classifying function candidates (FMCCAND) based on a performance criterion, to obtain at least one molecular classifying function (FMC) intended to be used to predict a risk of TCMR of the recipient.
[0096] By definition, a molecular classifying function FMC is adapted to take as input several gene expression level parameters of the subject and provide as output the biological state of the subject, a gene expression level parameter being a parameter representative of the expression level of said gene for said subject.
[0097] Gene expression level parameters should be construed as encompassing any parameter representative of the expression level, such as gene expression profiles, protein levels, or metabolite concentrations.
[0098] The molecular classifying function FMC can be qualified as a classifier and is here obtained by using an artificial intelligence technique.
[0099] An artificial intelligence technique consists in establishing a model (also named algorithm) based on data.
[0100] In particular, the artificial intelligence technique often implies learning the model. The term “machine learning” is thus employed to designate the fact that the model is learned by the machine based on data.
[0101] According to the case, the machine learning technique implies using a learning among a supervised learning, an unsupervised learning, a semi-supervised learning, a reinforcement learning, a self-learning, a feature learning, a sparse dictionary learning, an anomaly detection learning, a robot learning and association rules learning.
[0102] In particular, in the present example, the machine learning technique is a supervised learning technique, a semi-supervised learning technique or a reinforcement learning technique.
[0103] The model used in the artificial intelligence technique can be chosen from various models / algorithms, such as computational models and algorithms for classification,clustering, regression and dimensionality reduction, such as neural networks, genetic algorithms, support vector machines, k-means, kernel regression and discriminant analysis.
[0104] More generally, the artificial intelligence technique may imply the use of one or several of the following elements: sums, ratios, and regression operators, such as coefficients or exponents, biomarker value transformations and normalizations (including, without limitation, those normalization schemes based on clinical parameters, such as clinical data 58, gender, age or ethnicity), rules and guidelines, statistical classification models, and neural networks, structural and syntactic statistical classification algorithms, and methods of risk index construction, utilizing pattern recognition features, including established techniques such as cross-correlation, Principal Components Analysis (PCA), factor rotation, Logistic Regression (LogReg), Linear Discriminant Analysis (LDA), Eigengene Linear Discriminant Analysis (ELDA), Support Vector Machines (SVM), Random Forest (RF), Recursive Partitioning Tree (RPART), as well as other related decision tree classification techniques, Shrunken Centroids (SC), StepAIC, Kth-Nearest Neighbor, Boosting, Decision Trees, Neural Networks, Bayesian Networks, Support Vector Machines, and Hidden Markov Models, among others.
[0105] Alternatively or in complement, the artificial intelligence technique may imply the use of one or several of the following elements: Average One-Dependence Estimators (AODE), Artificial neural network (e.g., Backpropagation), Bayesian statistics (e.g., Naive Bayes classifier, Bayesian network, Bayesian knowledge base), Case-based reasoning, Decision trees, Inductive logic programming, Gaussian process regression, Group method of data handling (GMDH), Learning Automata, Learning Vector Quantization, Minimum message length (decision trees, decision graphs, etc.), Lazy learning, Instance-based learning Nearest Neighbor Algorithm, Analogical modeling, Probably approximately correct learning (PAC) learning, Ripple down rules, a knowledge acquisition methodology, Symbolic machine learning algorithms, Subsymbolic machine learning algorithms, Support vector machines, Random Forests, Ensembles of classifiers, Bootstrap aggregating (bagging), boosting, regression analysis, Information fuzzy networks (IFN), statistical classification, AODE, Linear classifiers (e.g., Fisher's linear discriminant, Logistic regression, Naive Bayes classifier, Perceptron, and Support vector machine), quadratic classifiers, k-nearest neighbor, Boosting, Decision trees (e.g., C4.5, Random forests), Bayesian networks, and Hidden Markov models.
[0106] Alternatively or in complement, the artificial intelligence technique may imply the use of one or several of the following elements: artificial neural network, Data clustering, Expectation-maximization algorithm, Self-organizing map, Radial basis function network, Vector Quantization, Generative topographic map, Information bottleneck method, andIBSEAD, rule learning algorithms such as Apriori algorithm, Eclat algorithm and FP-growth algorithm, hierarchical clustering, such as Single-linkage clustering and Conceptual clustering, partitional clustering such as K-means algorithm and Fuzzy clustering.
[0107] Alternatively or in complement, the artificial intelligence technique uses a reinforcement learning algorithm. Examples of reinforcement learning algorithms include, but are not limited to, temporal difference learning, Q-learning and Learning Automata.
[0108] Alternatively or in complement, the artificial intelligence technique uses a reinforcement learning algorithm. Examples of reinforcement learning algorithms include, but are not limited to, temporal difference learning, Q-learning and Learning Automata.
[0109] Alternatively or in complement, the artificial intelligence technique uses Data Preprocessing.
[0110] More specifically, the model is chosen among a linear model, a non-linear model, an ensemble model and a deep learning model.
[0111] A linear model is a model that uses linear relation(s) between the inputs and the outputs.
[0112] In the present case, the linear model is penalized multinomial regression or linear discriminant analysis
[0113] A non-linear model is a model that uses non-linear relation(s) between the inputs and the outputs.
[0114] As a specific example, it is hereinafter considered molecular classifying functions FMC forTCMR, assessed according to the international Banff 2019 classification.
[0115] Preferably, the other parameters relative to the recipient comprise at least one comorbidity parameter (COMP), at least one clinical parameter (CLIP), at least one biological parameter (BIOP) and / or at least one histological parameter (HISP).
[0116] Clinical relevance:
[0117] The prediction method as disclosed herein is useful for identifying patients (i.e. kidney transplant recipients) at risk of TCMR. Especially the prediction method of the invention is useful to predict patients with a high risk of TCMR. It is also useful to predict the severity and / or grade of TCMR.
[0118] The invention also relates to a method for diagnosing (especially early diagnosis) whether a kidney transplant recipient is afflicted by TCMR comprising:
[0119] a) quantifying the expression levels of a plurality of genes selected from the group consisting of CD96, CD45R0, SELL, MAPK3, CD27, IL16, CTLA4, CD84, LAIR1, TM4SF1,CD3E, CD247, HLADMB, MS4A6A, CD5, IKZF1, ZAP70, INPP5D, HLADMA, LEF1, ALOX5, STAT5A, PLAT, KDR, GZMK, ADAM8, ADAMDEC1, AGR2, AIM2, ANKRD22, AOAH, BATF, BCL2A1, BIRC3, BK large T Ag, BK VP1, BLK, BTK, BTLA, C1QA, C1QB, CALHM6, CAV1, CCL18, CCL19, CCL3 / L1, CCL4, CCL5, CCR2, CCR4, CCR5, CCR6, CCR7, CD163, CD19, CD1D, CD2, CD22, CD28, CD38, CD3D, CD3G, CD4, CD40LG, CD45RA, CD45RB, CD48, CD6, CD69, CD7, CD72, CD74, CD80, CD86, CD8A, CD8B, CIITA, CLEC4C, CSF2RB, CTSS, CTSW, CXCL13, CXCL9, CXCR3, CXCR4, CXCR5, CXCR6, DUSP2, EOMES, EZH2, FAM30A, FSLG, FCER1G, FCGRA1, FCGR3A / B, FCRL2, FGD2, FLT3, FOS, FOXP3, FPR1, GBP1, GBP2, GBP5, GIMAP5, GZMA, GZMB, GZMH, HLADPA1, HLADPB1, HLADQB1, HLADRA, HLADRB1, HLADRB3, HLAF, HLAG, IGOS, IDO1, IFI30, IFNG, IGHG1, IL10RA, IL21R, IL27RA, IL2RA, IL2RB, IL2RG, IL7R, IRF1, IRF4, IRF8, ISG20, ITGA4, ITGAX, ITGB2, JAK3, KLRB1, KLRC1, KLRG1, KLRK1, LAG3, LCK, LCP2, LILRB1, LILRB2, LILRB4, LST1, LTA, LTB, MIR155HG, MME, MMP9, MS4A1, MS4A4A, NFAM1, NFATC2, NKG7, NLRC5, NLRP3, NOD2, NPHS2, OR2 / 1P, PAX5, PDCD1, PDCD1LG2, PDPN, PHEX, PIK3CD, PIK3CG, PLAAT4, POU2AF1, PRF1, PSMB9, PSTPIP1, PTPN22, PTPN6, PTPN7, PTPRC, PTPRO, SAMHD1, SELPLG, SH2D1A, SIRPG, SLA, SLAMF6, SLAMF7, SLAM8, SLC4A1, SP140, SPIB, ST8SIA4, STAT1, STAT4, TAP1, TBX21, TCF7, TCL1A, TFF3, THEMIS, TIGIT, TLR2, TLR7, TLR8, TNFRSF18, TNFRSF1B, TNFRSF4, TNFRSF9, TNFRSF14, TNFSF8, TRAT1, TRDC and XCL1 / 2, in a sample obtained from the recipient;
[0120] b) implementing an algorithm on data comprising the quantified plurality of gene expression levels so as to obtain an algorithm output; and
[0121] c) determining the presence of TCMR from the algorithm output of step b).
[0122] The present invention relates to a method for identifying a biomarker for TCMR, the biomarker being a diagnosis biomarker of TCMR, a susceptibility biomarker of TCMR, a prognostic biomarker of TCMR or a predictive biomarker in response to the treatment of TCMR, the method comprising at least the steps of:
[0123] - carrying out the steps of the prediction method according to the invention for a first subject, to obtain a first predicted biological state, wherein the first subject is suffering from TCMR, - carrying out the steps of the prediction method according to the invention for a second subject, to obtain a second predicted biological state, wherein the second subject is not suffering from TCMR, and
[0124] - selecting a biomarker target based on the comparison of the first and second predicted biological states.In particular, the method of the present invention is particularly suitable for selecting a therapeutic regimen or determining if a certain therapeutic regimen is more appropriate for a patient identified as having a high risk of TCMR. Typically, protecting the allograft may be achieved using any suitable medical means known to those skilled in the art. In some embodiments, such reduction and protection comprise a therapeutic intervention with the subject such as administration of an immunosuppressive treatments, a corticosteroid and / or a T cell-depleting treatment. Said therapeutic intervention notably includes anti-thymocyte globulin (ATG), or a corticosteroid such as dexamethasone, prednisone or methylprednisone. Reciprocally, where the recipient is predicted a low risk of TCMR, the immunosuppressive therapy can be reduced in order to diminish the potential for drug toxicity.
[0125] Thus, the invention also relates to a method for selecting a therapeutic regimen for a kidney transplant recipient identified as having a high risk of TCMR comprising:
[0126] - performing the prediction method according to the invention, and
[0127] - if it is concluded that the kidney transplant recipient has a high risk of TCMR, then selecting a therapeutic regimen chosen from immunosuppressive treatments, corticosteroids and T cell-depleting treatments, preferably chosen from anti-thymocyte globulin, dexamethasone, prednisone and methylprednisone.
[0128] In some embodiments, the method of the present invention can be used to identify patients in need of frequent follow-up by a physician or clinician to monitor the therapeutic regimen. In some embodiments, a patient can be monitored using the method as disclosed herein, and if on a first testing (i.e. initial testing) the patient is identified as having a high risk of TCMR, the patient can be administered an appropriate therapeutic regimen, and on a second testing (i.e. follow-up testing), the patient is identified as having low risk of TCMR, the patient can be administered with a therapeutic regimen at a maintenance dose.
[0129] Thus, the invention also relates to a method for monitoring the efficacy of a therapeutic regimen for a kidney transplant recipient at risk of TCMR comprising:
[0130] i) performing the prediction method according to the invention with a first (i.e. initial) sample obtained from the recipient as a first testing (i.e. initial testing),
[0131] ii) if it is concluded from step i) that the kidney transplant recipient has a high risk of TCMR, then selecting a therapeutic regimen at a first dose, wherein said therapeutic regimen is chosen from immunosuppressive treatments, corticosteroids and T cell-depleting treatments, preferably chosen from anti-thymocyte globulin, dexamethasone, prednisone and methylprednisone, and
[0132] iii) performing the prediction method according to the invention with a second (i.e. follow-up) sample obtained from the recipient as a second testing (i.e. follow-up testing), after a time period after steps i) and ii),
[0133] wherein if it is concluded from step iii) that the kidney transplant recipient has a low risk of TCMR, then a maintenance dose of the therapeutic regimen ii) which is lower than the first dose, is selected.
[0134] Of course, the first and second samples obtained from the recipient are different.
[0135] Thus, the method of the present invention is particularly suitable for discriminating responder from non-responder.
[0136] As used herein the term “responder” in the context of the present disclosure refers to a subject that will achieve a response, i.e. the risk of TCMR does show a reduction. A non-responder subject includes subjects for whom the risk of TCMR does not show any reduction or improvement after the treatment.
[0137] The invention also relates to a method for discriminating a responder kidney transplant recipient from a non-responder kidney transplant recipient at risk of TCMR comprising: i) performing the prediction method according to the invention with a first sample obtained from the recipient before any treatment, as a first testing which results in a first probability of TCMR,
[0138] ii) performing the prediction method according to the invention with a second sample obtained from the recipient after a treatment, as a second testing which results in a second probability of TCMR,
[0139] iii) comparing the first and second probabilities of TCMR,
[0140] wherein if it is concluded from step iii) that the recipient has a low risk of TCMR, then concluding that said recipient is a responder,
[0141] and wherein if it is concluded from step iii) that the recipient has a high risk of TCMR, then concluding that said recipient is a non-responder.
[0142] Typically if the second probability is significantly lower than the first probability, then it can be concluded from step iii) that the recipient has a low risk of TCMR (and is a responder). Typically if the second probability is not significantly different from the first probability, or is higher than the first probability, then it can be concluded from step iii) that the recipient has a high risk of TCMR (and is a non-responder).In some embodiments, screening patients for identifying patients having a high risk of TCMR using the prediction method as disclosed herein is also useful to identify patients most suitable or amenable to be enrolled in clinical trial for assessing a therapy for management of allograft, which will permit more effective subgroup analyses and follow-up studies. Furthermore, the prediction method as disclosed herein can be suitable for monitoring patients enrolled in a clinical trial to provide a quantitative measure for the therapeutic efficacy of the therapy which is subject to the clinical trial. Accordingly, the output of the algorithm (e.g. the score) represents a suitable surrogate marker for use in a clinical trial for assessing the efficiency of a particular therapy.
[0143] In some embodiments, screening patients for identifying patients having a high risk of TCMR using the prediction method as disclosed herein is also useful to give biological information about said patients. It can also inform about the biological pathways that are involved for a given patient, and / or inform about cell types and / or immunological pathways that are involved.
[0144] In some embodiments, screening patients for identifying patients having a high risk of TCMR using the prediction method as disclosed herein is also useful to identify novel TCMR treatments (different from graft) and / or for drug repurposing.
[0145] The invention will be further illustrated by the following figures and examples. However, these examples and figures should not be interpreted in any way as limiting the scope of the present invention.
[0146] EXAMPLES
[0147] The following examples are to be considered illustrative and not limiting on the scope of the present disclosure described above.
[0148] Example 1: Methods
[0149] The inventors followed the STROBE (Strengthening the Reporting of Observational Studies in Epidemiology) statement checklist for the report of observational cohort studies12. Moreover, they adhered the SAGER (Sex and Gender Equity Research) guidelines for reporting sex and gender. Sex was self-reported by participants13.Clinical and biological data
[0150] At time of transplantation, clinical and biological data were collected related to i) recipient characteristics, ii) donor characteristics, iii) transplant characteristics, and iv) immunosuppressive treatment. At time of kidney allograft biopsies, a standardized transplant assessment was performed, comprising i) clinical examination, ii) immunosuppressive treatment, and iii) blood and urinary analyses for standard of care laboratory parameters.
[0151] Immunological phenotyping
[0152] Kidney transplant recipients were tested for the presence of circulating donor-specific anti-HLA-A, -B, and -DR antibodies at baseline and each follow-up visit with single antigen flow bead assays (One Lambda, Inc., Canoga Park, CA, USA). Beads with a normalized mean fluorescence intensity (MFI) of greater than 500 units were considered positive. Immunodominant donor-specific antibody (DSA) was defined as the DSA with the highest MFI.
[0153] HLA typing of recipients and donors were respectively performed by next-generation sequencing (NGS-go, GenDx, The Netherlands) and medium to high resolution sequencespecific primer (Linkage Biosciences, One Lambda, West Hills, CA, USA) in local HLA labs.
[0154] Histological and immunohistochemical phenotyping
[0155] Kidney allograft biopsies were routinely performed at each follow-up visit according to local centers’ practice. One kidney core was paraffin-embedded and formalin-fixed for histological analysis. C4d staining was performed by immunohistochemistry on paraffin-embedded tissue or by immunofluorescence on frozen tissue according to local practices. Biopsies were assessed by local pathologists, and centrally reviewed according to the international and standardized Banff 2019 classification for kidney allograft rejection14.
[0156] Bulk tissue transcriptomic profiling
[0157] The inventors performed bulk tissue transcriptomic profiling of formalin-fixed paraffin-embedded (FFPE) kidney allograft biopsies using the Banff Human Organ Transplant (B-HOT) panel on the NanoString nCounter® platform3. The B-HOT panel, approved by the international Banff consortium and developed by the Banff Molecular Diagnostics Working Group, targets 758 genes crucial for immune response and injury in solid organ transplants. It also includes housekeeping genes and both positive and negative controls to ensure data quality and normalization3.Following the quality assessment, the raw expression counts were normalized using the nSolver software in two steps15. First, positive control normalization was conducted using the geometric mean of positive control probes to adjust for technical variations in hybridization efficiency and purification across samples. Second, content normalization was performed using housekeeping genes to adjust for differences in RNA input amount and sample quality. All nCounter Gene Expression assays were normalized using nSolver software. The normalized gene expression data were not log transformed for the analysis.
[0158] Outcomes of interest
[0159] The primary outcomes of interest were T cell-mediated rejection, assessed according to the international Banff 2019 classification16. TCMR was defined as acute TCMR or chronic active TCMR.
[0160] Descriptive statistics
[0161] For continuous variables, means and standard deviations (SDs) or medians and interquartile ranges (IQRs) were used. The inventors compared means and proportions between groups using Student’s t-test, analysis of variance (ANOVA) (or Mann-Whitney test and Kruskal-Wallis if appropriate), or the chi-squared test (or Fisher’s exact test if appropriate). Values of p<0.05 were considered significant and all tests were two tailed.
[0162] Modelling pipeline
[0163] The study includes seven steps: 1) data pre-process, 2) original data-based molecular classifiers development, 3) conditional tabular GAN (CTGAN) model development, 4) synthetic data generation, 5) feature importance analysis, 6) learning curves generation, 7) machine learning classifiers developments, and 8) classifiers’ assessments. Figure 1 (diagram) demonstrates the overall flow of the study. Transcriptomic data was complete without any missing data. Phenotypic data had missing data and they were used as-is for CTGAN model development.
[0164] The inventors followed the TRIPOD (Transparent Reporting of a multivariable prediction model for Individual Prognosis Or Diagnosis) + Al (Artificial Intelligence) statement for the reporting of the development and validation of the classifiers17.
[0165] Development and validation of original data-based molecular classifiers The derivation cohort was split into development and test sets with 8:2 ratio. The development set was used to perform feature importance analysis, learning curvegeneration for the final feature selection, and machine learning-based classifiers generation as follows.
[0166] Feature importance analysis
[0167] The inventors performed feature selection and importance analysis using the Boruta algorithm, a wrapper and Monte Carlo method which uses multiple random forests and shadow features. Genes were included as features to predict the binarized AMR and TCMR18. The inventors performed a maximum of 2000 iterations of random forests. Among the selected features considered predictive by the Boruta algorithm, they calculated median importance to explore the feature importance changes by datasets.
[0168] Learning curves generation
[0169] Furthermore, they performed forward stepwise feature selection method to increment the number of features from one to the all Boruta selected features. For this process, they used random forest algorithm as well. Learning curves were generated to observe the changes of predictive performance of the random forest classifiers by adding features. A grid search on fixed one-fold cross-validation was performed to optimize the precision recall area under the curve (PRAUC) for each learning curve. The one-fold cross-validation was randomly selected from the development set and was used for hyperparameter optimization in order to tune the molecular classifiers based on the original data instead of synthetic data. The predictive performance was assessed in the internal validation test set using the PRAUC and the receiver operating characteristic area under the curve (ROCAUC). Features, which were chosen from the forward stepwise selection with the threshold of PRAUC reaching 0.9 quantile of the curve to make the classifiers have enough discriminative performance, were used for the final machine learning developments.
[0170] Machine learning developments
[0171] With the selected features for each set by the Boruta algorithm and the forward stepwise selection through the learning curves, the inventors generated machine learningbased classifiers (random forest) to predict the outcomes of TCMR.19The classifier hyperparameter optimization process was performed with grid search on the same fixed one-fold cross-validation used during the learning curves generation. During the optimization, PRAUC was used as a metric to maximize the performance. All continuous normalized gene expression data were standardized to have mean of zero and a standard deviation of one.Validation of molecular classifiers’ performances
[0172] Molecular classifiers’ performances were assessed in the external validation cohort. To measure the discrimination performance of the machine learning-based molecular classifiers, the inventors used multiple metrics for comprehensive assessments: 1) ROCAUC, 2) PRAUC, 3) accuracy, 4) balanced accuracy (average of sensitivity and specificity), 5) F1 score (harmonic mean of precision and recall), 6) negative predictive value (NPV), 7) positive predictive value (PPV), 8) sensitivity, 9) specificity, and 10) Brier Score. Calibration was assessed with calibration plots. For ROCAUC, PRAUC, and the Brier Score, probabilistic predictions were used to estimate the performance. For the others, cut-offs were calculated to maximize the F1 score on the test set for hard classification. For each assessment, mean of the estimates and 95% confidence intervals were calculated with 1,000 bootstrapping on the external validation cohort.
[0173] Conditional tabular GAN model development
[0174] The development set was used to train CTGAN model. 500 epochs (i.e. iteration over the entire development set) were performed to train the model. Parameters including genes (except for housekeeping genes), donor and recipient demographics, clinical history, biological and histological parameters, institutions, anonymized identifications, and the kidney allograft rejection diagnosis were used to train the CTGAN model. Loss values are measured to observe the model development. The CTGAN model was developed with SDV Python Package (version 1.9. O)20.
[0175] Synthetic data generation
[0176] Using the derived CTGAN model, a synthetic data was generated. The inventors generated synthetic cohort with twice the size of the development set (n>1,000) to increase the machine learning predictive power and robustness of the downstream analyses. One sample corresponds to a biopsy and no multiple biopsies were generated per patient. The synthetic cohort was sliced by the size of 25% of the development set to create eight synthetic datasets. These eight synthetic datasets are inclusive (i.e. the second dataset comprises the first and the second slices and the third dataset comprises the first, the second, and the third slices, and so on), allowing for a progressive accumulation of data. The eight synthetic datasets were attached to the end of the development set to generate augmented datasets. Comprising the development and the eight augmented datasets, the total of nine datasets were prepared and analyzed in the study. Each synthetic and augmented dataset is specified in tables, figures, and text using the notation GanN and GanAugN, respectively, where N refers to the percentage of sample size of thedevelopment set. For example, Gan50 refers to a synthetic set, which includes 50% of the sample size of the development set (the first two slices). Likewise, GanAug50 refers to an augmented set, which Gan50 set was attached to the end of the development set. Thus, the eight synthetic datasets are from Gan25 (the size of 25% of the development set) to Gan200 (the size of 200% of the development set). Likewise, the smallest augmented dataset is denoted as GanAug25 and the largest is denoted as GanAug200.
[0177] Assessment of similarities between synthetic and original data
[0178] The statistical similarities between the development set and the synthetic sets were assessed using SDV Python Package to assess Kolmogorov-Smirnov (KS) statistic complement and Correlation Similarity for continuous variables and Total Variance (TV) Distance complement and Contingency Similarity and for categorical variables20-22.
[0179] The definition of each metric is as follows:
[0180] For univariate comparison, column shapes were measured for all columns in KSComplement for numerical value and TVComplement for categorical value. For bivariate comparison (column pair trends), Correlationsimilarity was used for numerical pair of values and Contingencysimilarity was used for categorical pair of values or numerical and categorical pair of values.
[0181] Due to the rounding during calculation, some scores might be little bit out of ranges.
[0182] KSComplement
[0183] This metric uses the Kolmogorov-Smirnov (KS) statistic (Massey Jr. FJ. The Kolmogorov-Smirnov Test for Goodness of Fit. Journal of the American Statistical Association 1951; 46: 68-78). It is a measure of the difference between a sample distribution and a reference or theoretical distribution. In this study, the inventors compare the synthetic cohort to the original derivation cohort. A numerical distribution is converted into its cumulative distribution function (CDF). The KS statistic is the maximum difference between the two CDFs.
[0184] KSComplement is defined as 1 - KS statistic.
[0185] This metric ignores missing values and is meant for continuous and numerical data. The score ranges from 0 (worst, the two datasets are as different as they can be) to 1 (best, the two datasets are the same).
[0186] TVComplement
[0187] This metric computes the Total Variation Distance (TVD) between the real and synthetic variables by determining the probability of each category value and using it tocompare the differences in probabilities. The formula below demonstrates how the TVD statistic achieves this comparison.
[0188]
[0189] CL) 6 fl
[0190] TVComplement score is defined as 1 - 8 R, S Here, co describes all the possible categories in a column, Q. Meanwhile, R and S refer to the real and synthetic frequencies for those categories.
[0191] This metric computes the similarity of a real value variable and a synthetic This metric ignores missing values and is meant for discrete and categorical data. The score ranges from 0 (worst, the two datasets are as different as they can be) to 1 (best, the two datasets are the same).
[0192] CorrelationSimilarity
[0193] This metric computes a correlation coefficient on the real and synthetic data, R and S, respectively. A and B represents a pair of variables.
[0194] . |*^4, B—^4, B |
[0195] score = 1 —J-L
[0196]
[0197] 2
[0198] The metric supports both the Pearson correlation coefficient and the Spearman’s rank correlation coefficient. The score ranges from 0.5 (worst, the two correlations are as different as they can be) to 1 (best, the two correlations are the same).
[0199] ContingencySimilarity
[0200] This metric computes the similarity of a pair of categorical columns between the real and synthetic datasets.
[0201] The formula below summarizes the scoring process, where a and (8 describe all the possible categories in variable A and B, respectively and R and S refer to the real and synthetic frequencies for those categories, respectively.
[0202] score =1- 1 \SaiP- RaiP|
[0203]
[0204] aEA PEB
[0205] The score ranges from 0 (worst, the contingency table is as different as can be) to 1 (best, the contingency table is exactly the same between the real and synthetic data).
[0206] Development and validation of augmented data-based molecular classifiers With the eight augmented datasets, they inventors followed the same machine learning pipeline as the original data-based molecular classifiers: 1) feature importanceanalysis, 2) learning curves generation, 3) machine learning developments for the eight augmented datasets, and 4) validation of the eight molecular classifiers’ performances.
[0207] Software and packages
[0208] Descriptive analyses and machine learning classifier analyses were conducted using R (version 4.3.1, R Foundation for Statistical Computing) and RStudio (version 2023.9.0.463). Synthetic data were generated using Python (version 3.10) and JupyterLab (version 4.0.9). R packages used for descriptive, data, and machine learning analyses were: tidyverse (version 2.0.0), ggplot2 (version 3.4.3), forcats (version 1.0.0), Boruta (version 8.0.0), yardstick (version 1.2.0), rstatix (version 0.7.2), pROC (version 1.18.4), stringr (version 1.5.0), caret (version 6.0-94), caretEnsemble (version 2.0.3), randomForest (version 4.7-1.1), rsample (version 1.2.0), compareGroups (version 4.7.1), viridis (version 0.6.4), foreach (version 1.5.2), doParallel (version 1.0.17), CGPfunctions (version 0.6.3), and cutpointr (version 1.1.2). Python packages used for synthetic data generation were: numpy (version 1.26.2), torch (version 2.1.1), sdv (version 1.9.0), and pandas (version 2.1.3).
[0209] Example 2: Results
[0210] Baseline characteristics of the derivation cohort
[0211] The baseline characteristics of the derivation cohort are in Table 1.
[0212] Table 1. Baseline characteristics of the derivation cohort
[0213] **Characteristic** **N = 868** **Derivation** N = 441 **Validation** N = 427 **Internal Validation** N = 186 **External Validation** N = 241 N = 441 N = 427 N = 186 N = 241 Age [years, mean(sd)] 49 (16) 50 (14) 48 (18) 49 (15) 47 (20) Gender [n(%)]
[0214] F 337 (40%) 177 (40%) 160 (39%) 73 (39%) 87 (40%) M 510 (60%) 264 (60%) 246 (61%) 113 (61%) 133 (60%) Time to biopsy [months,
[0215] mean(sd)] 27 (47) 21 (39) 34 (53) 18 (31) 47 (63) Biopsy type [n(%)]
[0216] For cause 436 (51%) 192 (45%) 244 (58%) 88 (48%) 156 (65%) Protocol 419 (49%) 239 (55%) 180 (42%) 95 (52%) 85 (35%) Creatininemia at biopsy
[0217] [micromol / L, mean(sd)] 178 (138) 177 (146) 180 (128) 182 (130) 179 (126) DSA at biopsy [n(%)]
[0218] no 425 (49%) 259 (59%) 166 (39%) 109 (59%) 57 (24%) unknown 129 (15%) 19 (4.3%) 110 (26%) 11 (5.9%) 99 (41%)yes 314 (36%) 163 (37%) 151 (35%) 66 (35%) 85 (35%) Final Diagnosis [n(%)]
[0219] AMR 295 (34%) 144 (33%) 151 (35%) 61 (33%) 90 (37%) nBKv 17 (2.0%) 12 (2.7%) 5 (1.2%) 5 (2.7%) 0 (0%) Normal 204 (24%) 84 (19%) 120 (28%) 35 (19%) 85 (35%) Others 172 (20%) 104 (24%) 68 (16%) 44 (24%) 24 (10.0%) TCMR 180 (21%) 97 (22%) 83 (19%) 41 (22%) 42 (17%)
[0220] *N refers to the number of patients with available data
[0221] Baseline characteristics of the synthetic sets and statistical similarities
[0222] The development set was used to develop a CTGAN model to create synthetic sets. The details of the hyperparameters results are as follows: 'enforce_min_max_values': True; 'enforce_rounding': True; 'locales': None; 'embedding_dim': 128; 'generator_dim': (256, 256); 'discriminator_dim': (256, 256); 'generator_lr': 0.0002; 'generator_decay': 1e-06; 'discriminator_lr': 0.0002; 'discriminator_decay': 1e-06; 'batch_size': 500; 'discriminator_steps': 1; 'log_frequency': True; 'verbose': True; 'epochs': 500; 'pac': 10; 'cuda': True.
[0223] Eight synthetic sets were generated, sample sizes of 25%, 50%, 75%, 100%, 125%, 150%, 175%, and 200% of the development set. Overall, baseline characteristics show stable values across the synthetic sets.
[0224] Between the entire synthetic set (i.e. Gan200) and the original development set, the inventors assessed univariate and bivariate similarity scores. Overall, the two datasets show high similarity.
[0225] Original data-based molecular classifiers for kidney allograft rejection
[0226] The molecular classifiers based on the development set with selected genes were assessed on the external validation cohort. For TCMR, the molecular classifier showed discrimination performance with ROCAUC of 0.864 (0.813 – 0.906), PRAUC of 0.701 (0.597 – 0.785), and Brier Score of 0.122 (0.107 - 0.138).
[0227] Feature importance analysis
[0228] For TCMR, the genes associated with T cell activation and function considered important in all nine original development and augmented sets were: CD84, LAIR1, CD96, TM4SF1, CD3E, CD27, CD247, CD45R0, HLADMB, IL16, MS4A6A, SELL, MAPK3, CD5, IKZF1, CTLA4, ZAP70, INPP5D, HLADMA, LEF1, ALOX5, STAT5A, PLAT, KDR, and GZMK. Genes related to immune cell recruitment and chemotaxis (e.g., CCR2), immunecell interactions (e.g., CD40L), inflammatory and immune response (e.g., IL1B, IL7), and cell adhesion and migration (e.g., ITGA4, ITGB2) were only considered important features in the augmented datasets (Figure 2).
[0229] Learning curves analysis
[0230] With the selected features by Boruta algorithm, learning curves were generated with random forest algorithm. Incremental forward stepwise feature selection was performed and assessed on the test set (Figure 3).
[0231] For TCMR, the learning curve on the development set shows steep followed by stable increases in performance in both PRAUC and ROCAUC. At the performance reaching 0.9 quantile of PRAUC, numbers of features selected for the final molecular classifier derivation were 63 for the development set.
[0232] For TCMR, GanAug datasets show similar pattern in performance in both PRAUC and ROCAUC All learning curves show stable performance. At the performance reaching 0.9 quantile of PRAUC, numbers of features selected for the final molecular classifier derivation are detailed in Supplementary Table below.Supplementary Table Selected genes to develop TCMR molecular classifiers by dataset
[0233] Original GanAug25 GanAug50 GanAug75 GanAug100 GanAug125 GanAug150 GanAug175 GanAug200 1 LCK CD96 IPSMB8 PSMB8 IPSMB8 CD96 CD96 BATF CD27 2 CD84 SELL CD96 IL16 IL16 IL16 BATF CD84 ICAM2 3 PPM1F IIKZF1 CDH13 CD27 CD96 GZMK IL16 CD96 IIKZF1 4 LAIR1 CD8A CD8A SELL GZMK C1QA CD27 CD27 CD84 5 CD3D CDH13 SELL CD96 HLAB HLAB CD84 IIKZF1 ADGRL4 6 SPRY4 IPTPRC IIFNAR2 AOAH CD27 PSMB8 PSMB8 PDPN DNMT1 7 CD96 CD5 CTLA4 HLAB AOAH CD84 IKZF1 PSMB8 ST8SIA4 8 IGF1R IIFNAR2 GBP5 TM4SF1 SELL PDPN ST8SIA4 NPDC1 CD96 9 TM4SF1 MS4A6A IL18RAP CXCL9 BATF AOAHI C1QA ST8SIA4 CD45R0 10 CD3E LEF1 BMP6 CTLA4 GBP5 SELL GZMK ICAM2 PDPN 11 CD27 CTLA4 TM4SF1 HLADMA PDPN BATF ICAM2 CD247 CD247 12 CD247 CD45R0 GZMK GBP5 CTLA4 CD28 STAT5A IL16 IL1B 13 CD45R0 GBP5 CD3E PLAT CD3E PLAT AOAH MS4A6A BATF 14 BMP6 BCL2L1 MS4A6A ALOX5 MAPK3 CD45R0 CD247 DNMT1 IL16 15 HLADMB FGD2 CD163 PECAM1 STAT4 CTLA4 PDPN TM4SF1 TM4SF1 16 LTB SIH2D1A HLAB CD3E IIKZF1 ICAM2 TGFBR1 MMP12 PPM1F 17 LAG3 GZMK HLADMB PLAAT4 BCL2A1I ST8SIA4 MMP12 GZMK CD209 18 CXCR4 CD3E PSMB9 MS4A6A CDH13 CD27 CXCL13 STAT5A FCGR2A 19 IL16 BMP6 CD84 LAG3 PLAAT4 PLAAT4 MS4A6A AOAH MMP1I2 20 CD28 PPM1F MAPK3 CDH13 PECAM1 CD3G HLAB SELL PLA1IA 21 CD8A IGF1IR ADGRL4 LAIR1 IIFNAR2 PPM1F TM4SF1 CTLA4 PSMB8 22 PTPN7 IL18RAP CD5 CD3D PPM1F MAPK3 MAPK3 CD28 HLADMA 23 SOD2 ALOX5 BTLA CD163 CD28 TGFBR1 CD28 MAPK3 CD28 24 IL2RB CD3D IL16 SH2D1A ADGRL4 LAIIR1 PLAT CD45R0 MAPK3 25 BCL2L1 MAPK3 CD27 CD8A CD84 INPDC1 SLA SELL 26 MS4A6A KDR CD45R0 STAT5A CD45R0 SLA STAT1 CTLA4 27 SELL STAT5A ITGB2 IIFNAR2 ST8SIA4 LAIR1 CXCL13 TGFBR1 28 CD8B LAG3 GZMK HLADMA THBD PLAT 29 ADGRL4 ALOX5 CD247 MS4A6A CD45R0 HLADMA 30 MAPK3 NKG7 HLADMB CMKLR1 SELL CD3E 31 ROBO4 PLAAT4 NOD1 LAIR1 HLADMB CXCR3 32 ENG IKZF1 IL21R CXCL9 IL18RAP CD209 33 IRF8 PECAM1 MAPK3 HLADMB PLAAT4 C1QA 34 CDH13 IPTPRC IPSMB9 SLAMF6 STAT1 LAIIR1
[0234]
[0235] 35 CD5 STAT5A CD45R0 LCK KDR FCGR2A36 TCF7 RASSF9 KDR PLAT IL21R IL18BP 37 IKZF1 INPP5D CD247 ADGRL4 HLAB 38 IFNGR1 PLAT HAVCR2 AQP2 IL1B 39 CHCHD1O KRT19 IL21R TIGIT 40 TGFB2 SIH2D1A ALOX5 CD40LG 41 KLRK1 HLADMA CD5 HMGB1 42 CRIP2 HLADRA AQP2 CD5 43 IFNAR2 PTPN7 CD8A PLAAT4 44 AOAH IRF1 SLA IL21R 45 GNG11 CXCR3 PTPN7 BCL2A1 46 IL10RA CD3D NKG7 NKG7 47 BTLA EPAS1 TM4SF1 CD4 48 CTLA4 CD209 LAG3 LEF1 49 PTPRC LILRB4 C1QA DUSP2 50 KRT19 LEF1 BCL2L1 INPP5D 51 VEG FA STAT1 BTLA 52 CXCR3 ARG2 LEF1 53 FGD2 KDR BMP6 54 NLRC5 CD247 SH2D1A 55 IL7R TLR3 IL7R 56 ZAP70 SLA ZAP70 57 INPP5D THBD PTPRC 58 HLADRA ZAP70 IFNA1 59 HLADMA BMP2 KDR 60 IFI27 SPRY4 ATM 61 JAK2 TIGIT CXCL13 62 LEF1 PSMB9 6
[0236]
[0237] 3 RHOJ INPP5D 64 CD3DAugmented data-based molecular classifiers for TCMR
[0238] Synthetic biopsies were generated and merged to the development set to augment the data, for the GanAug25, GanAug50, GanAug75, GanAug100, GanAug125, GanAug150, GanAug175, and GanAug200 datasets, respectively. The performances of the molecular classifiers based on the augment data sets were assessed on the external validation cohort.
[0239] For TCMR, the molecular classifiers showed discrimination performance with ROCAUCs of 0.840 (0.784 - 0.887), 0.860 (0.807 - 0.904), 0.840 (0.781 - 0.886), 0.841 (0.786 - 0.886), 0.858 (0.808 - 0.901), 0.856 (0.803 - 0.900), 0.832 (0.778 - 0.878), and 0.841 (0.782 - 0.891), and PRAUCs of 0.684 (0.582 - 0.770), 0.704 (0.597 - 0.788), 0.641 (0.519 - 0.750), 0.675 (0.569 - 0.766), 0.690 (0.580 - 0.780), 0.693 (0.573 - 0.793), 0.646 (0.531 - 0.742), and 0.658 (0.535 - 0.757) for the GanAug25, GanAug50, GanAug75, GanAug100, GanAug125, GanAug150, GanAug175, and GanAug200 sets, respectively. The classifiers showed overall fit with Brier Scores of 0.125 (0.107 - 0.143), 0.114 (0.096 -0.133), 0.122 (0.104- 0.142), 0.124 (0.105- 0.144), 0.116 (0.098 -0.135), 0.121 (0.105 -0.139), 0.134 (0.115 - 0.153), and 0.129 (0.112 - 0.148) in the same order.
[0240] The calibration plots are available on Figure 4. Cut-offs were calibrated to maximize F1 score (harmonic mean of precision and recall) for hard classification (T cell-mediated rejection: Model based on Development set: 0.552; Model based on GanAug25 set: 0.536; Model based on GanAug50 set: 0.452; Model based on GanAug75 set: 0.550; Model based on GanAug100 set: 0.452; Model based on GanAug125 set: 0.484; Model based on GanAug150 set: 0.536; Model based on GanAug175 set: 0.498; Model based on GanAug200 set: 0.512).
[0241] Table 2. Performance in multimetrics in the external validation cohort
[0242] The table illustrates detailed performance metrics in the external validation cohort. The parentheses represent 95% confidence intervals calculated with 1000 bootstraps. PRAUC, precision recall area under the curve; ROCAUC, receiver operating characteristic area under the curve; NPV negative predictive values; PPV, positive predictive values.Model ROCAUC PRAUC Balanced Accuracy Accuracy F1 score NPV PPV Sensitivity Brier 0.864 0 701 0 839 0 700 0 550 0860 0 711 0.452 0 948 0 122 Development (0.813 - (0.597 - (0.802 - (0.639 - (0.432 - (0 821 - (0.577 - (0.341 - (0.920 - (0.107 - 0.906) 0 785) 0 878) 0 759) 0 652) 0.898) 0.837) 0.567) 0 974) 0.138) 0.840 0 684 0 845 0 718 0 580 0868 0 714 0491 0 945 0 125 GanAug25 (0.784 - (0.582 - (0.808 - (0.659 - (0.472 - (0 830 - (0.585 - (0.382 - (0.915 - (0.107 - 0 887) 0 770) 0 883) 0 774) 0 676) 0.907) 0 843) 0.600) 0 973) 0 143) 0.860 0 704 0 833 0 753 0 614 0891 0 622 0.611 0 896 0 114 GanAug50 (0.807 - (0.597 - (0.790 - (0.691 - (0.516 - (0 852 - (0.507 - (0.500 - (0.857 - (0.096 - 0 904) 0 788) 0 875) 0 810) 0 702) 0.926) 0 732) 0.718) 0 934) 0 133) 0.840 0.641 0.842 0.712 0 569 0.866 0 709 0.478 0.945 0.122 GanAug75 (0.781 - (0.519 - (0.805 - (0.654 - (0.464 - (0 828 - (0.580 - (0.366 - (0.916 - (0.104 - 0 886) 0 750) 0 880) 0 770) 0 667) 0.906) 0 828) 0.595) 0 971) 0 142) 0.841 0.675 0.807 0 742 0 585 0.891 0 553 0.625 0.858 0.124 GanAug100 (0.786 - (0.569 - (0.761 - (0.682 - (0.493 - (0 851 - (0.449 - (0.514 - (0.817 - (0.105 - 0 886) 0 766) 0 848) 0 797) 0 667) 0.926) 0 655) 0.729) 0 900) 0 144) 0.858 0.690 0.836 0 741 0 603 0.883 0 643 0.571 0.911 0.116 GanAug125 (0.808 - (0 580 - (0 796 - (0.681 - (0 504 - (0844 - (0.525 - (0 455 - (0.874 - (0.098 - 0 901) 0 780) 0 878) 0 800) 0 693) 0.918) 0 754) 0.682) 0 944) 0 135) 0.856 0 693 0 836 0 741 0 603 0883 0 642 0.572 0 911 0 121 GanAug150 (0.803 - (0.573 - (0.793 - (0.682 - (0.504 - (0846 - (0.525 - (0.458 - (0.876 - (0.105 - 0.900) 0 793) 0 878) 0 799) 0 691) 0.919) 0.750) 0.676) 0 944) 0.139) 0.832 0 646 0 810 0 749 0 595 0894 0 559 0.639 0 859 0 134 GanAug175 (0.776 - (0.531 - (0.767 - (0.687 - (0.500 - (0854 - (0.451 - (0.524 - (0.819 - (0.115 - 0.878) 0 742) 0 851) 0 805) 0 680) 0.930) 0.659) 0.740) 0 899) 0.153) 0.841 0 658 0 825 0 738 0 592 0884 0 604 0.583 0 892 0 129 GanAug200(0.782 - (0.535 - (0.781 - (0.678 - (0.493 - (0 844 - (0.486 - (0.466 - (0.852 - (0.112 -
[0243]
[0244] 0.891) 0 757) 0866) 0 795) 0 676) 0.920) 0 714) 0.690) 0 929) 0 148)DISCUSSION
[0245] In this international multicenter cohort study, the inventors developed and validated machine learning based molecular classifiers to predict T cell-mediated rejection. They showed that augmenting the original dataset with generative synthetic transcriptomic data can increase the performances of these molecular classifiers and capture additional biologically relevant predictive genes.
[0246] Original data-based molecular classifiers showed steep increase followed by stable performance during the learning curve analysis. In the external validation, the original data-based molecular classifiers showed good performance in multiple metrics including ROCAUC and PRAUC in the external validation cohort for predicting TCMR. Boruta algorithm captured TCMR associated genes as highly important features.
[0247] Molecular classifiers based on the synthetic data augmentation showed stable performance increment during the learning curve analysis. In the external validation, the derived molecular classifiers showed increased performance in the synthetically augmented data in TCMR (PRAUC 0.704 from 0.701). In other metrics, the molecular classifiers showed similar performance increase trend in the augmented datasets. Calibration plots showed comparable results across the development and augmented datasets.
[0248] The inventors found that the gene importance varies through the different sets of data by feature selection / importance analysis and learning curve generation. Compared to the original development set, the synthetically augmented sets allowed capturing genes biologically related to B cell activation, T cell activation and macrophage activation, NK-specific activation, and cellular injury and adaptive immune. This newly captured trend in the machine-generated augmented datasets enhances and expands upon previous studies, revealing additional biologically relevant genes for diagnosing kidney rejection23.
[0249] Incorporating GAN-based data augmentation into clinical practice alongside standard monitoring factors presents an opportunity to enhance cost-effectiveness and augment the breadth of data available for analysis24. By satisfying the growing demand for data and contributing to the development of more comprehensive predictive classifiers for TCMR diagnosis, this approach can lead to significant cost savings while simultaneously safeguarding the privacy of kidney allograft recipients. These advancements offer a promising avenue for improving the management of renal transplant patients by refining the clinical indications for biopsies and optimizing overall care protocols.
[0250] This study presents multiple strengths. First, the study cohort comprises a comprehensive set of multimodal datatypes including clinical, biological, histological, donor, recipient, and transplant parameters, in addition to transcriptomic data. This extensive dataset allows for a more holistic analysis and better understanding of kidney allograftrejection. Second, the inventors used multidisciplinary study methods. The study incorporates heuristic selection by kidney transplant experts from the Banff Human Organ Transplant (B-HOT) panel, ensuring the inclusion of the most relevant genes and markers3. Additionally, it integrates specialized expertise from nephrologists and pathologists for multimodal kidney rejection diagnosis. Advanced deep-learning techniques are employed to augment data, aiming to enhance the robustness and accuracy of the classifiers’ performance. Furthermore, multiple machine learning methods are utilized for feature selection, learning curve analysis, hyperparameter optimization, and classifiers’ generation. Finally, external validation cohorts are included to strengthen the reliability and generalizability of the study's findings. Validating classifiers on independent datasets ensures robust performance across different clinical settings.
[0251] In conclusion, the inventors derived and validated machine learning-based molecular classifiers for kidney allograft rejection and enhanced their performance with deep learning-driven synthetic data augmentation. Using synthetic transcriptomic data, they captured additional predictive genes and showed potential for reducing the necessary number of transcriptomic sequencings and costs of time-consuming procedures.References
[0252] 1. Solez, K. et al. International standardization of criteria for the histologic diagnosis of renal allograft rejection: The Banff working classification of kidney transplant pathology. Kidney International 44, 411-422 (1993).
[0253] 2. Schinstock, C. A. et al. Banff survey on antibody-mediated rejection clinical practices in kidney transplantation: Diagnostic misinterpretation has potential therapeutic implications.
[0254] 3. Mengel, M. et al. Banff 2019 Meeting Report: Molecular diagnostics in solid organ transplantation-Consensus for the Banff Human Organ Transplant (B-HOT) gene panel and open source multicenter validation.
[0255] 4. Halloran, P. F., Famulski, K. S. & Reeve, J. Molecular assessment of disease states in kidney transplant biopsy samples. Nature Reviews Nephrology 12, 534-548 (2016). 5. Loupy, A. et al. Molecular Microscope Strategy to Improve Risk...: Journal of the American Society of Nephrology.
[0256] 6. Frid-Adar, M. et al. GAN-based synthetic medical image augmentation for increased CNN performance in liver lesion classification. Neurocomputing 321, 321-331 (2018).
[0257] 7. Ju, L. et al. Leveraging Regular Fundus Images for Training UWF Fundus Diagnosis Models via Adversarial Learning and Pseudo-Labeling. IEEE Transactions on Medical Imaging 40, 2911-2925 (2021).
[0258] 8. Ktena, I. et al. Generative models improve fairness of medical classifiers under distribution shifts. Nature Medicine 30, 1166-1173 (2024).
[0259] 9. Goodfellow, I. et al. Generative Adversarial Nets, in Advances in Neural Information Processing Systems vol. 27 (Curran Associates, Inc., 2014).
[0260] 10. Xu, L. & Veeramachaneni, K. Synthesizing Tabular Data using Generative Adversarial Networks. Preprint at https: / / doi.org / 10.48550 / arXiv.1811.11264 (2018). 11. Xu, L., Skoularidou, M., Cuesta-Infante, A. & Veeramachaneni, K. Modeling Tabular data using Conditional GAN. in Advances in Neural Information Processing Systems vol.
[0261] 32 (Curran Associates, Inc., 2019).
[0262] 12. Elm, E. von et al. The Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) statement: guidelines for reporting observational studies. 13. Heidari, S., Babor, T. F., De Castro, P., Tort, S. & Curno, M. Sex and Gender Equity in Research: rationale for the SAGER guidelines and recommended use. Research Integrity and Peer Review 1, 2 (2016).
[0263] 14. Yoo, D. et al. An automated histological classification system for precision diagnostics of kidney allografts. Nature Medicine 29, 1211-1220 (2023).
[0264] 15. NanoString Technologies. nSolverTM 4.0 Analysis Software. 2018;5-98.. Loupy, A. et al. The Banff 2019 Kidney Meeting Report (I): Updates on and clarification of criteria for T cell- and antibody-mediated rejection.
[0265] . Collins, G. S. et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. (2024) doi: 10.1136 / bmj-2023-078378.
[0266] . Kursa, M. B. & Rudnicki, W. R. Feature Selection with the Boruta Package. Journal of Statistical Software 36, 1-13 (2010).
[0267] . Breiman, L. Random Forests. Machine Learning 45, 5-32 (2001).
[0268] . Patki, N., Wedge, R. & Veeramachaneni, K. The Synthetic Data Vault, in 2016 IEEE International Conference on Data Science and Advanced Analytics (DSAA) 399-410 (2016). doi: 10.1109 / DSAA.2016.49.
[0269] . Massey Jr., F. J. The Kolmogorov-Smirnov Test for Goodness of Fit. Journal of the American Statistical Association 46, 68-78 (1951).
[0270] . Gibbs, A. L. & Su, F. E. On Choosing and Bounding Probability Metrics. International Statistical Review / Revue Internationale de Statistique 70, 419-435 (2002).
[0271] . Halloran, P. F. et al. Real Time Central Assessment of Kidney Transplant Indication Biopsies by Microarrays: The INTERCOMEX Study.
[0272] . Motamed, S., Rogalla, P. & Khalvati, F. Data augmentation using Generative Adversarial Networks (GANs) for GAN-based detection of Pneumonia and COVID-19 in chest X-ray images.
[0273] . K, S. et al. Clinical validation and reproducibility of the Banff schema for renal allograft pathology. PubMed.
[0274] . Schinstock, C. A. et al. Banff survey on antibody-mediated rejection clinical practices in kidney transplantation: Diagnostic misinterpretation has potential therapeutic implications.
Claims
1. CLAIMS1. A method of predicting whether a kidney transplant recipient is at risk of T cell-mediated rejection (TCMR) comprising:a) quantifying the expression levels of a plurality of genes selected from the group consisting of CD96, CD45R0, SELL, MAPK3, CD27, IL16, CTLA4, CD84, LAIR1, TM4SF1, CD3E, CD247, HLADMB, MS4A6A, CD5, IKZF1, ZAP70, INPP5D, HLADMA, LEF1, ALOX5, STAT5A, PLAT, KDR, GZMK, ADAM8, ADAMDEC1, AGR2, AIM2, ANKRD22, AOAH, BATF, BCL2A1, BIRC3, BK large T Ag, BK VP1, BLK, BTK, BTLA, C1QA, C1QB, CALHM6, CAV1, CCL18, CCL19, CCL3 / L1, CCL4, CCL5, CCR2, CCR4, CCR5, CCR6, CCR7, CD163, CD19, CD1D, CD2, CD22, CD28, CD38, CD3D, CD3G, CD4, CD40LG, CD45RA, CD45RB, CD48, CD6, CD69, CD7, CD72, CD74, CD80, CD86, CD8A, CD8B, CIITA, CLEC4C, CSF2RB, CTSS, CTSW, CXCL13, CXCL9, CXCR3, CXCR4, CXCR5, CXCR6, DUSP2, EOMES, EZH2, FAM30A, FSLG, FCER1G, FCGRA1, FCGR3A / B, FCRL2, FGD2, FLT3, FOS, FOXP3, FPR1, GBP1, GBP2, GBP5, GIMAP5, GZMA, GZMB, GZMH, HLADPA1, HLADPB1, HLADQB1, HLADRA, HLADRB1, HLADRB3, HLAF, HLAG, ICOS, IDO1, IFI30, IFNG, IGHG1, IL10RA, IL21R, IL27RA, IL2RA, IL2RB, IL2RG, IL7R, IRF1, IRF4, IRF8, ISG20, ITGA4, ITGAX, ITGB2, JAK3, KLRB1, KLRC1, KLRG1, KLRK1, LAG3, LCK, LCP2, LILRB1, LILRB2, LILRB4, LST1, LTA, LTB, MIR155HG, MME, MMP9, MS4A1, MS4A4A, NFAM1, NFATC2, NKG7, NLRC5, NLRP3, NOD2, NPHS2, OR2 / 1P, PAX5, PDCD1, PDCD1LG2, PDPN, PHEX, PIK3CD, PIK3CG, PLAAT4, POU2AF1, PRF1, PSMB9, PSTPIP1, PTPN22, PTPN6, PTPN7, PTPRC, PTPRO, SAMHD1, SELPLG, SH2D1A, SIRPG, SLA, SLAMF6, SLAMF7, SLAM8, SLC4A1, SP140, SPIB, ST8SIA4, STAT1, STAT4, TAP1, TBX21, TCF7, TCL1A, TFF3, THEMIS, TIGIT, TLR2, TLR7, TLR8, TNFRSF18, TNFRSF1B, TNFRSF4, TNFRSF9, TNFRSF14, TNFSF8, TRAT1, TRDC and XCL1 / 2, in a sample obtained from the recipient;b) implementing an algorithm on data comprising the quantified plurality of gene expression levels so as to obtain an algorithm output; andc) determining the probability of TCMR from the algorithm output of step b).
2. The method according to claim 1, wherein step a) comprises quantifying the expression levels of a plurality of genes selected from the group consisting of CD96, CD45R0, SELL, MAPK3, CD27, IL16, CTLA4, CD84, LAIR1, TM4SF1, CD3E, CD247, HLADMB, MS4A6A, CD5, IKZF1, ZAP70, INPP5D, HLADMA, LEF1, ALOX5, STAT5A, PLAT, KDR and GZMK in a sample obtained from the recipient.
3. The method according to claim 1 or 2, wherein the expression levels of two or more genes selected from the group consisting of CD96, CD45R0, SELL, MAPK3, CD27, IL16, CTLA4, CD84, LAIR1, TM4SF1, CD3E, CD247, HLADMB, MS4A6A, CD5, IKZF1, ZAP70, INPP5D, HLADMA, LEF1, ALOX5, STAT5A, PLAT, KDR and GZMK, are quantified in step a).
4. The method according to any one of the preceding claims, wherein the expression levels of CD96 and one or more genes selected from the group consisting of CD45R0, SELL, MAPK3, CD27, IL16, CTLA4, CD84, LAIR1, TM4SF1, CD3E, CD247, HLADMB, MS4A6A, CD5, IKZF1, ZAP70, INPP5D, HLADMA, LEF1, ALOX5, STAT5A, PLAT, KDR and GZMK, are quantified in step a),preferably the expression levels of CD96, CD45R0, and one or more genes selected from the group consisting of SELL, MAPK3, CD27, IL16, CTLA4, CD84, LAIR1, TM4SF1, CD3E, CD247, HLADMB, MS4A6A, CD5, IKZF1, ZAP70, INPP5D, HLADMA, LEF1, ALOX5, STAT5A, PLAT, KDR and GZMK, are quantified in step a),preferably the expression levels of CD96, CD45R0, SELL and one or more genes selected from the group consisting of MAPK3, CD27, IL16, CTLA4, CD84, LAIR1, TM4SF1, CD3E, CD247, HLADMB, MS4A6A, CD5, IKZF1, ZAP70, INPP5D, HLADMA, LEF1, ALOX5, STAT5A, PLAT, KDR and GZMK, are quantified in step a),preferably the expression levels of CD96, CD45R0, SELL, MAPK3 and one or more genes selected from the group consisting of CD27, IL16, CTLA4, CD84, LAIR1, TM4SF1, CD3E, CD247, HLADMB, MS4A6A, CD5, IKZF1, ZAP70, INPP5D, HLADMA, LEF1, ALOX5, STAT5A, PLAT, KDR and GZMK, are quantified in step a).
5. The method according to any one of the preceding claims, wherein the expression levels of at least CD96, CD45R0, SELL and MAPK3, are quantified in step a), more preferably the expression levels of at least CD96, CD45R0, SELL, MAPK3, CD27, IL16 and CTLA4 are determined.
6. The method according to any one of the preceding claims, wherein the sample obtained from the recipient is a tissue sample that is obtained from the transplanted kidney, preferably chosen from biopsy samples and frozen sections taken for histological analysis.
7. The method according to any one of the preceding claims, wherein the expression levels of step a) are quantified by RNA-seq.
8. The method according to any one of the preceding claims, wherein the algorithm of step b) applies the quantified plurality of gene expression levels to a machine learning classifier designed to predict the risk of TCMR, preferably the machine learning classifier is a random forest algorithm.
9. The method according to any one of the preceding claims, wherein it predicts the risk of TCMR at 3, 5 and / or 7 years from the date of prediction.
10. A method for selecting a therapeutic regimen for a kidney transplant recipient identified as having a high risk of TCMR comprising:- performing the method according to any one of the preceding claims, and- if it is concluded that the kidney transplant recipient has a high risk of TCMR, then selecting a therapeutic regimen chosen from immunosuppressive treatments, corticosteroids and T cell-depleting treatments, preferably chosen from anti-thymocyte globulin, dexamethasone, prednisone and methylprednisone.
11. A method for monitoring the efficacy of a therapeutic regimen for a kidney transplant recipient at risk of TCMR comprising:i) performing the method according to any one of claims 1 to 9 with a first sample obtained from the recipient as a first testing,ii) if it is concluded from step i) that the kidney transplant recipient has a high risk of TCMR, then selecting a therapeutic regimen at a first dose, wherein said therapeutic regimen is chosen from immunosuppressive treatments, corticosteroids and T celldepleting treatments, preferably chosen from anti-thymocyte globulin, dexamethasone, prednisone and methylprednisone, andiii) performing the method according to any one of claims 1 to 9 with a second sample obtained from the recipient as a second testing, after a time period after steps i) and ii), wherein if it is concluded from step iii) that the kidney transplant recipient has a low risk of TCMR, then a maintenance dose of the therapeutic regimen ii) which is lower than the first dose, is selected.
12. A method for discriminating a responder kidney transplant recipient from a nonresponder kidney transplant recipient at risk of TCMR comprising:i) performing the method according to any one of claims 1 to 9 with a first sample obtained from the recipient before any treatment, as a first testing which results in a first probability of TCMR,ii) performing the method according to any one of claims 1 to 9 with a second sample obtained from the recipient after a treatment, as a second testing which results in a second probability of TCMR,iii) comparing the first and second probabilities of TCMR,wherein if it is concluded from step iii) that the recipient has a low risk of TCMR, then concluding that said recipient is a responder,and wherein if it is concluded from step iii) that the recipient has a high risk of TCMR, then concluding that said recipient is a non-responder.
13. A method for identifying a biomarker for TCMR, the biomarker being a diagnosis biomarker of TCMR, a susceptibility biomarker of TCMR, a prognostic biomarker of TCMR or a predictive biomarker in response to the treatment of TCMR, the method comprising at least the steps of:- carrying out the steps of the prediction method according to any one of claims 1 to 9 for a first subject, to obtain a first predicted biological state, wherein the first subject is suffering from TCMR,- carrying out the steps of the prediction method according to any one of claims 1 to 9 for a second subject, to obtain a second predicted biological state, wherein the second subject is not suffering from TCMR, and- selecting a biomarker target based on the comparison of the first and second predicted biological states.
14. Electronic device for obtaining at least one molecular classifying function (FMC) intended to be used to predict whether a kidney transplant recipient is at risk of TCMR, wherein the molecular classifying function (FMC) is adapted to take as input representative parameters of the expression levels of genes of the recipient and to provide as output the risk of TCMR of the recipient, the electronic device comprising:- an obtaining module configured to obtain an initial database (IDB), the initial database (IDB) comprising several elements, each element corresponding to a respective recipient and providing, for said respective recipient, several gene expression level parameters (GELP) and other parameters relative to the recipient,- a training module configured to train a conditional tabular generative adversarial network on the initial database (IDB) to output elements fulfilling at least one similarity condition with the elements of the initial database (IDB), to obtain a trained conditional tabular generative adversarial network,- a using module configured to use the trained conditional tabular generative adversarial network, to obtain synthetic elements, the set of the initial database (IDB) and the synthetic elements forming an augmented database (ADB),- a determining module configured to determine gene expression level parameters (GELP) of the augmented database (ADB) impacting a prediction of the risk of TCMR of the recipient by a molecular classifying function, to obtain determined gene expression level parameters (GELPDET),- a selecting module configured to select gene expression level parameters (GELP) among the determined gene expression level parameters (GELPDET), to obtain a set of selected gene expression level parameters (GELPSEL), the selecting module selecting said gene expression level parameters (GELPSEL) iteratively for each determined gene expression level parameter (GELPDET) by:- analyzing the impact of adding said determined gene expression level parameters (GELPDET) in a learning process of a molecular classifying function adapted to take as input at least one already selected gene expression level parameter (GELPSEL) and said determined gene expression level parameters (GELPDET) and to provide as output the risk of TCMR of the recipient, and- adding the determined gene expression level parameter (GELPDET) in case a selection criterion is fulfilled or discarding said determined gene expression level parameter (GELPDET) in case said selection criterion is not fulfilled,- a generating module configured to generate molecular classifying functions (FMCCAND) taking as input the set of selected gene expression level parameters (GELPSEL) and providing as output the risk of TCMR of the recipient, to obtain molecular classifying function candidates (FMCCAND), and- a choosing module configured to choose at least one molecular classifying function (FMC) among the molecular classifying function candidates (FMCCAND) based on a performance criterion, to obtain at least one molecular classifying function (FMC) intended to be used to predict a risk of TCMR of the recipient.