Method for obtaining a molecular classifying function for predicting a biological state of a subject and associated methods and devices

WO2026167187A1PCT designated stage Publication Date: 2026-08-13INST NAT DE LA SANTE & DE LA RECHERCHE MEDICALE (INSERM) +2
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2026-02-06
Publication Date
2026-08-13

Smart Images

  • Figure EP2026053214_13082026_PF_FP_ABST
    Figure EP2026053214_13082026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention concerns the field of molecular prediction, namely biological predictions based on the expression levels of a selection of genes of a subject. The availability of technical modules adapted to obtain the expression levels of a selection of genes in an easier way than before renders possible to develop molecular algorithms adapted to output a biological state of a subject based on the expression levels of a selection of genes. However, obtaining molecular algorithms is a cumbersome task. The inventors have therefore worked on a clever method for obtaining a molecular classifying function based on a cohort of patients. Such method is generic and can be used for any organ of the subject.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] METHOD FOR OBTAINING A MOLECULAR CLASSIFYING FUNCTION FOR PREDICTING A BIOLOGICAL STATE OF A SUBJECT AND ASSOCIATED METHODS AND DEVICES

[0002] FIELD OF THE INVENTION

[0003] The present invention concerns a method for obtaining a molecular classifying function intended to be used to predict a biological state of a subject.

[0004] The present invention also relates to a method for predicting a biological state of a subject using a molecular classifying function obtained by such a method for obtaining.

[0005] The present invention also relates to methods carrying out the steps of a method for predicting, the methods being selected among a method for predicting that a subject is at risk of suffering from a disease, a method for diagnosing a disease, a method for identifying a therapeutic target for preventing and / or treating a disease, a method for identifying a biomarker for a disease, a method for screening a compound useful as a medicament and a method for monitoring patients enrolled in a clinical trial.

[0006] The present invention also concerns associated devices involved in the previously methods, namely an electronic device, a system for predicting, a computer program product and a computer-readable medium.

[0007] BACKGROUND OF THE INVENTION

[0008] In the medical field, there is a growing interest for molecular classifications into routine pathologic and clinical practice.

[0009] Indeed, a molecular classification usually provides a practitioner with a link between gene expression level data and a disease. In other words, by studying the expression level of genes, the practitioner benefits from a precious tool to establish a diagnostic.

[0010] A known technique to establish such molecular classification is to collect multiple gene expression level data for a large number of patients and to carry out statistical analysis on the number of collected data is sufficient.

[0011] However, in practice, reaching such number is difficult because, on the one hand, the associated threshold strongly depends on the case and, on the other hand, data are generally split between several institutions and it is not possible to have an access to all the existing data. For rare illness, there is even not enough cases to reach the threshold.SUMMARY OF THE INVENTION

[0012] There is therefore a need for a method for obtaining a molecular classifying function intended to be used to predict a biological state of a subject, which can be operated even in cases the number of accessible data is low.

[0013] To this end, the specification describes a method for obtaining at least one molecular classifying function intended to be used to predict a biological state of a subject, a molecular classifying function being adapted to take as input several gene expression level parameters of the subject and provide as output the biological state of the subject, a gene expression level parameter being a parameter representative of the expression level of said gene for said subject, the method for obtaining being computer-implemented, the method comprising the steps of:

[0014] - obtaining an initial database, the initial database comprising several elements, each element corresponding to a respective subject and providing, for said respective subject, several gene expression level parameters, other parameters relative to the subject and the biological state of the subject,

[0015] - training a conditional tabular generative adversarial network on the initial database to output elements fulfilling at least one similarity condition with the elements of the initial database, to obtain a trained conditional tabular generative adversarial network,

[0016] - using the trained conditional tabular generative adversarial network, to obtain synthetic elements, the set of the initial database and the synthetic elements forming an augmented database,

[0017] - determining gene expression level parameters of the augmented database impacting a prediction of the biological state of the subject by a molecular classifying function, to obtain determined gene expression level parameters,

[0018] - selecting gene expression level parameters among the determined gene expression level parameters, to obtain a set of selected gene expression level parameters, the step of selecting comprising iteratively for several determined gene expression level parameters:

[0019] - analyzing the impact of adding said determined gene expression level parameters in a learning process of a molecular classifying function adapted to take as input at least one already selected gene expression level parameter and said determined gene expression level parameters and to provide as output the biological state of the subject, and

[0020] - adding the determined gene expression level parameter in case a selection criterion is fulfilled or discarding said determined gene expression level parameter in case said selection criterion is not fulfilled,- generating molecular classifying functions taking as input the set of selected gene expression level parameters and providing as output the biological state of the subject, to obtain molecular classifying function candidates, and

[0021] - choosing at least one molecular classifying function among the molecular classifying function candidates based on a performance criterion, to obtain at least one molecular classifying function intended to be used to predict a biological state of a subject.

[0022] According to further aspects of the method for obtaining, which are advantageous but not compulsory, the method for obtaining might incorporate one or several of the following features, taken in any technically admissible combination:

[0023] - the step of determining is carried out by applying a Boruta algorithm on the augmented database.

[0024] - the number of determined gene expression level parameters is comprised between 10 and 50, preferably comprised between 10 and 30.

[0025] - a ratio between a number of synthetic elements and a number of elements of the initial database is comprised between 0,1 % and 100%, preferably between 0.25% and 0.75% and advantageously equal to 50%.

[0026] - a ratio between a number of synthetic elements and a number of elements of the initial database is comprised between 100 % and 200%, preferably between 100% and 150% and advantageously equal to 125%.

[0027] - the molecular classifying function used in the step of determining and in the step of selecting is obtained by an hyperparameter optimization carried out on a same part of the initial database.

[0028] - the step of obtaining comprises receiving a cohort and extracting the initial database and a validating database from the received cohort, the performance criterion used at the step of choosing being assessed on the validating database.

[0029] - the extracting is carried out by splitting the received cohort into two parts, one part being the initial database, the other part being the validating database, said splitting being achieved with a splitting ratio comprised between 7:3 and 9:1.

[0030] - the other parameters relative to the subject comprise at least one comorbidity parameter, at least one clinical parameter, at least one biological parameter and at least one histological parameter.

[0031] - each parameter of the other parameters relative to the subject is chosen among a comorbidity parameter, a clinical parameter, a biological parameter and an histological parameter.The specification also relates a method for predicting a biological state of a subject, the method being carrying out by a system for predicting, the method for predicting comprising:

[0032] - providing gene expression level parameters of the subject,

[0033] - applying a molecular classifying function obtained by a method for obtaining as previously described on the provided gene expression level parameters, to output a predicted biological state of the subject.

[0034] The specification also concerns a method selected from the group consisting of: - a method for predicting that a subject is at risk of suffering from a disease, the method for predicting comprising at least the steps of:

[0035] - carrying out the steps of a method for predicting as previously described wherein the step of providing is achieved by receiving the gene expression level parameters, to obtain a predicted biological state of the subject, and

[0036] - predicting that the subject is at risk of suffering from the disease based on the predicted biological state of the subject,

[0037] - a method for diagnosing a disease to a subject, the method for diagnosing comprising at least the steps of:

[0038] - carrying out the steps of a method for predicting as previously described wherein the step of providing is achieved by receiving the gene expression level parameters, to obtain a predicted biological state of the subject, and

[0039] - diagnosing the disease based on the predicted biological state of the subject, - a method for identifying a therapeutic target for preventing and / or treating a disease, the method comprising at least the steps of:

[0040] - carrying out the steps of a method for predicting for a first subject, to obtain a first predicted biological state, wherein the first subject is suffering from the disease and the method for predicting is as previously described wherein the step of providing is achieved by receiving the gene expression level parameters,

[0041] - carrying out the steps of a method for predicting for a second subject, to obtain a second predicted biological state (S), wherein the second subject is not suffering from the disease and the method for predicting is as previously described wherein the providing is achieved by receiving the gene expression level parameters, and - selecting a therapeutic target based on the comparison of the first and second predicted biological states,

[0042] - a method for identifying a biomarker for a disease, the biomarker being a diagnosis biomarker of the disease, a susceptibility biomarker of the disease, a prognostic biomarkerof the disease or a predictive biomarker in response to the treatment of the disease, the method comprising at least the steps of:

[0043] - carrying out the steps of a method for predicting for a first subject, to obtain a first predicted biological state, wherein the first subject is suffering from the disease and the method for predicting is as previously described wherein the step of providing is achieved by receiving the gene expression level parameters,

[0044] - carrying out the steps of a method for predicting for a second subject, to obtain a second predicted biological state, wherein the second subject is not suffering from the disease and the method for predicting is as previously described wherein the providing is achieved by receiving the gene expression level parameters, and - selecting a biomarker target based on the comparison of the first and second predicted biological states,

[0045] - a method for screening a compound useful as a medicament, the compound having an effect on a known therapeutical target for preventing and / or treating a disease, the method comprising at least the steps of:

[0046] - carrying out the steps of a method for predicting for a first subject, to obtain a first predicted biological state, wherein the first subject is suffering from the disease and has received the compound and the method for predicting is as previously described wherein the step of providing is achieved by receiving the gene expression level parameters,

[0047] - carrying out the steps of a method for predicting for a second subject, to obtain a second predicted biological state, wherein the second subject is suffering from the disease and has not received the compound and the method for predicting is as previously described wherein the step of providing is achieved by receiving the gene expression level parameters, and

[0048] - selecting a biomarker target based on the comparison of the first and second predicted biological states, and

[0049] - a method for monitoring patients enrolled in a clinical trial to provide a quantitative measure for the therapeutic efficacy of the therapy which is subject to the clinical trial by carrying out the steps of a method for predicting for said patients, the method for predicting being as previously described wherein the step of providing is achieved by receiving the gene expression level parameters.

[0050] The specification also relates to an electronic device for obtaining at least one molecular classifying function intended to be used to predict a biological state of a subject, a molecular classifying function being adapted to take as input parameters representativeof the expression level of genes of the subject and provide as output the biological state of the subject, the electronic device comprising:

[0051] - an obtaining module configured to obtain an initial database, the initial database comprising several elements, each element corresponding to a respective subject and providing, for said respective subject, several gene expression level parameters, other parameters relative to the subject and the biological state of the subject,

[0052] - a training module configured to train a conditional tabular generative adversarial network on the initial database to output elements fulfilling at least one similarity condition with the elements of the initial database, to obtain a trained conditional tabular generative adversarial network,

[0053] - a using module configured to use the trained conditional tabular generative adversarial network, to obtain synthetic elements, the set of the initial database and the synthetic elements forming an augmented database,

[0054] - a determining module configured to determine gene expression level parameters of the augmented database impacting a prediction of the biological state of the subject by a molecular classifying function, to obtain determined gene expression level parameters, - a selecting module configured to select gene expression level parameters among the determined gene expression level parameters, to obtain a set of selected gene expression level parameters, the selecting module selecting said gene expression level parameters by iteratively for each determined gene expression level parameter:

[0055] - analyzing the impact of adding said determined gene expression level parameters in a learning process of a molecular classifying function adapted to take as input at least one already selected gene expression level parameter and said determined gene expression level parameters and to provide as output the biological state of the subject, and

[0056] - adding the determined gene expression level parameter in case a selection criterion is fulfilled or discarding said determined gene expression level parameter in case said selection criterion is not fulfilled,

[0057] - a generating module configured to generate molecular classifying functions taking as input the set of selected gene expression level parameters and providing as output the biological state of the subject, to obtain molecular classifying function candidates, and

[0058] - a choosing module configured to choose at least one molecular classifying function among the molecular classifying function candidates based on a performance criterion, to obtain at least one molecular classifying function intended to be used to predict a biological state of a subject.The specification also concerns a system for predicting a biological state of a subject, the system for predicting comprising:

[0059] - a providing module, the providing module being configured to provide gene expression level parameters of the subject, and

[0060] - an applying module, the applying module being configured to apply a molecular classifying function on the gene expression level parameters, which the providing module is configured to provide, to output a predicted biological state of the subject. The specification also relates to a computer program product comprising computer program instructions, the computer program instructions being loadable into a data-processing unit and adapted to cause execution of a method as previously described.

[0061] The specification also concerns a computer-readable medium comprising computer program instructions which, when executed by a data-processing unit, cause execution of a method as previously described.

[0062] BRIEF DESCRIPTION OF THE DRAWINGS

[0063] The invention will be better understood on the basis of the following description which is given in correspondence with the annexed figures and as an illustrative example, without restricting the object of the invention. In the annexed figures:

[0064] - figure 1 is a schematic representation of a system adapted to carry out a method for obtaining a molecular classifying function,

[0065] - figure 2 is a functional representation of the operating of the system of figure 1, - figure 3 is a flowchart illustrating an example of carrying out a method for obtaining a molecular classifying function,

[0066] - figure 4 schematically represents the elements of an augmented database used in the method illustrated by figure 3,

[0067] - figure 5 is a schematic representation of a system adapted to carry out a method for predicting a biological state of a subject,

[0068] - figure 6 is a flowchart illustrating an example of carrying out a method for predicting a biological state of a subject, and

[0069] - figures 7 to 13 are figures representative of the results obtained by the Applicant (see experimental section).

[0070] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0071] DESCRIPTION OF THE SYSTEMA system 20 and a computer program product 30 are represented on figure 1. The interaction between the computer program product 30 and the system 20 enables to carry out a method for obtaining a molecular classifying function FMC intended to be used to predict a biological state of a subject as will be described later. Such method for obtaining is named “obtaining method” in the remainder of the specification.

[0072] The obtaining method is a computer-implemented method.

[0073] The system 20 is a desktop computer. In variant, the system 20 is a rack-mounted computer, a laptop computer, a tablet computer, a PDA or a smartphone.

[0074] In specific embodiments, the system 20 is adapted to operate in real-time and / or is an embedded system.

[0075] In the case of figure 1, the system 20 comprises a calculator 32, a user interface 34 and a communication device 36.

[0076] The calculator 32 is electronic circuitry adapted to manipulate and / or transform data represented by electronic or physical quantities in registers of the calculator 32 and / or memories in other similar data corresponding to physical data in the memories of the registers or other kinds of displaying devices, transmitting devices or memoring devices.

[0077] As specific examples, the calculator 32 comprises a monocore or multicore processor (such as a CPU, a GPU, a microcontroller and a DSP), a programmable logic circuitry (such as an ASIC, a FPGA, a PLD and PLA), a state machine, gated logic and discrete hardware components.

[0078] The calculator 32 comprises a data-processing unit 38 which is adapted to process data, notably by carrying out calculations, memories 40 adapted to store data and a reader 42 adapted to read a computer-readable medium.

[0079] The user interface 34 comprises an input device 44 and an output device 46.

[0080] The input device 44 is a device enabling the user of the system 20 to input information or command to the system 20.

[0081] In figure 1, the input device 44 is a keyboard. Alternatively, the input device 44 is a pointing device (such as a mouse, a touch pad, a trackball and a digitizing tablet), a voicerecognition device, an eye tracker or a haptic device (motion gestures analysis).

[0082] The output device 46 is a graphical user interface, which is a display unit adapted to provide information to the user of the system 20.

[0083] In figure 1, the output device 46 is a display screen for visual presentation of output, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor.

[0084] In other embodiments, the output device is a printer, an augmented and / or virtual display unit, a speaker or another sound generating device for audible presentation ofoutput, a unit producing vibrations and / or odors or a unit adapted to produce electrical signal.

[0085] In other words, interaction with the user can also involve other sensory feedback mechanisms, including visual, auditory, or tactile feedback, and inputs can be received through various means, including acoustic, speech, or tactile input.

[0086] In a specific embodiment, the input device 44 and the output device 46 are the same component forming man-machine interfaces, such as an interactive screen.

[0087] The communication device 36 enables unidirectional or bidirectional communication between the components of the system 20. For instance, the communication device 36 is a bus communication system or an input / output interface.

[0088] The presence of the communication device 36 enables that, in some embodiments, the components of the system 20 be remote one from another.

[0089] The computer program product 30 comprises a computer-readable medium 48. The computer-readable medium 48 is a tangible device that can be read by the reader 42 of the calculator 32.

[0090] Notably, the computer-readable medium 48 is not transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, such as light pulses or electronic signals.

[0091] Such computer-readable storage medium 48 is, for instance, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device or any combination thereof.

[0092] As a non-exhaustive list of more specific examples, the computer-readable storage medium 48 is a mechanically encoded device such a punchcards or raised structures in a groove, a diskette, a hard disk, a ROM, a RAM, an EROM, an EEPROM, a magnetic-optical disk, a SRAM, a CD-ROM, a DVD, a memory stick, a floppy disk, a flash memory, a SSD or a PC card such as a PCMCIA.

[0093] A computer program is stored in the computer-readable storage medium 48. The computer program comprises one or more stored sequence of program instructions.

[0094] Such program instructions when run by the data-processing unit 38, cause the execution of steps of any method that will be described below.

[0095] For instance, the form of the program instructions is a source code form, a computer executable form or any intermediate forms between a source code and a computer executable form, such as the form resulting from the conversion of the source code via an interpreter, an assembler, a compiler, a linker or a locator. In variant, program instructions are a microcode, firmware instructions, state-setting data, configuration data for integrated circuitry (for instance VHDL) or an object code.Program instructions are written in any combination of one or more languages, such as an object oriented programming language (FORTRAN, C¨++, JAVA, HTML), procedural programming language (language C for instance).

[0096] Alternatively, the program instructions is downloaded from an external source through a network, as it is notably the case for applications. In such case, the computer program product comprises a computer-readable data carrier having stored thereon the program instructions or a data carrier signal having encoded thereon the program instructions.

[0097] In each case, the computer program product 30 comprises instructions, which are loadable into the data-processing unit 38 and adapted to cause execution of steps of any method described below when run by the data-processing unit 38. According to the embodiments, the execution is entirely or partially achieved either on the system 20, that is a single computer, or in a distributed system among several computers (notably via cloud computing).

[0098] In specific embodiments, the algorithm can be deployed in a computing system comprising a back-end component (e.g., a data server), middleware component (e.g., an application server), or front-end component (e.g., a client computer with a graphical user interface or Web browser) to enable user interaction. These components can interconnect via any form of digital communication network, such as a LAN (local area network) or WAN (wide area network), including the Internet. Such a computing system may consist of clients and servers, typically remote from each other, interacting through a communication network. The client-server relationship is defined by the computer programs running on the respective computers, establishing a client-server relationship.

[0099] In the specific embodiment of figure 2, the system 20 is implemented in the form of an electronic device 50 for obtaining a molecular classifying function FMC.

[0100] The electronic device 50 comprises an obtaining module 52, a training module 54, a using module 56, a determining module 58, a selecting module 60, a generating module 62 and a choosing module 64.

[0101] In the example of figure 2, the electronic device 50 comprises an information processing unit formed for example by a memory and a processor associated with the memory.

[0102] The obtaining module 52, the training module 54, the using module 56, the determining module 58, the selecting module 60, the generating module 62 and the choosing module 64 are each produced in the form of software, or a software brick, executable by the processor.The memory of the electronic device 50 is then able to store a software of obtaining, a software of training, a software of using, a software of determining, a software of selecting, a software of generating and a software of choosing.

[0103] The processor is then able to execute each of the software among the software of training, the software of using, the software of determining, the software of selecting, the software of generating and the software of choosing

[0104] In a variant not shown, the obtaining module 52, the training module 54, the using module 56, the determining module 58, the selecting module 60, the generating module 62 and the choosing module 64 are each produced in the form of a programmable logic component, such as an FPGA (abbreviation for Field Programmable Gate Array), or an integrated circuit, such as an ASIC (abbreviation for Application Specific Integrated Circuit).

[0105] When the electronic device 50 is produced in the form of one or more software programs, i.e. in the form of a computer program, also called a computer program product, it is also capable of being recorded on a medium, not shown, readable by a computer. The computer-readable medium is, for example, a medium capable of storing electronic instructions and of being coupled to a bus of a computer system. For example, the readable medium is an optical disk, a magneto-optical disk, a ROM memory, a RAM memory, any type of non-volatile memory (for example FLASH or NVRAM) or a magnetic card. A computer program comprising software instructions is then stored on the readable medium.

[0106] The actions that each module is configured to implement are now exemplified in reference to figure 3, which is a flowchart illustrating an implementation of the obtaining method.

[0107] Such obtaining method aims at obtaining a molecular classifying function, FMC which is intended to be used to predict a biological state of a subject.

[0108] In this specific example, the subject is a human being.

[0109] More generally, the subject is a living subject and notably an animal.

[0110] For instance, the subject is a mammal, and more specifically a rodent such a mouse. A biological state can be defined as a specific condition or set of characteristics of a living organism at a given time.

[0111] In other words, a biological state typically refers to the condition or status of an organism. This can include various physiological, biochemical, or molecular characteristics that define the current functioning or condition of the organism, it reflects how an organism is functioning and responding to its surroundings at any particular moment.

[0112] As a specific example, for the heart, specific examples of biological states may include normal functioning, hypertrophy, heart failure, ischemia, arrhythmia, inflammation and regeneration.For the kidney, biological states may reflect kidney function and health. Specific examples may be normal functioning, acute kidney injury, chronic kidney disease, glomerulonephritis, nephrotic syndrome, polycystic kidney disease, renal tubular acidosis and kidney stone formation.

[0113] As another case, for liver, one may cite as examples of biological state: normal functioning, hepatitis, cirrhosis, fatty liver disease, liver failure, liver regeneration, cholestasis and liver cancer.

[0114] By definition, a molecular classifying function FMC is adapted to take as input several gene expression level parameters of the subject and provide as output the biological state of the subject, a gene expression level parameter being a parameter representative of the expression level of said gene for said subject.

[0115] Gene expression level parameters should be construed as encompassing any parameter representative of the expression level, such as gene expression profiles, protein levels, or metabolite concentrations.

[0116] The molecular classifying function FMC can be qualified as a classifier and is here obtained by using an artificial intelligence technique.

[0117] An artificial intelligence technique consists in establishing a model (also named algorithm) based on data.

[0118] In particular, the artificial intelligence technique often implies learning the model. The term “machine learning” is thus employed to designate the fact that the model is learned by the machine based on data.

[0119] According to the case, the machine learning technique implies using a learning among a supervised learning, an unsupervised learning, a semi-supervised learning, a reinforcement learning, a self-learning, a feature learning, a sparse dictionary learning, an anomaly detection learning, a robot learning and association rules learning.

[0120] In particular, in the present example, the machine learning technique is a supervised learning technique, a semi-supervised learning technique or a reinforcement learning technique.

[0121] The model used in the artificial intelligence technique can be chosen from various models / algorithms, such as computational models and algorithms for classification, clustering, regression and dimensionality reduction, such as neural networks, genetic algorithms, support vector machines, k-means, kernel regression and discriminant analysis.

[0122] More generally, the artificial intelligence technique may imply the use of one or several of the following elements: sums, ratios, and regression operators, such ascoefficients or exponents, biomarker value transformations and normalizations (including, without limitation, those normalization schemes based on clinical parameters, such as clinical data 58, gender, age or ethnicity), rules and guidelines, statistical classification models, and neural networks, structural and syntactic statistical classification algorithms, and methods of risk index construction, utilizing pattern recognition features, including established techniques such as cross-correlation, Principal Components Analysis (PCA), factor rotation, Logistic Regression (LogReg), Linear Discriminant Analysis (LDA), Eigengene Linear Discriminant Analysis (ELDA), Support Vector Machines (SVM), Random Forest (RF), Recursive Partitioning Tree (RPART), as well as other related decision tree classification techniques, Shrunken Centroids (SC), StepAIC, Kth-Nearest Neighbor, Boosting, Decision Trees, Neural Networks, Bayesian Networks, Support Vector Machines, and Hidden Markov Models, among others.

[0123] Alternatively or in complement, the artificial intelligence technique may imply the use of one or several of the following elements: Average One-Dependence Estimators (AODE), Artificial neural network (e.g., Backpropagation), Bayesian statistics (e.g., Naive Bayes classifier, Bayesian network, Bayesian knowledge base), Case-based reasoning, Decision trees, Inductive logic programming, Gaussian process regression, Group method of data handling (GMDH), Learning Automata, Learning Vector Quantization, Minimum message length (decision trees, decision graphs, etc.), Lazy learning, Instance-based learning Nearest Neighbor Algorithm, Analogical modeling, Probably approximately correct learning (PAC) learning, Ripple down rules, a knowledge acquisition methodology, Symbolic machine learning algorithms, Subsymbolic machine learning algorithms, Support vector machines, Random Forests, Ensembles of classifiers, Bootstrap aggregating (bagging), boosting, regression analysis, Information fuzzy networks (IFN), statistical classification, AODE, Linear classifiers (e.g., Fisher's linear discriminant, Logistic regression, Naive Bayes classifier, Perceptron, and Support vector machine), quadratic classifiers, k-nearest neighbor, Boosting, Decision trees (e.g., C4.5, Random forests), Bayesian networks, and Hidden Markov models.

[0124] Alternatively or in complement, the artificial intelligence technique may imply the use of one or several of the following elements: artificial neural network, Data clustering, Expectation-maximization algorithm, Self-organizing map, Radial basis function network, Vector Quantization, Generative topographic map, Information bottleneck method, and IBSEAD, rule learning algorithms such as Apriori algorithm, Eclat algorithm and FP-growth algorithm, hierarchical clustering, such as Single-linkage clustering and Conceptual clustering, partitional clustering such as K-means algorithm and Fuzzy clustering.Alternatively or in complement, the artificial intelligence technique uses a reinforcement learning algorithm. Examples of reinforcement learning algorithms include, but are not limited to, temporal difference learning, Q-learning and Learning Automata.

[0125] Alternatively or in complement, the artificial intelligence technique uses a reinforcement learning algorithm. Examples of reinforcement learning algorithms include, but are not limited to, temporal difference learning, Q-learning and Learning Automata.

[0126] Alternatively or in complement, the artificial intelligence technique uses Data Preprocessing.

[0127] More specifically, the model is chosen among a linear model, a non-linear model, an ensemble model and a deep learning model.

[0128] A linear model is a model that uses linear relation(s) between the inputs and the outputs.

[0129] In the present case, the linear model is penalized multinomial regression or linear discriminant analysis

[0130] A non-linear model is a model that uses non-linear relation(s) between the inputs and the outputs.

[0131] As a specific example, it is hereinafter considered molecular classifying functions FMC for kidney allograft rejection.

[0132] More specifically, it is assumed that the obtaining method aims at obtaining molecular classifying functions FMC for antibody-mediated (AMR) and T cell-mediated rejection (TCMR), assessed according to the international Banff 2019 classification.

[0133] As used herein, the term “kidney transplantation” refers to the process of taking a kidney (i.e. “graft”) from one individual and placing it or them into a different individual. The individual who provides for the transplant is called the “donor” and the individual who receives the transplant is called the “recipient”. As used herein, the term “allograft” refers to a tissue or organ that is transplanted from one individual to another within the same species.

[0134] As used herein, the term "antibody-mediated rejection” or “AMR” has its general meaning in the art and refers to the pathological process that is associated with pathogenic donor specific anti-HLA antibodies (DSA). Banff classification recognizes three diagnostic AMR categories: active AMR, chronic active AMR and chronic (inactive) AMR.

[0135] As used herein, the term " T cell-mediated rejection” or “TCMR” has its general meaning in the art and refers to the pathological process characterized by infiltration of the interstitium by T cells and macrophages, intense IFNγ and TGFβ effects, and epithelial deterioration. The three primary sites of acute TC R in transplanted kidneys are indeed the tubular epithelial cells, interstitium, and the vascular endothelial cells. TCMR can be acuteTCMR or chronic acitve TCMR. Banff classification recognizes three gradings: tubulitis (the t score), interstitial inflammation (the i score) and endarteritis (the v score).

[0136] However, such obtaining method may apply to any molecular classifying functions FMC.

[0137] The obtaining method comprises a step of obtaining E70, a step of training E72, a step of using E74, a step of determining E76, a step of selecting E78, a step of generating E80 and a step of choosing E82.

[0138] During the step of obtaining E70, the obtaining module 52 obtains an initial database IDB.

[0139] As an example, the obtaining module 52 simply reads the values of the initial database IDB in a memory, for instance a memory of an external secured server.

[0140] In the present case, the obtaining module 52 reads a cohort C0 from which the obtaining module 52 extracts the initial database IDB and a validating database VDB.

[0141] The cohort CO is a set of subjects for which several parameters are known.

[0142] Such extraction is carried out by splitting the cohort DB0 in two with a splitting ratio. Advantageously, the splitting ratio is comprised between 7:3 and 9:1. The notation X: Y means that for (X+Y) elements present in the cohort DB0, X elements are inserted in the initial database IDB and Y elements are inserted in the validating database VDB.

[0143] Preferably, the splitting ratio is equal to 8:2.

[0144] As schematically represented on figure 4, the initial database IDB comprises data relative to several patients.

[0145] In the specific example, the initial database IDB comprises parameters for real subjects RS.

[0146] As apparent in figure 4, the initial database IDB gathers the data of a first number N1 of real subjects. Such initial database IDB can thus be seen as a cohort of N1 subjects.

[0147] The first number N1 is here an integer superior to 2. The first number N1 is preferably superior to 100, more preferably superior to 300 and advantageously superior to 500.

[0148] In the present case, for each real subject, the initial database IDB comprises at least one comorbidity parameter COMP, at least one clinical parameter CLIP, at least one biological parameter BIOP, at least one histological parameter HISP, at least one gene expression level parameter GELP and the biological state S of the subject.

[0149] By definition, a comorbidity is the presence of one or more additional conditions cooccurring with a primary condition. A comorbidity parameter CO P is a parameter relative to a comorbidity.

[0150] In the present example, the comorbidity parameters COMP are a binary response to several questions. The questions are the following:• whether the donor is living or deceased,

[0151] • if the donor is deceased, whether the cause of the death is due to a circulatory illness, such as a cardiac illness,

[0152] • if the donor is decease, whether the cause of the death is due to a cerebrovascular cause,

[0153] • whether the donor suffers from hypertension,

[0154] • whether the donor suffers from diabetes,

[0155] • whether the donor suffers from proteiurina, and

[0156] • whether the donor suffers from Hepatisis C virus.

[0157] In variant or in complement, the comorbidities parameters COMP comprises any one of the previous parameters chosen among the previous list of items.

[0158] The term “proteinuria” refers to a condition in which excess protein is present in the urine of a subject. In human subjects, proteinuria is often diagnosed by urinalysis. Clinically, proteinuria is expressed as a ratio for urinary protein / creatinine (g / g of creatinine) and said ratio is normally comprised between 0 and 0.3. When such ratio is out of this range, it is considered that the subject suffers from proteinuria.

[0159] In the described example, the clinical parameters CLIP are:

[0160] • the age of the donor,

[0161] • the gender of the donor, and

[0162] • the body mass index of the donor. The body mass index is the donor's weight in kilograms divided by the square of height in meters.

[0163] As an illustration, a biological parameter BIOP is the creatinine rate in milligrams (mg) per deciliter (dL). Normal levels of creatinine in the blood are approximately 0.7 to 1.2 milligram (mg) per deciliter (dL) in adult males and 0.5 to 1.0 milligram per deciliter in adult females.

[0164] By definition, a histological parameter HISP is a data / information, which concerns the study of biological tissues.

[0165] In the present example, the histological parameters HISP are allograft histological pieces of information.

[0166] In this specific example, the histological parameters HISP are assessed according to the Banff Classification that is an international consensus classification for the reporting of biopsies from solid organ transplants. Banff Lesion Scores indeed assess the presence and the degree of histopathological changes in the different compartments of renal transplant biopsies, focusing primarily but not exclusively on the diagnostic features seen in rejection.In the current illustrated case, the histological parameters HISP are four parameters, which are:

[0167] • the value of the percentage of the glomerusclerosis,

[0168] • the stage of arteriosclerosis,

[0169] • the stage of the arteriolar hyalinosis,

[0170] • the stage of interstitial fibrosis, and

[0171] • the stage of tubular atrophy.

[0172] In this case, the stages are the class of the Banff classification of Renal Allograft Pathology predefined classes.

[0173] More specifically, glomerulosclerosis is hardening of the glomeruli in the kidney. It is a general term to describe scarring of the kidneys' tiny blood vessels, the glomeruli, the functional units in the kidney that filter urea from the blood. By definition, the value of the glomerusclerosis is a percentage obtained by dividing the number of sclerosed glomeruli by the number of glomeruli found on the biopsy.

[0174] The arteriosclerosis is evaluated by Banff Lesion Score cv. This score reflects the extent of arterial intimal thickening in the most severely affected artery. The score is evaluated according to four stages which are:

[0175] cv0: no chronic vascular changes;

[0176] cv1: vascular narrowing of up to 25% luminal area by fibrointimal thickening; cv2: vascular narrowing of 26 to 50% luminal area by fibrointimal thickening, and cv3: vascular narrowing of more than 50% luminal area y fibrointimal thickening. The arteriolar hyalinosis is evaluated by Banff Lesion Score ah. This score evaluates the extent of arteriolar hyalinosis. The score is evaluated according to four stages which are:

[0177] stage 0: no PAS (PAS)-positive hyaline arteriolar thickening;

[0178] stage 1: mild to moderate PAS-positive hyaline thickening in at least 1 arteriole; stage 2: moderate to severe PAS-positive hyaline thickening in more than 1 arteriole, and

[0179] stage 3: severe PAS-positive hyaline thickening in many arterioles.

[0180] The interstitial fibrosis is evaluated by Banff Lesion Score ci. This score evaluates the extent of cortical fibrosis. The score is evaluated according to four stages which are:

[0181] ci0: interstitial fibrosis in up to 5% of cortical area;

[0182] ci 1: interstitial fibrosis in 6 to 25% of cortical area (mild interstitial fibrosis); ci2: interstitial fibrosis in 26 to 50% of cortical area (moderate interstitial fibrosis), andci3: interstitial fibrosis in more than 50% of cortical area (severe interstitial fibrosis). The interstitial fibrosis is evaluated by Banff Lesion Score ct. This score evaluates the extent of cortical tubular atrophy which is usually tightly associated with the areas affected with interstitial fibrosis. The score is evaluated according to four stages which are:

[0183] ct0: no tubular atrophy;

[0184] ct1: tubular atrophy involving up to 25% of the area of cortical tubules;

[0185] ct2: tubular atrophy involving 26 to 50% of the area of cortical tubules, and ct3: tubular atrophy involving more than 50% of the area of cortical tubules.

[0186] By definition, a gene expression level parameter GELP is a parameter representative of the expression level of said gene for said subject.

[0187] As used herein, the term “gene” has its general meaning in the art and refers to a sequence of nucleotides in DNA or RNA that encodes the synthesis of a gene product, either RNA or protein, which performs a specific function in the organism. Genes are the fundamental units of heredity and are transferred from parent to offspring during reproduction. They are responsible for the inherited characteristics and functions of living organisms, and their expression levels can be quantitatively assessed to provide insights into various biological processes and disease states. In the present specification, the name of each of the genes of interest refers to the internationally recognised name of the corresponding gene, as found in internationally recognised gene sequences and protein sequences databases, in particular in the database from the HUGO Gene Nomenclature Committee, that is available notably at the following Internet address http: / / www.gene.ucl.ac.uk / nomenclature / index.html. In the present specification, the name of each of the various biological markers of interest may also refer to the internationally recognised name of the corresponding gene, as found in the internationally recognised gene sequences and protein sequences databases ENTRE ID, Genbank, TrEMBL or ENSEMBL. Through these internationally recognised sequence databases, the nucleic acid sequences corresponding to each of the gene of interest described herein may be retrieved by the one skilled in the art.

[0188] As used herein, the term “expression level” refers to the quantity of mRNA produced from a gene, which serves as an indicator of gene activity within a sample.

[0189] As used herein, the term “sample” refers to any biological sample obtained for evaluation in vitro. In the context of transplantation, this typically includes tissue samples. The term “tissue sample” encompasses sections of tissues such as biopsy or autopsy samples and frozen sections taken for histological analysis. According to the present invention the sample is obtained from the transplanted organ. Thus, in some embodiments,the tissue sample may result from a biopsy performed on the transplanted organ to assess its health and function.

[0190] Methods for determining the expression level of a gene product such as nucleic acid (e.g. RNA) are also well known in the art. Conventional methods typically involve polymerase chain reaction (PCR). For instance, U. S. Pat. Nos. 4,683,202, 4,683,195, 4,800,159, and 4,965,188 disclose conventional PCR techniques. PCR typically employs two oligonucleotide primers that bind to a selected target nucleic acid sequence. Primers useful in the present invention include oligonucleotides capable of acting as a point of initiation of nucleic acid synthesis within the target nucleic acid sequence. A primer can be purified from a restriction digest by conventional methods, or it can be produced synthetically. PCR involves use of a thermostable polymerase. The term “thermostable polymerase” refers to a polymerase enzyme that is heat stable, i.e., the enzyme catalyzes the formation of primer extension products complementary to a template and does not irreversibly denature when subjected to the elevated temperatures for the time necessary to effect denaturation of double-stranded template nucleic acids. Thermostable polymerases have been isolated from Thermus fiavus, T. ruber, T. thermophilus, T. aquaticus, T. lacteus, T. rubens, Bacillus stearothermophilus, and Methanothermus fervidus. Nonetheless, polymerases that are not thermostable also can be employed in PCR assays provided the enzyme is replenished. Typically, the polymerase is a Taq polymerase (i.e. Thermus aquaticus polymerase). Quantitative PCR is typically carried out in a thermal cycler with the capacity to illuminate each sample with a beam of light of a specified wavelength and detect the fluorescence emitted by the excited fluorophore. The thermal cycler is also able to rapidly heat and chill samples, thereby taking advantage of the physicochemical properties of the nucleic acids and thermal polymerase. In order to detect and measure the amount of amplicon (i.e. amplified target nucleic acid sequence) in the sample, a measurable signal has to be generated, which is proportional to the amount of amplified product. All current detection systems use fluorescent technologies. Some of them are non-specific techniques, and consequently only allow the detection of one target at a time. Alternatively, specific detection chemistries can distinguish between non- specific amplification and target amplification. These specific techniques can be used to multiplex the assay, i.e. detecting several different targets in the same assay. For example, SYBR® Green I probes, High Resolution Melting probes, TaqMan® probes, LNA® probes and Molecular Beacon probes can be suitable. TaqMan® probes are the most widely used type of probes. They were developed by Roche (Basel, Switzerland) and ABI (Foster City, USA) from an assay that originally used a radio-labelled probe (Holland et al. 1991), which consisted of a single-stranded probe sequence that was complementary to one of thestrands of the amplicon. A fluorophore is attached to the 5’ end of the probe and a quencher to the 3’ end. The fluorophore is excited by the machine and passes its energy, via FRET (Fluorescence Resonance Energy Transfer) to the quencher. Traditionally, the FRET pair has been conjugated to FAM as the fluorophore and TAMRA as the quencher. In a well-designed probe, FAM does not fluoresce as it passes its energy onto TAMRA. As TAMRA fluorescence is detected at a different wavelength to FAM, the background level of FAM is low. The probe binds to the amplicon during each annealing step of the PCR. When the Taq polymerase extends from the primer which is bound to the amplicon, it displaces the 5’ end of the probe, which is then degraded by the 5’-3’ exonuclease activity of the Taq polymerase. Cleavage continues until the remaining probe melts off the amplicon. This process releases the fluorophore and quencher into solution, spatially separating them (compared to when they were held together by the probe). This leads to an irreversible increase in fluorescence from the FAM and a decrease in the TAMRA.

[0191] In some embodiments, the expression level is determined by RNA-seq. As used, the term " RNA-Seq" or "transcriptome sequencing" refers to sequencing performed on RNA (or cDNA) instead of DNA, where typically, the primary goal is to measure expression levels, detect fusion transcripts, alternative splicing, and other genomic alterations that can be better assessed from RNA. RNA-Seq typically includes whole transcriptome sequencing. As used herein, the term “whole transcriptome sequencing” refers to the use of high throughput sequencing technologies to sequence the entire transcriptome in order to get information about a sample's RNA content. Whole transcriptome sequencing can be done with a variety of platforms for example, the Genome Analyzer (Illumina, Inc., San Diego, Calif.) and the SOLiD™ Sequencing System (Life Technologies, Carlsbad, Calif.), However, any platform useful for whole transcriptome sequencing may be used. Typically, the RNA is extracted, and ribosomal RNA may be deleted as described in U. S. Pub, No.

[0192] 2011 / 0111409. cDNA sequencing libraries may be prepared that are directional and single or paired-end using commercially available kits such as the ScriptSeq™ M mRNA-Seq Library Preparation Kit (Epicenter Biotechnologies, Madison, Wis.). The libraries may also be barcoded for multiplex sequencing using commercially available barcode primers such as the RNA-Seq Barcode Primers from Epicenter Biotechnologies (Madison, Wis.). PCR is then carried out to generate the second strand of cDNA to incorporate the barcodes and to amplify the libraries. After the libraries are quantified, the sequencing libraries may be sequenced. Nucleic acid sequencing technologies are suitable methods for expression analysis. The principle underlying these methods is that the number of times a (DNA sequence is detected in a sample is directly related to the relative RNA levels corresponding to that sequence. These methods are sometimes referred to by the term Digital GeneExpression (DOE) to reflect the discrete numeric property of the resulting data. Early methods applying this principle were Serial Analysis of Gene Expression (SAGE) and Massively Parallel Signature Sequencing (MPSS). See, e.g., S. Brenner, et al., Nature Biotechnology 18(6):630-634 (2000). Typically RNA-seq uses Next Generation Sequencing or NGS. As used herein, the term " Next Generation Sequencing" (NGS) refers to a relatively new sequencing technique as compared to the traditional Sanger sequencing technique. For review, see Shendure et al., Nature Biotech., 26(10): 1135-45 (2008), which is hereby incorporated by reference into this disclosure. For purpose of this disclosure, NGS may include cyclic array sequencing, microelectrophoretic sequencing, sequencing by hybridization, among others. By way of example, in a typical NGS using cyclic-array methods, genomic DNA or cDNA library is first prepared, and common adaptors may then be ligated to the fragmented genomic DNA or cDNA. Different protocols may be used to generate jumping libraries of mate-paired tags with controllable distance distribution. An array of millions of spatially immobilized PCR colonies or "polonies" is generated with each polonies consisting of many copies of a single shotgun library fragment. Because the polonies are tethered to a planar array, a single microliter- scale reagent volume can be applied to manipulate the array features in parallel, for example, for primer hybridization or for enzymatic extension reactions. Imaging-based detection of fluorescent labels incorporated with each extension may be used to acquire sequencing data on all features in parallel. Successive iterations of enzymatic interrogation and imaging may also be used to build up a contiguous sequencing read for each array feature.

[0193] In some embodiments, the nCounter® Analysis system is used to detect intrinsic gene expression. The basis of the nCounter® Analysis system is the unique code assigned to each nucleic acid target to be assayed (International Patent Application Publication No. WO 08 / 124847, U. S. Patent No. 8,415,102 and Geiss et al. Nature Biotechnology. 2008. 26(3): 317-325; the contents of which are each incorporated herein by reference in their entireties). The code is composed of an ordered series of colored fluorescent spots which create a unique barcode for each target to be assayed. A pair of probes is designed for each DNA or RNA target, a biotinylated capture probe and a reporter probe carrying the fluorescent barcode. This system is also referred to, herein, as the nanoreporter code system. Specific reporter and capture probes are synthesized for each target. The reporter probe can comprise at a least a first label attachment region to which are attached one or more label monomers that emit light constituting a first signal; at least a second label attachment region, which is non-over-lapping with the first label attachment region, to which are attached one or more label monomers that emit light constituting a second signal; and a first target- specific sequence. Preferably, each sequence specific reporter probe comprisesa target specific sequence capable of hybridizing to no more than one gene and optionally comprises at least three, or at least four label attachment regions, said attachment regions comprising one or more label monomers that emit light, constituting at least a third signal, or at least a fourth signal, respectively. The capture probe can comprise a second targetspecific sequence; and a first affinity tag. In some embodiments, the capture probe can also comprise one or more label attachment regions. Preferably, the first target- specific sequence of the reporter probe and the second target- specific sequence of the capture probe hybridize to different regions of the same gene to be detected. Reporter and capture probes are all pooled into a single hybridization mixture, the "probe library". The relative abundance of each target is measured in a single multiplexed hybridization reaction. The method comprises contacting the tissue sample with a probe library, such that the presence of the target in the sample creates a probe pair - target complex. The complex is then purified. More specifically, the sample is combined with the probe library, and hybridization occurs in solution. After hybridization, the tripartite hybridized complexes (probe pairs and target) are purified in a two-step procedure using magnetic beads linked to oligonucleotides complementary to universal sequences present on the capture and reporter probes. This dual purification process allows the hybridization reaction to be driven to completion with a large excess of target-specific probes, as they are ultimately removed, and, thus, do not interfere with binding and imaging of the sample. All post hybridization steps are handled robotically on a custom liquid-handling robot (Prep Station, NanoString Technologies). Purified reactions are typically deposited by the Prep Station into individual flow cells of a sample cartridge, bound to a streptavidin-coated surface via the capture probe, electrophoresed to elongate the reporter probes, and immobilized. After processing, the sample cartridge is transferred to a fully automated imaging and data collection device (Digital Analyzer, NanoString Technologies). The level of a target is measured by imaging each sample and counting the number of times the code for that target is detected. For each sample, typically 600 fields-of-view (FOV) are imaged (1376 X 1024 pixels) representing approximately 10 mm2 of the binding surface. Typical imaging density is 100- 1200 counted reporters per field of view depending on the degree of multiplexing, the amount of sample input, and overall target abundance. Data is output in simple spreadsheet format listing the number of counts per target, per sample. This system can be used along with nanoreporters. Additional disclosure regarding nanoreporters can be found in International Publication No. WO 07 / 076129 and W007 / 076132, and US Patent Publication No.

[0194] 2010 / 0015607 and 2010 / 0261026, the contents of which are incorporated herein in their entireties. Further, the term nucleic acid probes and nanoreporters can include the rationally designed (e.g. synthetic sequences) described in International Publication No. WO2010 / 019826 and US Patent Publication No.2010 / 0047924, incorporated herein by reference in its entirety.

[0195] Typically, the expression level of a gene may be expressed as absolute level or normalized level. For instance, levels are normalized by correcting the absolute level of a gene by comparing its expression to the expression of a gene that is not a relevant for determining the risk. This normalization allows the comparison of the level in one sample, e.g., a subject sample, to another sample, or between samples from different sources.

[0196] Each of these levels are example of gene expression level parameters GLEP, which may be part of the initial database IDB.

[0197] The number of gene involved at this stage of the method is quite high, notably superior to 100, preferably superior to 200, advantageously superior to 500.

[0198] An element of the initial database IDB may comprise other data, such as identification data. For instance, the element may provide, for each subject, with anonymized identifier or the institution from which the data originates.

[0199] During the step of training E72, the training module 54 trains a conditional tabular generative adversarial network on the initial database IDB to output elements fulfilling at least one similarity condition with the elements of the initial database IDB.

[0200] Conditional tabular generative adversarial network is often named by the abbreviation CTGAN, which will be used hereinafter.

[0201] A conditional tabular generative adversarial network is a deep learning-based model specifically designed for generating synthetic tabular data.

[0202] More specifically, a conditional tabular generative adversarial network is a specific generative Adversarial Network (more often named GAN). Such network uses a unique structure that includes two neural network architectures: the ‘generator’ and the ‘discriminator’. They show adversarial behavior to sample and model from the training dataset. Ultimately, the modeled GAN generates a synthetic dataset similar to the original dataset.

[0203] More precisely, the generator creates synthetic data samples, while the discriminator evaluates them against real data samples to distinguish between the two. In the context of tabular data, the generator is trained to produce synthetic data that mimics the statistical properties and distribution of the real tabular dataset. It learns to generate rows of data with realistic feature values. The discriminator's role is to differentiate between real and synthetic data samples. It provides feedback to the generator, helping it improve the quality of the generated data over time.

[0204] A similarity condition is that a value of a similarity metrics be inferior to a predefined threshold.A first kind of similarity metric is a measure of the difference of the distribution between the numerical values of two databases to be compared. The KSComplement is a specific example of such first kind.

[0205] A second kind of similarity metric is the Total Variation Distance between the real and synthetic variables. This kind of similarity metric is notably appropriate for categorical value. The TVComplement is a specific example of such second kind.

[0206] Such kinds of similarity metrics rely on univariate comparison.

[0207] Similarity metrics enabling bivariate comparison may also be considered.

[0208] For such comparison, a third kind of similarity metric is a measure a correlation coefficient on the real and synthetic data. The Correlationsimilarity is a specific example of such third kind.

[0209] A fourth kind of similarity metric is the calculation of a similarity of a pair of categorical columns between the real and synthetic datasets. The Contingencysimilarity is a specific example of such fourth kind.

[0210] This results in a trained conditional tabular generative adversarial network During the step of using E74, the using module 56 uses the trained CTGAN to obtain synthetic elements.

[0211] These synthetic elements each correspond to generated subjects, in other words to virtual subjects, for which data similar to the data present for the real subject of the initial database IDB are present.

[0212] As apparent in figure 4, the resulting generated database GDB gathers the data of a second number N2 of virtual or generated subjects GS. Such initial database GDB can thus be seen as a cohort of N2 subjects.

[0213] The second number N2 is here an integer superior to 2.

[0214] The Applicant has found that a ratio between the second number N2 and the first number N1 is comprised between 0,1 % and 100%, preferably between 0,25% and 0,75% and advantageously equal to 50%.

[0215] As can be seen in figure 4, an augmented database ADB can be defined as the set of the initial database IDB and the synthetic elements, namely, the set of RS1, RS2, …, RSN1, GS1, GS2, … and GSN2.

[0216] During the step of determining E76, the determining module 58 determines gene expression level parameters GLEP of the augmented database ADB impacting a prediction of the biological state S of the subject by a molecular classifying function FMCE76.

[0217] The molecular classifying function FMCE76used in the step of determining E76 is obtained by an hyperparameter optimization carried out on a part of the initial database IDB, named determining database DDB.According to the present example, for such determination, the determining module 58 applies a Boruta algorithm on the augmented database ADB.

[0218] In short, Boruta algorithm is based on the same idea which forms the foundation of the random forest classifier, namely, that by adding randomness to the system and collecting results from the ensemble of randomized samples one can reduce the misleading impact of random fluctuations and correlations. Here, this extra randomness shall provide us with a clearer view of which attributes are really important.

[0219] More specifically, Boruta algorithm is a wrapper built around the random forest classification algorithm implemented in the R package randomForest.

[0220] This implies that the molecular classifying function FMCE76considered for this step of determining E76 is a random forest one.

[0221] As used herein, the term “random forest” refers to an ensemble learning method used for classification and regression tasks. It operates by constructing a multitude of decision trees during training time and outputting the mode of the classes (classification) or mean prediction (regression) of the individual trees. Random forests correct for decision trees' habit of overfitting to their training set by introducing randomness into the tree-building process. Each tree is trained on a random subset of the training data, and features are randomly selected at each split, ensuring that the resulting model is both robust and generalizable.

[0222] The importance measure of a gene expression level parameters GLEP is obtained as the loss of accuracy of classification caused by the random permutation of attribute values between objects. It is computed separately for all trees in the forest which use a given gene expression level parameters GLEP for classification. Then the average and standard deviation of the accuracy loss are computed.

[0223] Alternatively, the Z score computed by dividing the average loss by its standard deviation can be used as the importance measure.

[0224] The determining database DDB is enriched with gene expression level parameters GLEP that are random by design. For each gene expression level parameters GLEP, the determining module 58 creates a corresponding ‘shadow’ parameters, whose values are obtained by shuffling values of the original gene expression level parameter across objects.

[0225] The determining module 58 then performs a classification using all gene expression level parameters GLEP of this extended system and computes the importance of all gene expression level parameters GLEP.

[0226] The importance of a shadow parameter can be nonzero only due to random fluctuations.Thus the set of importances of shadow parameters is used as a reference for deciding which attributes are truly important.

[0227] The importance measure itself varies due to stochasticity of the random forest classifier. Additionally it is sensitive to the presence of non important parameters in the determining database DDB (also the shadow ones). Moreover it is dependent on the particular realization of shadow parameter. Therefore the re-shuffling procedure is shortto obtain statistically valid results.

[0228] According to the present example, applying the Boruta algorithm comprises the following operations:

[0229] - an operation of extending the determining database DDB by adding copies of all variables (the determining database DDB is extended by at least 5 shadow parameters, even if the number of gene expression level parameters GLEP in the original set is lower than 5).

[0230] - an operation of shuffling the added gene expression level parameters GLEP to remove their correlations with the response.

[0231] - an operation of running a random forest classifier (corresponding to a molecular classifying function FMCE76) a on the extended determining database DDB and gather the Z scores computed.

[0232] - an operation of finding the maximum Z score among shadow parameters (MZSP), and then assigning a hit to every gene expression level parameter GLEP that scored better than MZSP.

[0233] - an operation of performing a two-sided test of equality with the MZSP, for each parameter with undetermined importance.

[0234] - an operation of deeming the parameters which have importance significantly lower than MZSP as ‘unimportant’ and permanently removing them from the determining database DDB.

[0235] - an operation of deeming the attributes which have importance significantly higher than MZSP as ‘important’.

[0236] - an operation of removing all shadow parameters.

[0237] - an operation of repeating the procedure until the importance is assigned for all the attributes, or the Boruta algorithm has reached the previously set limit of the random forest runs.

[0238] In practice this algorithm is preceded with three start-up rounds, with less restrictive importance criteria. The startup rounds are introduced to cope with high fluctuations of Z scores when the number of attributes is large at the beginning of the procedure. During these initial rounds, parameters are compared respectively to the fifth, third and secondbest shadow parameter; the test for rejection is performed only at the end of each initial round, while the test for confirmation is not performed at all.

[0239] This enables to obtain determined gene expression level parameters GLEPDET.

[0240] The number of determined gene expression level parameters GLEPDET is comprised between 10 and 50, preferably comprised between 10 and 30.

[0241] This amounts to a drastic reduction compared to the original gene expression level parameters GLEP, with a reduction factor of the order of 10.

[0242] During the step of selecting E78, the selecting module 60 selects gene expression level parameters GLEP among the determined gene expression level parameters GLEPDET resulting from the step of determining E76.

[0243] The output of the step of selecting E78 is a set of selected gene expression level parameters GLEPSEL.

[0244] For this, the step of selecting E78 comprises iteratively for, several determined gene expression level parameters GLEPDET carrying out two operations, namely an operation of analyzing and an operation of adding / discarding.

[0245] During the operation of analyzing, the selecting module 60 analyses the impact of adding said determined gene expression level parameters GLEPDET in a learning process of a molecular classifying function FMCETS adapted to take as at least one already selected gene expression level parameter GLEPSEL and said determined gene expression level parameters GLEPDET and to provide as output the biological state S of the subject.

[0246] During the operation of adding / discarding, the selecting module 60 adds or discard said gene expression level parameter GLEP depending from the fulfillment of not of a selection criterion.

[0247] In other words, the selecting module 60 performs forward stepwise feature selection method to increment the number of features from one to the all determined gene expression level parameters GLEPDET.

[0248] According to a preferred embodiment, the molecular classifying function FMCETS used is also here a random forest algorithm.

[0249] The selecting module 60 generates learning curves to observe the changes of predictive performance of the random forest classifying functions FMCE78 by adding determined gene expression level parameters GLEPDET.

[0250] The selecting module 60 then optimizes each of said generated learning curves according to an optimization criterion by using a grid search on fixed one-fold cross-validation.

[0251] The one-fold cross-validation is a part of the initial database IDB and is used for hyperparameter optimization in order to tune the molecular classifying functions FMCE76(and not the augmented database ADB). Said part is advantageously the same as the part used in the step of determining E76, namely the determining database DDB.

[0252] Said optimization criterion is, for instance, to optimize the PRAUC value and / or the ROCAUC value (which are defined in reference to the step of choosing E82 with other example of performance metrics that may be used also in this context). The value of the optimization criterion is determined based on the validating database VDB.

[0253] Based on the optimized learning curves, the selecting module 60 is adapted to select the appropriate optimized learning curve, which corresponds to the optimized set of selected gene expression level parameters GLEPSEL, which should be used by a molecular classifying function FMC.

[0254] During the step of generating E80, the generating module 62 generates molecular classifying functions FMCCANDtaking as input the set of selected gene expression level parameters GLEPSEL and providing as output the biological state S of the subject.

[0255] For this, the generating module 62 carries out a training process on the augmented database ADB and preferably with the same performance metric, such as the value of PRAUC.

[0256] The assessment of the performance metric is preferably also carried on the same determining database DDB.

[0257] This enables to obtain molecular classifying function candidates FMCCAND.

[0258] During the step of choosing E82, the choosing module 64 chooses at least one molecular classifying function FMC among the molecular classifying function candidates FMCCANDbased on a performance criterion.

[0259] According to an embodiment, the choosing module 64 selects each molecular classifying function FMCCANDfulfilling the performance criterion.

[0260] In another embodiment, the choosing module 64 selects only the molecular classifying function FMCCAND, which better fulfills the performance criterion.

[0261] The performance criterion is, for instance, evaluated by determining the value of a performance metric.

[0262] In such example, according to the case, the performance criterion is that the value of the performance metric be inferior to a threshold (several chosen molecular classifying functions FMCCAND) or the value of the performance metric itself (one molecular classifying function FMCCANDchosen only).

[0263] The performance metric is chosen in the list consisting of the positive predictive value, the negative predictive value, the sensitivity, the specificity, the accuracy, the balanced accuracy, the F1 score, the ROCAUC, the PRAUC and the Brier score.Hereinafter, TP refers to a true positive prediction, FP to a false positive prediction, FN to a false negative prediction and TN to a true negative prediction.

[0264] The positive predictive value, also named by its abbreviation PPV, is defined as the ratio of the number true positive predictions to the number of positive predictions. Precision is another name of this metric.

[0265] In other words, this metric PPV can be written mathematically as:

[0266] #TP

[0267] PPV = - #TP + #FP

[0268] wherein #X designates the number of X.

[0269] Similarly, the negative predictive value, also named by its abbreviation NPV, is defined as the ratio of the number of true negative predictions to the number of negative predictions. This can be expressed as follows:

[0270] #TN

[0271] NPV = - #TN + #FN

[0272] The sensitivity is defined as the ratio of the number of true positive predictions to the sum of the number of true positive predictions and the number of false negative predictions. The sensitivity is sometimes named true positive rate (TPR) or recall and can be expressed as follows:

[0273] #TP

[0274] TPR =

[0275] #TP + #FN

[0276] The specificity is the ratio of the number of true negative predictions to the sum of the number of true negative predictions and the number of false positive predictions. This corresponds to the following relation:

[0277] #TN

[0278] Specificity = #TN / (#TN + #FP)

[0279]

[0280] The accuracy quantifies the proportion of correct predictions, either negative or positive. This leads to the following formula:

[0281] #TN + #TP

[0282] Accuracy = (#TN + #TP) / (#TN + #FN + #TP + #FP)

[0283]

[0284] The balanced accuracy is the arithmetic average of sensitivity and specificity. This can be expressed as follows:

[0285] 1 / #TP #TN \

[0286] Balanced Accuracy = - — — - — — ■ + — — — - — — -

[0287]

[0288] 72 \#TP + #FN #TN + #FP / The F1 score is defined as the harmonic mean of precision and recall. Mathematically, this implies that:

[0289] 2 #TP

[0290] F1 score = 2 / (1 / precision + 1 / recall) = #TP / (#TP + (1 / 2)(#FN + #FP))

[0291] precision recall 2yROCAUC stands for the receiver operating characteristic area under the curve. The ROCAUC is the area of the ROC (receiver operating characteristic) space that lies below the ROC curve, the ROC curve being the plot of the true positive rate (TPR) against the false positive rate (FPR) at each threshold setting of the molecular classifying function FMCCANDto be evaluated.

[0292] The false positive rate FPR is the ratio between the number of negative events wrongly categorized as positive (false positives) and the total number of actual negative events (regardless of classification), meaning that:

[0293] #FP

[0294] FPR = - #FP + #TN

[0295] PRAUC designates the precision recall area under the curve. The PRAUC is thus the area of the space that lies below the precision recall curve (the curve of recall in function of the precision).

[0296] The Brier score measures the mean squared difference between the predicted probability and the real prediction.

[0297] For a binary prediction, the Brier score BS can be expressed as:

[0298] n

[0299] BS = (1 / n) Σ(pi- oi)2

[0300] i=1

[0301]

[0302] 1=1

[0303] Where:

[0304] • n is the total number of prediction,

[0305] • pi is the prediction probability, and

[0306] • oi is equal to 1 if the event occurred and 0 if not.

[0307] Alternatively, the choosing module 64 uses simultaneously several or all of the previous performance metrics.

[0308] For this, the choosing module 64 calculates the value of said performance metrics for each molecular classifying function candidate FMCCAND.

[0309] Said assessment of the performance criterion is carried out on the validating database VDB.

[0310] Then, according to an embodiment, the choosing module 64 compares said calculated values to a respective threshold (defined previously for each performance metrics) and selects the molecular classifying function candidates FMCCANDfor which all the comparisons are positive (the molecular classifying function candidate FMCCANDfulfills all the performance criteria).

[0311] Alternatively, the choosing module 64 aggregates the values of the different performance metrics and then compares the aggregated value to a predefined threshold toobtain one or several chosen molecular classifying functions FMC among the molecular classifying function candidates FMCCAND.

[0312] In another embodiment, the choosing module 64 selects the molecular classifying functions FMC for which the aggregated value is the highest.

[0313] In the present case, it has been found that using the combination of the ROCAUC, the PRAUC and the Brier score as performance metrics enables to an appropriate choice of the molecular classifying functions FMC

[0314] This leads to obtaining at least one molecular classifying function FMC intended to be used to predict a biological state of a subject, which exhibits reliable performance.

[0315] Such new way of obtaining a relevant molecular classifying function FMC is advantageously used in a method for predicting carried out by a system for predicting 100 schematized on figure 5.

[0316] Such system for predicting 100 comprising a providing module 102 and an applying module 104.

[0317] In the example of figure 5, the system for predicting 100 comprises an information processing unit formed for example by a memory and a processor associated with the memory. In that sense, the system for predicting 100 is an electronic system.

[0318] The providing module 102 and the applying module 104 are each produced in the form of software, or a software brick, executable by the processor.

[0319] The memory of the system for predicting 100 is then able to store a software of obtaining, a software of training, a software of using, a software of determining, a software of selecting, a software of generating and a software of choosing.

[0320] The processor is then able to execute each of the software among the software of training, the software of using, the software of determining, the software of selecting, the software of generating and the software of choosing

[0321] In a variant not shown, the providing module 102 and the applying module 104 are each produced in the form of a programmable logic component, such as an FPGA (abbreviation for Field Programmable Gate Array), or an integrated circuit, such as an ASIC (abbreviation for Application Specific Integrated Circuit).

[0322] When the system for predicting 100 is produced in the form of one or more software programs, i.e. in the form of a computer program, also called a computer program product, it is also capable of being recorded on a medium, not shown, readable by a computer. The computer-readable medium is, for example, a medium capable of storing electronic instructions and of being coupled to a bus of a computer system. For example, the readable medium is an optical disk, a magneto-optical disk, a ROM memory, a RAM memory, anytype of non-volatile memory (for example FLASH or NVRAM) or a magnetic card. A computer program comprising software instructions is then stored on the readable medium.

[0323] The actions that the providing module 102 and the applying module 104 are configured to implement are now exemplified in reference to figure 6, which is a flowchart illustrating an implementation of a method for predicting a biological state of a subject (predicting method hereinafter).

[0324] The predicting method comprises a step of providing E110, a step of applying E112 and outputs the state of a subject as apparent from the box 114.

[0325] During the step of providing E110, the providing module 102 provides parameters representative of the expression level of genes of the subject.

[0326] In this example, the providing module 102 is receiving the parameters, for instance transmitted from an external server.

[0327] Such server may be in relation with a laboratory in which the expression level of genes of the subject are obtained.

[0328] According to another embodiment, a user enters data in an input device, which is part of the providing module 102.

[0329] Alternatively, the providing module 102 is further adapted to measure the expression level of genes, so that the step of providing E110 is achieved by measuring the parameters representative of the expression level of genes of the subject.

[0330] During the step of applying E112, the applying module 104 applies a molecular classifying function FMC obtained by the obtaining method on the obtained parameters, to output the biological state of the subject.

[0331] This enables to obtain an accurate prediction of the biological state of the subject. It can be observed that the predicting method can be construed as an inference phase of a method, which further comprises a training phase corresponding to the obtaining method.

[0332] With such construction, the obtaining method is generally carried off-line and predicting method is carried out by a local system, typically a diagnostic tool of an hospital.

[0333] Such predicted biological state of the subject may be used specific applications. Some examples are provided hereinafter for many applications linked to diseases. According to the context and in accordance with the previously mentioned examples, such disease can, for instance, be a kidney disease or a heart disease.

[0334] Other examples of diseases are acute cellular rejection, antibody mediated rejection, recurrence of the original disease (amyloidosis, diabetes notably) and poliomavirus nephropathy.Graft loss, graft rejection, graft versus host disease, stenosis, thrombosis, acute tubulonephritis, chronic transplant nephropathy, kidney failure, atherosclerosis, arterial hypertension, coronary artery disease are other examples of such kind of diseases.

[0335] In each application, there exists a link between the disease of the application in which the predicting method is used and the organ to which the predicting method is applied. In other words, such disease is a disease or a disorder related to this organ.

[0336] It can thus be considered to use the predicting method in a method for predicting that a subject is at risk of suffering from a disease.

[0337] The term “risk” relates to the probability that an event will occur over a specific time period, and can mean a subject's “absolute” risk or “relative” risk. Absolute risk can be measured with reference to either actual observation post-measurement for the relevant time cohort, or with reference to index values developed from statistically valid historical cohorts that have been followed for the relevant time period. Relative risk refers to the ratio of absolute risks of a subject compared either to the absolute risks of low risk cohorts or an average population risk, which can vary by how clinical risk factors are assessed. Odds ratios, the proportion of positive events to negative events for a given test result, are also commonly used (odds are according to the formula p / (1 — p) where p is the probability of event and (1 — p) is the probability of no event).

[0338] The method for predicting comprises at least the steps of carrying out the steps of the predicting method on the subject, to obtain a predicted biological state, and predicting that the subject is at risk of suffering from the disease based on the predicted biological state.

[0339] Alternatively, it can be considered a method for diagnosing a disease wherein the method for diagnosing comprises at least the steps of carrying out the steps of the predicting method, to obtain a predicted biological state, and diagnosing the disease based on the predicted biological state.

[0340] In both embodiments, the predicted biological state is for instance the state of the graft or the state of a vein of the heart.

[0341] The predicting method can also be advantageously used in a method for identifying a therapeutic target for preventing and / or treating a disease, the method comprising at least the steps of carrying out the steps of the predicting method for a first subject, to obtain a first predicted biological state, the first subject being a subject suffering from the disease, carrying out the steps of the predicting method for a second subject, to obtain a second predicted biological state, the second subject being a subject not suffering from the disease, and selecting a therapeutic target based on the comparison of the first and second predicted biological states.Alternatively, it can be considered a method for identifying a biomarker for a disease, the biomarker being a diagnosis biomarker of the disease, a susceptibility biomarker of the disease, a prognostic biomarker of the disease or a predictive biomarker in response to the treatment of the disease, the method comprising at least the steps of carrying out the steps of the predicting method for a first subject, to obtain a first predicted biological state, the first subject being a subject suffering from the disease, carrying out the steps of the predicting method for a second subject, to obtain a second predicted biological state, the second subject being a subject not suffering from disease, and selecting a biomarker target based on the comparison of the first and second predicted biological states.

[0342] The predicting method can also be advantageously used in a method for screening a compound useful as a medicament, the compound having an effect on a known therapeutical target for preventing and / or treating a disease, the method comprising at least the steps of carrying out the steps of the predicting method for a first subject, to obtain a first predicted biological state, the first subject being a subject suffering from the disease and having received the compound, carrying out the steps of the predicting method for a second subject, to obtain a second predicted biological state, the second subject being a subject suffering from the disease and not having received the compound, and selecting a biomarker target based on the comparison of the first and second predicted biological states.

[0343] The predicting method is also advantageous in a method for monitoring patients enrolled in a clinical trial to provide a quantitative measure for the therapeutic efficacy of the therapy, which is subject to the clinical trial by carrying out the steps of the predicting method on said patients.

[0344] More generally, the predicting method can advantageously be used in any context where the biological state is useful and, even more in the case where such biological state can only be obtained in an invasive way.

[0345] In addition, the person skilled in the art can consider any combination of the features of the previously mentioned embodiment of a method to obtain new embodiments when the features are technically compatible. This applies both to the obtaining method and the predicting method.

[0346] EXPERIMENTAL SECTION

[0347] The following examples are to be considered illustrative and not limiting on the scope of the present disclosure described above. These examples illustrate the results obtained for AMR and TCMR.

[0348] These results are supported by the following figures:Figure 7. Development of the antibody-mediated rejection classifiers on augmented transcriptomic data using generative model in kidney transplant patients.

[0349] Schematic diagram illustrating the development of generative model, generating synthetic data, augmenting the datasets, developing machine learning classifiers, and assessing them in the external validation cohort processes to observe the impact of synthetic data augmentation and to construct the molecular classifiers in kidney transplant patients.

[0350] Figure 8. Flow of the importance of the top 30 genes from the development set to synthetically augmented GanAug200 set. Using the generative adversarial networks (GAN), the inventors generated eight synthetic cohorts resembling the original development set. They were merged to the development set to augment. The inventors performed Boruta algorithm, Monte Carlo method which uses multiple random forests and shadow features. Genes were included as features to predict AMR. The inventors performed 2000 iterations of random forests. Among the selected features considered predictive by the Boruta algorithm, they calculated median importance to explore the feature importance changes by datasets. The slope graph depicts the flow of the top 30 gene features’ importance. The median importances of genes from the development set to incrementally augmented sets were represented from the left to the right side. Lines represent genes considered important from the development set or which appeared to be important in the augmented sets.

[0351] Figure 9. Learning curves of the machine learning models. Using the generative adversarial networks (GAN), the inventors generated eight synthetic cohorts resembling the original development set. They were merged to the development set to augment. After the Boruta feature importance analysis, the inventors incrementally increased the number of selected features to observe the performance change for all datasets. The performance was measured in precision recall area under the curve (PRAUC) and receiver operating characteristic area under the curve (ROCAUC) in the split test set.

[0352] Figure 10. Calibration plots on the external validation cohort. This figure details the performance of final molecular classifiers assessed on the external validation cohort. The calibration performance is demonstrated as plots in this figure (observed probabilities are in ordinate).

[0353] Figure 11. Flow of the importance of the top 30 genes from the development set to synthetically augmented GanAug200 set. Using the generative adversarial networks(GAN), the inventors generated eight synthetic cohorts resembling the original development set. They were merged to the development set to augment. The inventors performed Boruta algorithm, Monte Carlo method which uses multiple random forests and shadow features. Genes were included as features to predict TCMR. The inventors performed 2000 iterations of random forests. Among the selected features considered predictive by the Boruta algorithm, they calculated median importance to explore the feature importance changes by datasets. The slope graph depicts the flow of the top 30 gene features’ importance. The median importances of genes from the development set to incrementally augmented sets were represented from the left to the right side. Lines represent genes considered important from the development set or which appeared to be important in the augmented sets.

[0354] Figure 12. Learning curves of the machine learning models. Using the generative adversarial networks (GAN), the inventors generated eight synthetic cohorts resembling the original development set. They were merged to the development set to augment. After the Boruta feature importance analysis, the inventors incrementally increased the number of selected features to observe the performance change for all datasets. The performance was measured in precision recall area under the curve (PRAUC) and receiver operating characteristic area under the curve (ROCAUC) in the split test set.

[0355] Figure 13. Calibration plots on the external validation cohort. This figure details the performance of final molecular classifiers assessed on the external validation cohort. The calibration performance is demonstrated as plots in this figure.

[0356] Example 1 (AMR)

[0357] Methods

[0358] The inventors followed the STROBE (Strengthening the Reporting of Observational Studies in Epidemiology) statement checklist for the report of observational cohort studies12. Moreover, they adhered the SAGER (Sex and Gender Equity Research) guidelines for reporting sex and gender. Sex was self-reported by participants13.

[0359] Clinical and biological data

[0360] At time of transplantation, clinical and biological data were collected related to i) recipient characteristics, ii) donor characteristics, iii) transplant characteristics, and iv) immunosuppressive treatment. At time of kidney allograft biopsies, a standardized transplant assessment was performed, comprising i) clinical examination, ii)immunosuppressive treatment, and iii) blood and urinary analyses for standard of care laboratory parameters.

[0361] Immunological phenotyping

[0362] Kidney transplant recipients were tested for the presence of circulating donor-specific anti-HLA-A, -B, and -DR antibodies at baseline and each follow-up visit with single antigen flow bead assays (One Lambda, Inc., Canoga Park, CA, USA). Beads with a normalized mean fluorescence intensity (MFI) of greater than 500 units were considered positive. Immunodominant donor-specific antibody (DSA) was defined as the DSA with the highest MFI.

[0363] HLA typing of recipients and donors were respectively performed by next-generation sequencing (NGS-go, GenDx, The Netherlands) and medium to high resolution sequencespecific primer (Linkage Biosciences, One Lambda, West Hills, CA, USA) in local HLA labs.

[0364] Histological and immunohistochemical phenotyping

[0365] Kidney allograft biopsies were routinely performed at each follow-up visit according to local centers’ practice. One kidney core was paraffin-embedded and formalin-fixed for histological analysis. C4d staining was performed by immunohistochemistry on paraffin-embedded tissue or by immunofluorescence on frozen tissue according to local practices. Biopsies were assessed by local pathologists, and centrally reviewed according to the international and standardized Banff 2019 classification for kidney allograft rejection14.

[0366] Bulk tissue transcriptomic profiling

[0367] The inventors performed bulk tissue transcriptomic profiling of formalin-fixed paraffin-embedded (FFPE) kidney allograft biopsies using the Banff Human Organ Transplant (B-HOT) panel on the NanoString nCounter® platform3. The B-HOT panel, approved by the international Banff consortium and developed by the Banff Molecular Diagnostics Working Group, targets 758 genes crucial for immune response and injury in solid organ transplants. It also includes housekeeping genes and both positive and negative controls to ensure data quality and normalization3.

[0368] Following the quality assessment, the raw expression counts were normalized using the nSolver software in two steps15. First, positive control normalization was conducted using the geometric mean of positive control probes to adjust for technical variations in hybridization efficiency and purification across samples. Second, content normalization was performed using housekeeping genes to adjust for differences in RNA input amount andsample quality. All nCounter Gene Expression assays were normalized using nSolver software. The normalized gene expression data were not log transformed for the analysis.

[0369] Outcomes of interest

[0370] The primary outcomes of interest were antibody-mediated rejection, assessed according to the international Banff 2019 classification16. AMR was defined as either active AMR, chronic active AMR or chronic inactive AMR.

[0371] Descriptive statistics

[0372] For continuous variables, means and standard deviations (SDs) or medians and interquartile ranges (IQRs) were used. The inventors compared means and proportions between groups using Student’s t-test, analysis of variance (ANOVA) (or Mann-Whitney test and Kruskal-Wallis if appropriate), or the chi-squared test (or Fisher’s exact test if appropriate). Values of p<0.05 were considered significant and all tests were two tailed.

[0373] Modelling pipeline

[0374] The study includes seven steps: 1) data pre-process, 2) original data-based molecular classifiers development, 3) conditional tabular GAN (CTGAN) model development, 4) synthetic data generation, 5) feature importance analysis, 6) learning curves generation, 7) machine learning classifiers developments, and 8) classifiers’ assessments. Figure 7 (diagram) demonstrates the overall flow of the study. Transcriptomic data was complete without any missing data. Phenotypic data had missing data and they were used as-is for CTGAM model development.

[0375] The inventors followed the TRIPOD (Transparent Reporting of a multivariable prediction model for Individual Prognosis Or Diagnosis) + Al (Artificial Intelligence) statement for the reporting of the development and validation of the classifiers17.

[0376] Development and validation of original data-based molecular classifiers The derivation cohort was split into development and test sets with 8:2 ratio. The development set was used to perform feature importance analysis, learning curve generation for the final feature selection, and machine learning-based classifiers generation as follows.

[0377] Feature importance analysis

[0378] The inventors performed feature selection and importance analysis using the Boruta algorithm, a wrapper and Monte Carlo method which uses multiple random forests andshadow features. Genes were included as features to predict the binarized AMR and TCMR18. The inventors performed a maximum of 2000 iterations of random forests. Among the selected features considered predictive by the Boruta algorithm, they calculated median importance to explore the feature importance changes by datasets.

[0379] Learning curves generation

[0380] Furthermore, they performed forward stepwise feature selection method to increment the number of features from one to the all Boruta selected features. For this process, they used random forest algorithm as well. Learning curves were generated to observe the changes of predictive performance of the random forest classifiers by adding features. A grid search on fixed one-fold cross-validation was performed to optimize the precision recall area under the curve (PRAUC) for each learning curve. The one-fold cross-validation was randomly selected from the development set and was used for hyperparameter optimization in order to tune the molecular classifiers based on the original data instead of synthetic data. The predictive performance was assessed in the internal validation test set using the PRAUC and the receiver operating characteristic area under the curve (ROCAUC). Features, which were chosen from the forward stepwise selection with the threshold of PRAUC reaching 0.9 quantile of the curve to make the classifiers have enough discriminative performance, were used for the final machine learning developments.

[0381] Machine learning developments

[0382] With the selected features for each set by the Boruta algorithm and the forward stepwise selection through the learning curves, the inventors generated machine learningbased classifiers (random forest) to predict the outcomes of AMR.19The classifier hyperparameter optimization process was performed with grid search on the same fixed one-fold cross-validation used during the learning curves generation. During the optimization, PRAUC was used as a metric to maximize the performance. All continuous normalized gene expression data were standardized to have mean of zero and a standard deviation of one.

[0383] Validation of molecular classifiers’ performances

[0384] Molecular classifiers’ performances were assessed in the external validation cohort. To measure the discrimination performance of the machine learning-based molecular classifiers, the inventors used multiple metrics for comprehensive assessments: 1) ROCAUC, 2) PRAUC, 3) accuracy, 4) balanced accuracy (average of sensitivity and specificity), 5) F1 score (harmonic mean of precision and recall), 6) negative predictivevalue (NPV), 7) positive predictive value (PPV), 8) sensitivity, 9) specificity, and 10) Brier Score. Calibration was assessed with calibration plots. For ROCAUC, PRAUC, and the Brier Score, probabilistic predictions were used to estimate the performance. For the others, cut-offs were calculated to maximize the F1 score on the test set for hard classification. For each assessment, mean of the estimates and 95% confidence intervals were calculated with 1,000 bootstrapping on the external validation cohort.

[0385] Conditional tabular GAN model development

[0386] The development set was used to train CTGAN model. 500 epochs (i.e. iteration over the entire development set) were performed to train the model. Parameters including genes (except for housekeeping genes), donor and recipient demographics, clinical history, biological and histological parameters, institutions, anonymized identifications, and the kidney allograft rejection diagnosis were used to train the CTGAN model. Loss values are measured to observe the model development. The CTGAN model was developed with SDV Python Package (version 1.9.0)20.

[0387] Synthetic data generation

[0388] Using the derived CTGAN model, a synthetic data was generated. The inventors generated synthetic cohort with twice the size of the development set (n>1,000) to increase the machine learning predictive power and robustness of the downstream analyses. One sample corresponds to a biopsy and no multiple biopsies were generated per patient. The synthetic cohort was sliced by the size of 25% of the development set to create eight synthetic datasets. These eight synthetic datasets are inclusive (i.e. the second dataset comprises the first and the second slices and the third dataset comprises the first, the second, and the third slices, and so on), allowing for a progressive accumulation of data. The eight synthetic datasets were attached to the end of the development set to generate augmented datasets. Comprising the development and the eight augmented datasets, the total of nine datasets were prepared and analyzed in the study. Each synthetic and augmented dataset is specified in tables, figures, and text using the notation GanN and GanAugN, respectively, where N refers to the percentage of sample size of the development set. For example, Gan50 refers to a synthetic set, which includes 50% of the sample size of the development set (the first two slices). Likewise, GanAug50 refers to an augmented set, which Gan50 set was attached to the end of the development set. Thus, the eight synthetic datasets are from Gan25 (the size of 25% of the development set) to Gan200 (the size of 200% of the development set). Likewise, the smallest augmented dataset is denoted as GanAug25 and the largest is denoted as GanAug200.Assessment of similarities between synthetic and original data The statistical similarities between the development set and the synthetic sets were assessed using SDV Python Package to assess Kolmogorov-Smirnov (KS) statistic complement and Correlation Similarity for continuous variables and Total Variance (TV) Distance complement and Contingency Similarity and for categorical variables20-22.

[0389] The definition of each metric is as follows:

[0390] For univariate comparison, column shapes were measured for all columns in KSComplement for numerical value and TVComplement for categorical value. For bivariate comparison (column pair trends), Correlationsimilarity was used for numerical pair of values and Contingencysimilarity was used for categorical pair of values or numerical and categorical pair of values.

[0391] Due to the rounding during calculation, some scores might be little bit out of ranges.

[0392] KSComplement

[0393] This metric uses the Kolmogorov-Smirnov (KS) statistic (Massey Jr. FJ. The Kolmogorov-Smirnov Test for Goodness of Fit. Journal of the American Statistical Association 1951; 46: 68-78). It is a measure of the difference between a sample distribution and a reference or theoretical distribution. In this study, the inventors compare the synthetic cohort to the original derivation cohort. A numerical distribution is converted into its cumulative distribution function (CDF). The KS statistic is the maximum difference between the two CDFs.

[0394] KSComplement is defined as 1 - KS statistic.

[0395] This metric ignores missing values and is meant for continuous and numerical data. The score ranges from 0 (worst, the two datasets are as different as they can be) to 1 (best, the two datasets are the same).

[0396] TVComplement

[0397] This metric computes the Total Variation Distance (TVD) between the real and synthetic variables by determining the probability of each category value and using it to compare the differences in probabilities. The formula below demonstrates how the TVD statistic achieves this comparison.

[0398]

[0399] δ(R,S) = ½ Σ |R_ω − S_ω|TVComplement score is defined as 1 − δ(R, S). Here, co describes all the possible categories in a column, Q. Meanwhile, R and S refer to the real and synthetic frequencies for those categories.

[0400] This metric computes the similarity of a real value variable and a synthetic This metric ignores missing values and is meant for discrete and categorical data. The score ranges from 0 (worst, the two datasets are as different as they can be) to 1 (best, the two datasets are the same).

[0401] Correlationsimilarity

[0402] This metric computes a correlation coefficient on the real and synthetic data, R and S, respectively. A and B represents a pair of variables.

[0403] |S_A,B − R_A,B|

[0404] score = 1 -

[0405]

[0406] 2

[0407] The metric supports both the Pearson correlation coefficient and the Spearman’s rank correlation coefficient. The score ranges from 0.5 (worst, the two correlations are as different as they can be) to 1 (best, the two correlations are the same).

[0408] ContinoencySimilarity

[0409] This metric computes the similarity of a pair of categorical columns between the real and synthetic datasets.

[0410] The formula below summarizes the scoring process, where a and ft describe all the possible categories in variable A and B, respectively and R and S refer to the real and synthetic frequencies for those categories, respectively.

[0411] score = 1 − ½ Σ Σ |S_{α,β} − R_{α,β}|

[0412]

[0413] aEA PEB

[0414] The score ranges from 0 (worst, the contingency table is as different as can be) to 1 (best, the contingency table is exactly the same between the real and synthetic data).

[0415] Development and validation of augmented data-based molecular classifiers With the eight augmented datasets, they inventors followed the same machine learning pipeline as the original data-based molecular classifiers: 1) feature importance analysis, 2) learning curves generation, 3) machine learning developments for the eight augmented datasets, and 4) validation of the eight molecular classifiers’ performances.Software and packages

[0416] Descriptive analyses and machine learning classifier analyses were conducted using R (version 4.3.1, R Foundation for Statistical Computing) and RStudio (version 2023.9.0.463). Synthetic data were generated using Python (version 3.10) and JupyterLab (version 4.0.9). R packages used for descriptive, data, and machine learning analyses were: tidyverse (version 2.0.0), ggplot2 (version 3.4.3), forcats (version 1.0.0), Boruta (version 8.0.0), yardstick (version 1.2.0), rstatix (version 0.7.2), pROC (version 1.18.4), stringr (version 1.5.0), caret (version 6.0-94), caretEnsemble (version 2.0.3), randomForest (version 4.7-1.1), rsample (version 1.2.0), compareGroups (version 4.7.1), viridis (version 0.6.4), foreach (version 1.5.2), doParallel (version 1.0.17), CGPfunctions (version 0.6.3), and cutpointr (version 1.1.2). Python packages used for synthetic data generation were: numpy (version 1.26.2), torch (version 2.1.1), sdv (version 1.9.0), and pandas (version 2.1.3).

[0417] Results

[0418] Baseline characteristics of the derivation cohort

[0419] In total, 639 biopsies (363 [57.9%] protocol and 264 [42.1%] clinically indicated biopsies) from 578 kidney transplant recipients with gene expression data were included in the derivation cohort. The mean recipient age was 49.6 ± 14.6 (SD) with 228 (39.4%) females. The mean donor age was 53.9 ± 16.9 with 451 (78.2%) deceased at donation. 199 (35.3%) donors were expanded criteria donor and 13 (2.2%) were ABO incompatible transplant. The baseline characteristics of the patients and biopsies from the derivation cohort are described in Tables 1a and 1b.Table 1a. Characteristics of the patients of the derivation cohort

[0420]

[0421] N* Derivation cohort (n=578) Recipient's characteristics

[0422] Age (years), Mean (SD) 578 49.6 (14.6) Sex female, N (%) 578 228 (39.4%) Etiology of recipient's end-stage kidney disease, N 578

[0423] (%)

[0424] Autosomal dominant polycystic kidney disease 85 (14.7%) Diabetes 42 (7.3%) Glomerulonephritis 161 (27.9%) Polycystic kidney disease 1 (0.2%) Tubulo-interstitial 56 (9.7%) Vascular 38 (6.6%) Unknown 129 (22.3%) Others 66 (11.4%) Donor's characteristics

[0425] Age (years), Mean (SD) 564 53.9 (16.9) Deceased donor, N (%) 577 451 (78.2%) Expanded criteria donor, N (%) 564 199 (35.3%) Transplant baseline characteristics

[0426] Prior kidney transplantation, N (%) 578 105 (18.2%) Cold ischemia time (hours), Median [25th;75th] 571 15.2 [8.1;22.0] Delayed graft function, N (%) 561 114 (20.3%) Number of HLA A / B / DR mismatches, Mean (SD) 572 3.4 (1.5) Donor-specific antibody at the time of transplant, N 537 138 (25.7%) (%)

[0427] ABO incompatible transplant, N (%) 578 13 (2.2%) Immunosuppressive protocol at time of

[0428] transplantation

[0429] Induction therapy, N (%) 571

[0430] Anti-thymocyte globulin 260 (45.5%) Basiliximab 294 (51.5%) Other 3 (0.5%) Baseline immunosuppression, N (%)

[0431] Steroids 562 525 (93.4%) Mycophenolate mofetil or mycophenolic acid / 560 474 (84.6%) / 6 (1.1 %) azathioprine

[0432] Tacrolimus / cyclosporine 566 469 (82.9%) / 80 (14.1%) Everolimus / sirolimus 560 50 (8.9%) / 0 (0%) Belatacept 561 6 (1.1%)

[0433] *N refers to the number of patients with available dataTable 1b. Characteristics of the biopsies of the derivation cohort

[0434] N* Derivation cohort

[0435]

[0436] (n=639) Indication of biopsy, N (%) 627

[0437] Protocol 363 (57.9%)

[0438] For cause 264 (42.1 %)

[0439] Time after transplantation (months), Median [25th;75th] 639 12.0 [3.0; 13.0] Diagnosis, N (%) 639

[0440] Active AMR 129 (20.2%)

[0441] Acute TCMR 99 (15.5%)

[0442] ATI without rejection 16 (2.5%)

[0443] BK virus nephropathy 25 (3.9%)

[0444] Chronic active AMR 61 (9.5%)

[0445] Chronic active TCMR 24 (3.8%)

[0446] Chronic inactive AMR 3 (0.5%) Glomerulonephritis (recurrent or de novo) 13 (2.0%)

[0447] Isolated IFTA > 2 113 (17.7%)

[0448] Normal or minimal changes 120 (18.8%)

[0449] Pristine 9 (1.4%)

[0450] Others 27 (4.2%)

[0451] Biology at time of biopsy

[0452] Serum creatinine (μmol / l), Median [25th;75th] 631 142.0 [110.0;198.0] Proteinuria to creatinine ratio (g / g), Median [25th;75th] 594 0.1 [0.1;0.5] Immunology at time of biopsy

[0453] Positive anti-HLA antibodies, N (%) 561 278 (49.6%) Presence of donor-specific antibodies, N (%) 604 232 (38.4%)

[0454] Class of donor-specific antibodies, N (%) 576

[0455] I 58 (10.1%)

[0456] II 115 (20.0%)

[0457]

[0458] I and II 31 (5.4%)

[0459] Class of immunodominant DSA, N (%) 575

[0460] I 64 (11.1%)

[0461] II 139 (24.2%) Immunosuppressive therapy at the time of biopsy,

[0462] N (%)

[0463] Steroids maintenance 621 547 (88.1%) Mycophenolate mofetil or mycophenolic acid / 617 507 (82.2%) / 19 (3.1%) azathioprine maintenance

[0464] Tacrolimus / cyclosporine maintenance 616 489 (79.4%) / 83 (13.5%) Everolimus / sirolimus maintenance 613 52 (8.5%) / 10 (1.6%) Belatacept maintenance 609 20 (3.3%)

[0465] *N refers to the number of biopsies with available data

[0466] AMR was diagnosed in 193 biopsies (129 active AMR, 61 chronic active AMR, and 3 chronic inactive AMR). Other diagnoses were in 323 biopsies (16 acute tubular injury [ATI] without rejection, 25 BK virus nephropathy, 13 glomerulonephritis [recurrent or de novo], 113 isolated IFTA > 2, 120 normal or minimal changes, 9 pristine, and 27 others), representing the full spectrum of kidney allograft phenotypes that can be encountered inreal-life. The baseline characteristics of the histology of biopsies of the derivation cohort are detailed in Supplementary Table 1.

[0467] Supplementary Table 1. Characteristics of the histology of biopsies of the derivation cohort

[0468] Glomerulitis (g) score

[0469] Category % (n=627)

[0470] 0 69.2%

[0471] >1 31.3%

[0472] Peritubular capillaritis (ptc) score

[0473] Category % (n=627)

[0474] 0 61.2%

[0475] >1 33.5%

[0476] Double contour (eg) score

[0477] Category % (n=627)

[0478] 0 88.4%

[0479] >1 11.2%

[0480] Interstitial inflammation (i) score

[0481] Category % (n=627)

[0482] 0 73.4%

[0483]

[0484] 26.5%

[0485] Tubulitis (t) score

[0486] Category % (n=627)

[0487] 0 58.1%

[0488] >1 41.8%

[0489] Intimal arteritis (v) score

[0490] Category % (n=627)

[0491] 0 84.7%

[0492] >1 9.1%

[0493] Interstitial fibrosis (ci) score

[0494] Category % (n=627)

[0495] 0 31.1%

[0496]

[0497] 68.3%

[0498] Tubular atrophy (ct) score

[0499] Category % (n=627)

[0500] 0 30.3%

[0501] >1 69.1%

[0502] IFTA score

[0503] Category % (n=627)

[0504] 0 31.9%61.9%

[0505] Total inflammation (ti) score

[0506] Category % (n=627)

[0507] 0 58.4%

[0508] >1 34.3%

[0509] Inflammation in areas of IFTA (i-IFTA) score

[0510] Category % (n=627)

[0511] 0 49.0%

[0512] >1 28.4%

[0513] Tubulitis in scarred cortex (t-IFTA) score

[0514] Category % (n=627)

[0515] 0 4.5%

[0516]

[0517] 9.7%

[0518] Vascular fibrous intimal thickening (cv) score

[0519] Category % (n=627)

[0520] 0 28.1%

[0521] >1 61.6%

[0522] Arteriolar hyalinosis (ah) score

[0523] Category % (n=627)

[0524] 0 33.5%

[0525] >1 63.3%

[0526] Hyaline arteriolar thickening (aah) score

[0527] Category % (n=627)

[0528] 0 7.5%

[0529]

[0530] 2.4%

[0531] Mesangial matrix expansion (mm) score

[0532] Category % (n=627)

[0533] 0 73.5%

[0534] >1 21.1%

[0535] C4d deposition in peritubular capillaritis

[0536] Category % (n=627)

[0537] No 67.9%

[0538] Yes 23.9%

[0539] C4d type

[0540] Category % (n=627)

[0541] IF 12.4%

[0542] IHC 89.5%

[0543] Among the biopsies, a development set and a test set are randomly split.Baseline characteristics of the synthetic sets and statistical similarities The development set was used to develop a CTGAN model to create synthetic sets. The details of the hyperparameters results are as follows: -'enforce_min_max_values': True; 'enforce_rounding': True; 'locales': None; 'embedding_dim': 128; 'generator_dim': (256, 256); 'discriminator_dim': (256, 256); 'generator_lr': 0.0002; 'generator_decay': 1e-06; 'discriminator_lr': 0.0002; 'discriminator_decay': 1e-06; 'batch_size': 500; 'discriminator_steps': 1; 'log_frequency': True; 'verbose': True; 'epochs': 500; 'pac': 10; 'cuda': True.

[0544] Eight synthetic sets were generated, sample sizes of 25%, 50%, 75%, 100%, 125%, 150%, 175%, and 200% of the development set. Overall, baseline characteristics show stable values across the synthetic sets.

[0545] Between the entire synthetic set (i.e. Gan200) and the original development set, the inventors assessed univariate and bivariate similarity scores. Overall, the two datasets show high similarity.

[0546] Original data-based molecular classifiers for kidney allograft rejection

[0547] The molecular classifiers based on the development set with selected genes were assessed on the external validation cohort. For AMR, the molecular classifier showed discrimination performance with ROCAUC of 0.829 (95% Cl 0.778 - 0.872) and PRAUC of 0.711 (0.618 - 0.797). The classifier also showed Brier Score of 0.165 (0.146 - 0.186).

[0548] Feature importance analysis

[0549] The genes associated with AMR were consistently considered important in all nine original development and augmented sets: CXCL11, GNG11, GBP1, PLA1A, GBP4, WARS, CXCR3, PPM1F, CXCL9, ROBO4, SH2D1B, CD5, CD96, CD3E, GNLY, CXCL10, APOL1, ADGRL4, MAPK11, PRF1, LEF1, FCGR3AB, SELL, TM4SF18, CD247, HYAL2, CD27, TAP1, IRF1, LAP3, EGFR, APOL2, HLAF, PLAAT4 and CDH13. Genes related to B cell activation (e.g., CD209, CD38), T cell activation (e.g., CD4, STAT1), macrophage activation (e.g., CD74), NK specific activation (e.g., KIR3DL1, HLAE), endothelial activation (e.g., PECAM1), and cellular injury and adaptive immune responses (e.g., GZMB, CTLA4) were only considered important features in the augmented datasets (Figure 8).

[0550] Learning curves analysis

[0551] With the selected features by Boruta algorithm, learning curves were generated with random forest algorithm. Incremental forward stepwise feature selection was performed and assessed on the test set (Figure 9).For AMR, the learning curve on the original development set shows steep increase in performance up to 24 features followed by unimproved performance with fluctuations. At the performance reaching 0.9 quantile of PRAUC, numbers of features selected for the final molecular classifier derivation were 26 for the development set. Selected genes to develop a molecular classifier for the development set are detailed in Supplementary Table 2.

[0552] On the other hand, GanAug datasets require more features to learn to classify the outcome AMR, yet, with constant improvements in performance in both PRAUC and ROCAUC, achieve outperforming the original development set. The learning curves based on the augmented sets do not fluctuate and show stable performance. At the performance reaching 0.9 quantile of PRAUC, numbers of features selected for the final molecular classifier derivation are detailed in Supplementary Table 2.Supplementary Table Selected genes to develop AMR molecular classifiers by dataset

[0553] GanAug25 GanAug50 GanAug75 GanAug100 GanAug125 GanAug150 GanAug175 GanAug200 1 CXCL11 CXCL11 CXCL11 GBP1 GBP1 GBP1 GBP1 GBP1 GBP1 2 GNG11 GNG11 GBP1 CXCL11 GNG11 GNG11 CXCL11 CXCL11 CXCL11 3 GBP1 GBP1 GNG11 GNG11 CXCL11 GNLY GNLY GNLY GNLY 4 PLA1A ROBO4 GNLY PLA1A PLA1A PLA1A GNG11 CXCL9 CXCL9 5 GBP4 PLA1A PLA1A CXCL9 PPM1F CXCL11 PPM1F GNG11 GNG11 6 WARS SIH2D1B CXCL9 PPM IF ROBO4 CXCL9 PLA1A PPM1F PPM1F 7 CXCR3 CXCL9 WARS ROBO4 CXCL9 CXCL9 SELL SELL 8 PPM1F CXCR3 GNLY GNLY SELL PLA1A PLA1A 9 PTPN7 GBP4 WARS JAK3 WARS WARS ADGRL4 10 CXCL9 PPM1F SIH2D1B SH2D1B ADGRL4 SH2D1B 11 ROBO4 CXCL10 CXCL10 SELL CXCR3 CXCR3 12 SH2D1B WARS IL16 WARS FCGR3AB FCGR3AB 13 CD5 GNLY SELL APOL2 CD209 APOL2 14 CD96 PTPN7 GBP4 CXCL10 SH2D1B CALHM6 15 CD3E CD5 EGFR EGFR APOL1 JAK3 16 GNLY CD96 CD209 CD27 CD5 CD209 o 17 CXCL10 MAPK11 CXCR3 CALHM6 CD27 CD27 18 APOL1 CD3E LEF1 PTPN7 JAK3 TAP1 19 ICOS CD209 APOL1 LEF1 GZMB CD5 20 IDO1 CCL3L1 JAK3 IL16 CD96 CD3E 21 CXCL13 ADGRL4 CALHM6 GBP4 CALHM6 ROBO4 22 ADGRL4 EGFR CD5 CXCR3 APOL2 PLAAT4 23 CD28 APOL1 CD27 FCGR3AB GBP4 ST8SIA4 24 HSPA12B CD28 PTPN7 ADGRL4 PLAAT4 WARS 25 PSMB9 TRDC CD247 CD3E ROBO4 CD96 26 LCK CD27 ADGRL4 CD5 LEF1 CD247 27 XAF1 CD28 TRDC CD247 CXCL10 28 TAP1 FCGR3AB CD209 CD3E APOL1 29 IL16 PRF1 ICOS CXCL10 GBP4 30 PRF1 TRDC CD96 TAP1 PRF1 31 LEF1 APOL2 APOL1 EGFR CD28 32 IL21IR HYAL2 LCK LCK LEF1 33 TCF7 TAP1 CD247 PRF1 EGFR

[0554]

[0555] 34 LAP3 ICOS PRF1 MAPK11 HLAF35 LCK IRF1 CD28 IL21R CIITA 36 SELL CD3E PL AT4 PTPN7 IRF1 37 IL2RB HLAF HLAF BTLA FCER1G 38 NOS3 IL21R LCP2 HLAF GZMB 39 FCGR3AB CD96 IRF1 ITGA4 MAPK3 40 IGOS LCK NFKB2 ST8SIA4 MAPK11 41 SIRPG PL AT4 HLADPB1 IRF1 HLADPA1 42 IRF1 MAPK11 TAP1 CDH5 KIT 43 TM4SF18 CD45R0 IL21R FCER1G PSMB10 44 MIR155HG TM4SF18 GZMB CCL3L1 SPIB 45 PSME2 CXCL13 MAPK11 IL16 CCL3L1 46 PSMB9 BTLA SLAMF6 NFKB2

[0556] 47 APOL2 IDO1 HYAL2 IGOS

[0557] 48 MAPK3 LAP3 TM4SF18 SOX7

[0558] 49 HLAF CCL3L1 SIRPG LAIR1

[0559] 50 IDO1 CD3G LAP3 LAP3

[0560] 51 CXCL13 NFKB2 CXCL13 LCP2

[0561] 52 CD247 LCP2 PSMB10 HLAE

[0562] 53 BTLA KLRG1 PECAM1 PSMB10

[0563] 54 NLRC5 HLADPB1I TNFRSF1B PECAM1

[0564] 55 GBP5 CTSS NLRC5 RNF149

[0565] 56 CALHM6 NOD1 ITGA4 NLRC5

[0566] 57 JAK2 PSME2 TRAT1

[0567] 58 PHEX IL17A IDO1

[0568] 59 CX3CL1 SLAMF6 CDH13

[0569] 60 PECAM1 NKG7

[0570] 61 PSMB10 CDH5

[0571] 62 TCF7 NFATC2

[0572] 63 GZMB IKZF1

[0573] 64 TM4SF1 CTSS

[0574] 65 NKG7 CCL3L1

[0575] 66 IL2RB ITGAX

[0576] 67 IL2RG IGF2R

[0577] 68 TNFRSF1B CD3G

[0578] 69 XAF1 ISG20

[0579] 70 NFATC2

[0580] 71 CDH13Augmented data-based molecular classifiers for AMR

[0581] Synthetic biopsies were generated and merged to the development set to augment the data. The performances of the molecular classifiers based on the augment data sets were assessed on the external validation cohort.

[0582] For AMR, the molecular classifiers showed discrimination performance with ROCAUCs of 0.772 (95% Cl 0.716 - 0.821), 0.861 (0.816 - 0.903), 0.806 (0.754 - 0.852), 0.807 (0.756 - 0.852), 0.866 (0.823 - 0.905), 0.851 (0.805 - 0.892), 0.817 (0.770 - 0.861), 0.829 (0.780 - 0.871), and PRAUCs of 0.636 (0.537 - 0.734), 0.801 (0.731 - 0.859), 0.672 (0.577 - 0.763), 0.671 (0.580 - 0.760), 0.796 (0.721 - 0.860), 0.755 (0.663 - 0.834), 0.679 (0.583 - 0.774), and 0.717 (0.627 - 0.798) for the GanAug25, GanAug50, GanAug75, GanAug100, GanAug125, GanAug150, GanAug175, and GanAug200 sets, respectively. The classifiers showed overall fit with Brier Scores of 0.188 (0.172 - 0.206), 0.171 (0.150-0.194), 0.178 (0.162 - 0.196), 0.177 (0.162 - 0.193), 0.168 (0.145 -0.190), 0.174 (0.153 -0.193), 0.182 (0.166 - 0.198), and 0.171 (0.158 - 0.187) in the same order.

[0583] The calibration plots are available on Figure 10. Cut-offs were calibrated to maximize F1 score (harmonic mean of precision and recall) for hard classification (antibody-mediated rejection: Model based on Development set: 0.402; Model based on GanAug25 set: 0.270; Model based on GanAug50 set: 0.476; Model based on GanAug75 set: 0.240; Model based on GanAug100 set: 0.280; Model based on GanAug125 set: 0.344; Model based on GanAug150 set: 0.386; Model based on GanAug175 set: 0.198; Model based on GanAug200 set: 0.440).

[0584] Table 3. Performance in multimetrics in the external validation cohort

[0585] The table illustrates detailed performance metrics in the external validation cohort. The parentheses represent 95% confidence intervals calculated with 1000 bootstraps. PRAUC, precision recall area under the curve; ROCAUC, receiver operating characteristic area under the curve; NPV negative predictive values; PPV, positive predictive values.lA^KMAreiectionJAMRl

[0586] Model ROCAUC PRAUC Acm®cv W«£ei F1 score NPV PPV Sensitivity

[0587]

[0588] Brier tairaa:

[0589]

[0590] 0.829 0.711 0.729 0.754 0 680 0.884 0.575 0.833 0.674 0 165. QffitaUKlt (0.778 - (0.618 - (0.682 - (0.709 - (0.620 - (0.836 - (0.500 - (0.763 - (0.614 - (0.146 - 0.872) 0.797) 0.778) 0.799) 0.740) 0.929) 0.648) 0.893) 0.734) 0.186) 0.772 0.636 0 542 0.632 0 582 0.894 0.426 0.924 0.340 0 188 GanAug25 (0.716 - (0.537 - (0.490 - (0.595 - (0.524 - (0.826 - (0.369 - (0.873 - (0.278 - (0.172 - 0.821) 0.734) 0.598) 0.669) 0.643) 0.957) 0.492) 0.971) 0.401) 0.206) 0.861 0.801 0.740 0.770 0 697 0.904 0.584 0.865 0.674 0 171 GanAug50 (0.816 - (0.731 - (0.697 - (0.729 - (0.637 - (0.861 - (0.515 - (0.805 - (0.613 - (0.150 - 0.903) 0.859) 0.784) 0.812) 0.753) 0.945) 0.657) 0.923) 0.732) 0.194) 0.806 0.672 0 507 0.613 0 573 0.923 0.410 0.957 0.269 0 178 GanAug75 (0.754 - (0.577 - (0.455 - (0.581 - (0.517 - (0.859 - (0.355 - (0.920 - (0.212 - (0.162 - 0.852) 0.763) 0.560) 0.646) 0.632) 0.983) 0.471) 0.991) 0.326) 0.196) 0.807 0.671 0 595 0.676 0.616 0.929 0.459 0.941 0.411 0 177 GanAug100 (0.756 - (0.580 - (0.545 - (0.640 - (0.557 - (0.879 - (0.400 - (0.897 - (0.351 - (0.162 - 0.852) 0.760) 0.650) 0.714) 0.677) 0.978) 0.525) 0.981) 0.477) 0.193) 0.866 0.796 0.671 0.722 0 651 0.905 0.514 0.891 0.554 0 168 GanAug125 (0.823 - (0.721 - (0.621 - (0.680 - (0.591 - (0.856 - (0.448 - (0.835 - (0.491 - (0.145 - 0.905) 0.860) 0.720) 0.763) 0.708) 0.949) 0.587) 0.942) 0.619) 0.190) 0.851 0 755 0 677 0 727 0 656 0.907 0 519 0.891 0 563 0 174 GanAug150 (0.805 - (0.663 - (0.630 - (0.687 - (0.595 - (0.859 - (0 453 - (0.835 - (0 504 - (0.153 - W 0 892) 0.834) 0.726) 0 769) 0.716) 0.951) 0 591) 0.944) 0 628) 0.193) 0.817 0 679 0 419 0 550 0 537 0.903 0 371 0.975 0 125 0 182 GanAug175 (0770 - (0 583 - (0.364 - (0 524 - (0.481 - (0.788 - (0 319 - (0.944 - (0.081 - (0.166 - 0.861) 0 774) 0.472) 0 575) 0.592) 1.000) 0 426) 1.000) 0.170) 0.198) 0.829 0 717 0 751 0 757 0 682 0.861 0 610 0.776 0 737 0 171 GanAug200 (0780 - (0.627 - (0.708 - (0.713 - (0.617 - (0.812 - (0 533 - (0.701 - (0.682 - (0.158 -

[0591]

[0592] 0.871) 0 798) 0.793) 0 803) 0.744) _ 0.908) _ 0 685) _ 0.851) _ 0 790) _ 0.187)DISCUSSION

[0593] In this international multicenter cohort study, the inventors developed and validated machine learning based molecular classifiers to predict antibody-mediated rejection. They showed that augmenting the original dataset with generative synthetic transcriptomic data can increase the performances of these molecular classifiers and capture additional biologically relevant predictive genes.

[0594] Original data-based molecular classifiers showed steep increase followed by stable performance during the learning curve analysis. In the external validation, the original data-based molecular classifiers showed good performance in multiple metrics including ROCAUC and PRAUC in the external validation cohort for predicting AMR. Boruta algorithm captured AMR associated genes as highly important features.

[0595] Molecular classifiers based on the synthetic data augmentation showed stable performance increment during the learning curve analysis. In the external validation, the derived molecular classifiers showed increased performance in the synthetically augmented data in AMR (ROCAUC 0.866 from 0.829 and PRAUC 0.801 from 0.711). In other metrics, the molecular classifiers showed similar performance increase trend in the augmented datasets. Calibration plots showed comparable results across the development and augmented datasets.

[0596] The inventors found that the gene importance varies through the different sets of data by feature selection / importance analysis and learning curve generation. Compared to the original development set, the synthetically augmented sets allowed capturing genes biologically related to B cell activation, T cell activation and macrophage activation, NK-specific activation, and cellular injury and adaptive immune. This newly captured trend in the machine-generated augmented datasets enhances and expands upon previous studies, revealing additional biologically relevant genes for diagnosing kidney rejection23.

[0597] Incorporating GAN-based data augmentation into clinical practice alongside standard monitoring factors presents an opportunity to enhance cost-effectiveness and augment the breadth of data available for analysis24. By satisfying the growing demand for data and contributing to the development of more comprehensive predictive classifiers for AMR diagnosis, this approach can lead to significant cost savings while simultaneously safeguarding the privacy of kidney allograft recipients. These advancements offer a promising avenue for improving the management of renal transplant patients by refining the clinical indications for biopsies and optimizing overall care protocols.

[0598] This study presents multiple strengths. First, the study cohort comprises a comprehensive set of multimodal datatypes including clinical, biological, histological, donor, recipient, and transplant parameters, in addition to transcriptomic data. This extensivedataset allows for a more holistic analysis and better understanding of kidney allograft rejection. Second, the inventors used multidisciplinary study methods. The study incorporates heuristic selection by kidney transplant experts from the Banff Human Organ Transplant (B-HOT) panel, ensuring the inclusion of the most relevant genes and markers3. Additionally, it integrates specialized expertise from nephrologists and pathologists for multimodal kidney rejection diagnosis. Advanced deep-learning techniques are employed to augment data, aiming to enhance the robustness and accuracy of the classifiers’ performance. Furthermore, multiple machine learning methods are utilized for feature selection, learning curve analysis, hyperparameter optimization, and classifiers’ generation. Finally, external validation cohorts are included to strengthen the reliability and generalizability of the study's findings. Validating classifiers on independent datasets ensures robust performance across different clinical settings.

[0599] In conclusion, the inventors derived and validated machine learning-based molecular classifiers for kidney allograft rejection and enhanced their performance with deep learning-driven synthetic data augmentation. Using synthetic transcriptomic data, they captured additional predictive genes and showed potential for reducing the necessary number of transcriptomic sequencings and costs of time-consuming procedures.

[0600] Example 2: TCMR

[0601] Methods

[0602] The inventors followed the STROBE (Strengthening the Reporting of Observational Studies in Epidemiology) statement checklist for the report of observational cohort studies12. Moreover, they adhered the SAGER (Sex and Gender Equity Research) guidelines for reporting sex and gender. Sex was self-reported by participants13.

[0603] Clinical and biological data

[0604] At time of transplantation, clinical and biological data were collected related to i) recipient characteristics, ii) donor characteristics, iii) transplant characteristics, and iv) immunosuppressive treatment. At time of kidney allograft biopsies, a standardized transplant assessment was performed, comprising i) clinical examination, ii) immunosuppressive treatment, and iii) blood and urinary analyses for standard of care laboratory parameters.Immunological phenotyping

[0605] Kidney transplant recipients were tested for the presence of circulating donor-specific anti-HLA-A, -B, and -DR antibodies at baseline and each follow-up visit with single antigen flow bead assays (One Lambda, Inc., Canoga Park, CA, USA). Beads with a normalized mean fluorescence intensity (MFI) of greater than 500 units were considered positive. Immunodominant donor-specific antibody (DSA) was defined as the DSA with the highest MFI.

[0606] HLA typing of recipients and donors were respectively performed by next-generation sequencing (NGS-go, GenDx, The Netherlands) and medium to high resolution sequencespecific primer (Linkage Biosciences, One Lambda, West Hills, CA, USA) in local HLA labs.

[0607] Histological and immunohistochemical phenotyping

[0608] Kidney allograft biopsies were routinely performed at each follow-up visit according to local centers’ practice. One kidney core was paraffin-embedded and formalin-fixed for histological analysis. C4d staining was performed by immunohistochemistry on paraffin-embedded tissue or by immunofluorescence on frozen tissue according to local practices. Biopsies were assessed by local pathologists, and centrally reviewed according to the international and standardized Banff 2019 classification for kidney allograft rejection14.

[0609] Bulk tissue transcriptomic profiling

[0610] The inventors performed bulk tissue transcriptomic profiling of formalin-fixed paraffin-embedded (FFPE) kidney allograft biopsies using the Banff Human Organ Transplant (B-HOT) panel on the NanoString nCounter® platform3. The B-HOT panel, approved by the international Banff consortium and developed by the Banff Molecular Diagnostics Working Group, targets 758 genes crucial for immune response and injury in solid organ transplants. It also includes housekeeping genes and both positive and negative controls to ensure data quality and normalization3.

[0611] Following the quality assessment, the raw expression counts were normalized using the nSolver software in two steps15. First, positive control normalization was conducted using the geometric mean of positive control probes to adjust for technical variations in hybridization efficiency and purification across samples. Second, content normalization was performed using housekeeping genes to adjust for differences in RNA input amount and sample quality. All nCounter Gene Expression assays were normalized using nSolver software. The normalized gene expression data were not log transformed for the analysis.Outcomes of interest

[0612] The primary outcomes of interest were T cell-mediated rejection, assessed according to the international Banff 2019 classification16. TCMR was defined as acute TCMR or chronic active TCMR.

[0613] Descriptive statistics

[0614] For continuous variables, means and standard deviations (SDs) or medians and interquartile ranges (IQRs) were used. The inventors compared means and proportions between groups using Student’s t-test, analysis of variance (ANOVA) (or Mann-Whitney test and Kruskal-Wallis if appropriate), or the chi-squared test (or Fisher’s exact test if appropriate). Values of p<0.05 were considered significant and all tests were two tailed.

[0615] Modelling pipeline

[0616] The study includes seven steps: 1) data pre-process, 2) original data-based molecular classifiers development, 3) conditional tabular GAN (CTGAN) model development, 4) synthetic data generation, 5) feature importance analysis, 6) learning curves generation, 7) machine learning classifiers developments, and 8) classifiers’ assessments. Figure 7 (diagram) demonstrates the overall flow of the study. Transcriptomic data was complete without any missing data. Phenotypic data had missing data and they were used as-is for CTGAM model development.

[0617] The inventors followed the TRIPOD (Transparent Reporting of a multivariable prediction model for Individual Prognosis Or Diagnosis) + Al (Artificial Intelligence) statement for the reporting of the development and validation of the classifiers17.

[0618] Development and validation of original data-based molecular classifiers The derivation cohort was split into development and test sets with 8:2 ratio. The development set was used to perform feature importance analysis, learning curve generation for the final feature selection, and machine learning-based classifiers generation as follows.

[0619] Feature importance analysis

[0620] The inventors performed feature selection and importance analysis using the Boruta algorithm, a wrapper and Monte Carlo method which uses multiple random forests and shadow features. Genes were included as features to predict the binarized AMR and TCMR18. The inventors performed a maximum of 2000 iterations of random forests. Amongthe selected features considered predictive by the Boruta algorithm, they calculated median importance to explore the feature importance changes by datasets.

[0621] Learning curves generation

[0622] Furthermore, they performed forward stepwise feature selection method to increment the number of features from one to the all Boruta selected features. For this process, they used random forest algorithm as well. Learning curves were generated to observe the changes of predictive performance of the random forest classifiers by adding features. A grid search on fixed one-fold cross-validation was performed to optimize the precision recall area under the curve (PRAUC) for each learning curve. The one-fold cross-validation was randomly selected from the development set and was used for hyperparameter optimization in order to tune the molecular classifiers based on the original data instead of synthetic data. The predictive performance was assessed in the internal validation test set using the PRAUC and the receiver operating characteristic area under the curve (ROCAUC). Features, which were chosen from the forward stepwise selection with the threshold of PRAUC reaching 0.9 quantile of the curve to make the classifiers have enough discriminative performance, were used for the final machine learning developments.

[0623] Machine learning developments

[0624] With the selected features for each set by the Boruta algorithm and the forward stepwise selection through the learning curves, the inventors generated machine learningbased classifiers (random forest) to predict the outcomes of TCMR.19The classifier hyperparameter optimization process was performed with grid search on the same fixed one-fold cross-validation used during the learning curves generation. During the optimization, PRAUC was used as a metric to maximize the performance. All continuous normalized gene expression data were standardized to have mean of zero and a standard deviation of one.

[0625] Validation of molecular classifiers’ performances

[0626] Molecular classifiers’ performances were assessed in the external validation cohort. To measure the discrimination performance of the machine learning-based molecular classifiers, the inventors used multiple metrics for comprehensive assessments: 1) ROCAUC, 2) PRAUC, 3) accuracy, 4) balanced accuracy (average of sensitivity and specificity), 5) F1 score (harmonic mean of precision and recall), 6) negative predictive value (NPV), 7) positive predictive value (PPV), 8) sensitivity, 9) specificity, and 10) Brier Score. Calibration was assessed with calibration plots. For ROCAUC, PRAUC, and theBrier Score, probabilistic predictions were used to estimate the performance. For the others, cut-offs were calculated to maximize the F1 score on the test set for hard classification. For each assessment, mean of the estimates and 95% confidence intervals were calculated with 1,000 bootstrapping on the external validation cohort.

[0627] Conditional tabular GAN model development

[0628] The development set was used to train CTGAN model. 500 epochs (i.e. iteration over the entire development set) were performed to train the model. Parameters including genes (except for housekeeping genes), donor and recipient demographics, clinical history, biological and histological parameters, institutions, anonymized identifications, and the kidney allograft rejection diagnosis were used to train the CTGAN model. Loss values are measured to observe the model development. The CTGAN model was developed with SDV Python Package (version 1.9.0)20.

[0629] Synthetic data generation

[0630] Using the derived CTGAN model, a synthetic data was generated. The inventors generated synthetic cohort with twice the size of the development set (n>1,000) to increase the machine learning predictive power and robustness of the downstream analyses. One sample corresponds to a biopsy and no multiple biopsies were generated per patient. The synthetic cohort was sliced by the size of 25% of the development set to create eight synthetic datasets. These eight synthetic datasets are inclusive (i.e. the second dataset comprises the first and the second slices and the third dataset comprises the first, the second, and the third slices, and so on), allowing for a progressive accumulation of data. The eight synthetic datasets were attached to the end of the development set to generate augmented datasets. Comprising the development and the eight augmented datasets, the total of nine datasets were prepared and analyzed in the study. Each synthetic and augmented dataset is specified in tables, figures, and text using the notation GanN and GanAugN, respectively, where N refers to the percentage of sample size of the development set. For example, Gan50 refers to a synthetic set, which includes 50% of the sample size of the development set (the first two slices). Likewise, GanAug50 refers to an augmented set, which Gan50 set was attached to the end of the development set. Thus, the eight synthetic datasets are from Gan25 (the size of 25% of the development set) to Gan200 (the size of 200% of the development set). Likewise, the smallest augmented dataset is denoted as GanAug25 and the largest is denoted as GanAug200.Assessment of similarities between synthetic and original data The statistical similarities between the development set and the synthetic sets were assessed using SDV Python Package to assess Kolmogorov-Smirnov (KS) statistic complement and Correlation Similarity for continuous variables and Total Variance (TV) Distance complement and Contingency Similarity and for categorical variables20-22.

[0631] The definition of each metric is as follows:

[0632] For univariate comparison, column shapes were measured for all columns in KSComplement for numerical value and TVComplement for categorical value. For bivariate comparison (column pair trends), Correlationsimilarity was used for numerical pair of values and Contingencysimilarity was used for categorical pair of values or numerical and categorical pair of values.

[0633] Due to the rounding during calculation, some scores might be little bit out of ranges.

[0634] KSComplement

[0635] This metric uses the Kolmogorov-Smirnov (KS) statistic (Massey Jr. FJ. The Kolmogorov-Smirnov Test for Goodness of Fit. Journal of the American Statistical Association 1951; 46: 68-78). It is a measure of the difference between a sample distribution and a reference or theoretical distribution. In this study, the inventors compare the synthetic cohort to the original derivation cohort. A numerical distribution is converted into its cumulative distribution function (CDF). The KS statistic is the maximum difference between the two CDFs.

[0636] KSComplement is defined as 1 - KS statistic.

[0637] This metric ignores missing values and is meant for continuous and numerical data. The score ranges from 0 (worst, the two datasets are as different as they can be) to 1 (best, the two datasets are the same).

[0638] TVComplement

[0639] This metric computes the Total Variation Distance (TVD) between the real and synthetic variables by determining the probability of each category value and using it to compare the differences in probabilities. The formula below demonstrates how the TVD statistic achieves this comparison.

[0640]

[0641] 6 fl

[0642] TVComplement score is defined as 1 - 8 R, S Here, co describes all the possible categories in a column, Q. Meanwhile, R and S refer to the real and synthetic frequencies for those categories.This metric computes the similarity of a real value variable and a synthetic This metric ignores missing values and is meant for discrete and categorical data. The score ranges from 0 (worst, the two datasets are as different as they can be) to 1 (best, the two datasets are the same).

[0643] Correlationsimilarity

[0644] This metric computes a correlation coefficient on the real and synthetic data, R and S, respectively. A and B represents a pair of variables.

[0645] . |*^4, B—^4, B |

[0646] score = 1 —J-L

[0647]

[0648] 2

[0649] The metric supports both the Pearson correlation coefficient and the Spearman’s rank correlation coefficient. The score ranges from 0.5 (worst, the two correlations are as different as they can be) to 1 (best, the two correlations are the same).

[0650] ContingencySimilarity

[0651] This metric computes the similarity of a pair of categorical columns between the real and synthetic datasets.

[0652] The formula below summarizes the scoring process, where a and ft describe all the possible categories in variable A and B, respectively and R and S refer to the real and synthetic frequencies for those categories, respectively.

[0653] score = 1 - ½ |Sα,β- Rα,β|

[0654]

[0655] aEA PEB

[0656] The score ranges from 0 (worst, the contingency table is as different as can be) to 1 (best, the contingency table is exactly the same between the real and synthetic data).

[0657] Development and validation of augmented data-based molecular classifiers With the eight augmented datasets, they inventors followed the same machine learning pipeline as the original data-based molecular classifiers: 1) feature importance analysis, 2) learning curves generation, 3) machine learning developments for the eight augmented datasets, and 4) validation of the eight molecular classifiers’ performances.

[0658] Software and packages

[0659] Descriptive analyses and machine learning classifier analyses were conducted using R (version 4.3.1, R Foundation for Statistical Computing) and RStudio (version 2023.9.0.463). Synthetic data were generated using Python (version 3.10) and JupyterLab (version 4.0.9). R packages used for descriptive, data, and machine learning analyses were:tidyverse (version 2.0.0), ggplot2 (version 3.4.3), forcats (version 1.0.0), Boruta (version 8.0.0), yardstick (version 1.2.0), rstatix (version 0.7.2), pROC (version 1.18.4), stringr (version 1.5.0), caret (version 6.0-94), caretEnsemble (version 2.0.3), randomForest (version 4.7-1.1), rsample (version 1.2.0), compareGroups (version 4.7.1), viridis (version 0.6.4), foreach (version 1.5.2), doParallel (version 1.0.17), CGPfunctions (version 0.6.3), and cutpointr (version 1.1.2). Python packages used for synthetic data generation were: numpy (version 1.26.2), torch (version 2.1.1), sdv (version 1.9.0), and pandas (version 2.1.3).

[0660] Example 2: Results

[0661] Baseline characteristics of the derivation cohort

[0662] The baseline characteristics of the derivation cohort are in Table 1.

[0663] Table 1. Baseline characteristics of the derivation cohort

[0664] **Internal **External **N = **Derivation** **Validation** Validation** Validation** **Characteristic** 868** N = 441 N = 427 N = 186 N = 241 Age [years, mean(sd)] 49 (16) 50 (14) 48 (18) 49 (15) 47 (20) Gender [n(%)]

[0665] F 337 (40%) 177 (40%) 160 (39%) 73 (39%) 87 (40%) M 510 (60%) 264 (60%) 246 (61%) 113 (61%) 133 (60%) Time to biopsy [months,

[0666] mean(sd)] 27 (47) 21 (39) 34 (53) 18 (31) 47 (63) Biopsy type [n(%)]

[0667] For cause 436 (51%) 192 (45%) 244 (58%) 88 (48%) 156 (65%) Protocol 419 (49%) 239 (55%) 180 (42%) 95 (52%) 85 (35%) Creatininemia at biopsy

[0668] [micromol / L, mean(sd)] 178 (138) 177 (146) 180 (128) 182 (130) 179 (126) DSA at biopsy [n(%]

[0669] no 425 (49%) 259 (59%) 166 (39%) 109 (59%) 57 (24%) unknown 129 (15%) 19 (4.3%) 110 (26%) 11 (5.9%) 99 (41%) yes 314 (36%) 163 (37%) 151 (35%) 66 (35%) 85 (35%) Final Diagnosis [n(%)]

[0670] AMR 295 (34%) 144 (33%) 151 (35%) 61 (33%) 90 (37%) nBKv 17 (2.0%) 12 (2.7%) 5 (1.2%) 5 (2.7%) 0 (0%) Normal 204 (24%) 84 (19%) 120 (28%) 35 (19%) 85 (35%) Others 172 (20%) 104 (24%) 68 (16%) 44 (24%) 24 (10.0%) TCMR 180 (21%) 97 (22%) 83 (19%) 41 (22%) 42 (17%)N refers to the number of patients with available data

[0671] Baseline characteristics of the synthetic sets and statistical similarities The development set was used to develop a CTGAN model to create synthetic sets. The details of the hyperparameters results are as follows: -'enforce_min_max_values': True; 'enforce_rounding': True; 'locales': None; 'embedding_dim': 128; 'generator_dim': (256, 256); 'discriminator_dim': (256, 256); 'generator_lr': 0.0002; 'generator_decay': 1e-06; 'discriminator_lr': 0.0002; 'discriminator_decay': 1e-06; 'batch_size': 500; 'discriminator_steps': 1; 'log_frequency': True; 'verbose': True; 'epochs': 500; 'pac': 10; 'cuda': True.

[0672] Eight synthetic sets were generated, sample sizes of 25%, 50%, 75%, 100%, 125%, 150%, 175%, and 200% of the development set. Overall, baseline characteristics show stable values across the synthetic sets.

[0673] Between the entire synthetic set (i.e. Gan200) and the original development set, the inventors assessed univariate and bivariate similarity scores. Overall, the two datasets show high similarity.

[0674] Original data-based molecular classifiers for kidney allograft rejection

[0675] The molecular classifiers based on the development set with selected genes were assessed on the external validation cohort. For TCMR, the molecular classifier showed discrimination performance with ROCAUC of 0.864 (0.813 - 0.906), PRAUC of 0.701 (0.597 - 0.785), and Brier Score of 0.122 (0.107 - 0.138).

[0676] Feature importance analysis

[0677] For TCMR, the genes associated with T cell activation and function considered important in all nine original development and augmented sets were: CD84, LAIR1, CD96, TM4SF1, CD3E, CD27, CD247, CD45R0, HLADMB, IL16, MS4A6A, SELL, MAPK3, CD5, IKZF1, CTLA4, ZAP70, INPP5D, HLADMA, LEF1, ALOX5, STAT5A, PLAT, KDR, and GZMK. Genes related to immune cell recruitment and chemotaxis (e.g., CCR2), immune cell interactions (e.g., CD40L), inflammatory and immune response (e.g., IL1B, IL7), and cell adhesion and migration (e.g., ITGA4, ITGB2) were only considered important features in the augmented datasets (Figure 11).Learning curves analysis

[0678] With the selected features by Boruta algorithm, learning curves were generated with random forest algorithm. Incremental forward stepwise feature selection was performed and assessed on the test set (Figure 12).

[0679] For TCMR, the learning curve on the development set shows steep followed by stable increases in performance in both PRAUC and ROCAUC. At the performance reaching 0.9 quantile of PRAUC, numbers of features selected for the final molecular classifier derivation were 63 for the development set.

[0680] For TCMR, GanAug datasets show similar pattern in performance in both PRAUC and ROCAUC All learning curves show stable performance. At the performance reaching 0.9 quantile of PRAUC, numbers of features selected for the final molecular classifier derivation are detailed in Supplementary Table below.Supplementary Table 8. Selected genes to develop TCMR molecular classifiers by dataset

[0681] GanAug25 GanAug50 GanAug75 GanAug100 GanAug125 GanAug150 GanAug175 GanAug200 LCK CD96 PSMB8 PSMB8 PSMB8 CD96 CD96 BATF CD27 2 CD84 SELL CD96 IL16 IL16 IL16 BATF CD84 ICAM2 3 PPM1F IKZF1 CDH13 CD27 CD96 GZMK IL16 CD96 IKZF1 4 LAIR1 CD8A CD8A SELL GZMK C1QA CD27 CD27 CD84 5 CD3D CDH13 SELL CD96 HLAB HLAB CD84 IKZF1 ADGRL4 6 SPRY4 PTPRC IFNAR2 AOAH CD27 PSMB8 PSMB8 PDPN DNMT1 7 CD96 CD5 CTLA4 HLAB AOAH CD84 IKZF1 PSMB8 ST8SIA4 8 IGF1R IFNAR2 GBP5 TM4SF1 SELL PDPN ST8SIA4 NPDC1 CD96 9 TM4SF1 MS4A6A IL18RAP CXCL9 BATF AOAH C1QA ST8SIA4 CD45R0 10 CD3E LEF1 BMP6 CTLA4 GBP5 SELL GZMK ICAM2 PDPN 11 CD27 CTLA4 TM4SF1 HLADMA PDPN BATF ICAM2 CD247 CD247 12 CD247 CD45R0 GZMK GBP5 CTLA4 CD28 STAT5A IL16 IL1B 13 CD45R0 GBP5 CD3E PLAT CD3E PLAT AOAH MS4A6A BATF 14 BMP6 BCL2L1 MS4A6A ALOX5 MAPK3 CD45R0 CD247 DNMT1 IL16 15 HLADMB FGD2 CD163 PECAM1 STAT4 CTLA4 PDPN TM4SF1 TM4SF1 16 LTB SH2D1A HLAB CD3E IKZF1 ICAM2 TGFBR1 MMP12 PPM1F CD 17 LAG3 GZMK HLADMB PLAAT4 BCL2A1 ST8SIA4 MMP12 GZMK CD209 18 CXCR4 CD3E PSMB9 MS4A6A CDH13 CD27 CXCL13 STAT5A FCGR2A 19 IL16 BMP6 CD84 LAG3 PLAAT4 PLAAT4 MS4A6A AOAH MMP12 20 CD28 PPM1F MAPK3 CDH13 PECAM1 CD3G HLAB SELL PLA1A 21 CD8A IGF1R ADGRL4 LAIR1 IFNAR2 PPM1F TM4SF1 CTLA4 PSMB8 22 PTPN7 IL18RAP CD5 CD3D PPM1F MAPK3 MAPK3 CD28 HLADMA 23 SOD2 ALOX5 BTLA CD163 CD28 TGFBR1 CD28 MAPK3 CD28 24 IL2RB CD3D IL16 SH2D1A ADGRL4 LAIR1 PLAT CD45R0 MAPK3 25 BCL2L1 MAPK3 CD27 CD8A CD84 NPDC1 SLA SELL 26 MS4A6A KDR CD45R0 STAT5A CD45R0 SLA STAT1 CTLA4 27 SELL STAT5A ITGB2 IFNAR2 ST8SIA4 LAIR1 CXCL13 TGFBR1 28 CD8B LAG3 GZMK HLADMA THBD PLAT 29 ADGRL4 ALOX5 CD247 MS4A6A CD45R0 HLADMA 30 MAPK3 NKG7 HLADMB CMKLR1 SELL CD3E 31 ROBO4 PLAAT4 NOD1 LAIR1 HLADMB CXCR3 32 ENG IKZF1 IL21R CXCL9 IL18RAP CD209 33 IRF8 PECAM1 MAPK3 HLADMB PLAAT4 C1QA 34 CDH13 PTPRC PSMB9 SLAMF6 STAT1 LAIR1

[0682]

[0683] 35 CD5 STAT5A CD45R0 LCK KDR FCGR2A

[0684]

[0685] Augmented data-based molecular classifiers for TCMR

[0686] Synthetic biopsies were generated and merged to the development set to augment the data, for the GanAug25, GanAug50, GanAug75, GanAug100, GanAug125, GanAug150, GanAug175, and GanAug200 datasets, respectively. The performances of the molecular classifiers based on the augment data sets were assessed on the external validation cohort.

[0687] For TCMR, the molecular classifiers showed discrimination performance with ROCAUCs of 0.840 (0.784 - 0.887), 0.860 (0.807 - 0.904), 0.840 (0.781 - 0.886), 0.841 (0.786 - 0.886), 0.858 (0.808 - 0.901), 0.856 (0.803 - 0.900), 0.832 (0.778 - 0.878), and 0.841 (0.782 - 0.891), and PRAUCs of 0.684 (0.582 - 0.770), 0.704 (0.597 - 0.788), 0.641 (0.519 - 0.750), 0.675 (0.569 - 0.766), 0.690 (0.580 - 0.780), 0.693 (0.573 - 0.793), 0.646(0.531 - 0.742), and 0.658 (0.535 - 0.757) for the GanAug25, GanAug50, GanAug75, GanAug100, GanAug125, GanAug150, GanAug175, and GanAug200 sets, respectively. The classifiers showed overall fit with Brier Scores of 0.125 (0.107 - 0.143), 0.114 (0.096 -0.133), 0.122 (0.104- 0.142), 0.124 (0.105- 0.144), 0.116 (0.098 -0.135), 0.121 (0.105 -0.139), 0.134 (0.115 - 0.153), and 0.129 (0.112 - 0.148) in the same order.

[0688] The calibration plots are available on Figure 13. Cut-offs were calibrated to maximize F1 score (harmonic mean of precision and recall) for hard classification (T cell-mediated rejection: Model based on Development set: 0.552; Model based on GanAug25 set: 0.536; Model based on GanAug50 set: 0.452; Model based on GanAug75 set: 0.550; Model based on GanAug100 set: 0.452; Model based on GanAug125 set: 0.484; Model based on GanAug150 set: 0.536; Model based on GanAug175 set: 0.498; Model based on GanAug200 set: 0.512).

[0689] Table 2. Performance in multimetrics in the external validation cohort

[0690] The table illustrates detailed performance metrics in the external validation cohort. The parentheses represent 95% confidence intervals calculated with 1000 bootstraps. PRAUC, precision recall area under the curve; ROCAUC, receiver operating characteristic area under the curve; NPV negative predictive values; PPV, positive predictive values.

[0691] T cell-mediated rejection (TCMR)Model ROCAUC PRAUC taMHKM WitGfil F1 score NPV PPV

[0692]

[0693] . SwaafttiM Brier

[0694]

[0695] fyaauaa^

[0696] 0.864 0.701 0.839 0.700 0.550 0.860 0.711 0.452 0.948 0.122 Development (0.813- (0.597- (0.802- (0.639- (0.432- (0.821- (0.577- (0.341- (0.920- (0.107-

[0697]

[0698] 0.906) 0785) 0878) 0.759) 0.652) 0.898) 0.837) 0.567) 0974) 0138)

[0699] 0.840 0.684 0.845 0.718 0.580 0.868 0.714 0.491 0.945 0.125 GanAug25 (0.784- (0.582- (0.808- (0.659- (0.472- (0.830- (0.585- (0.382- (0.915- (0.107- 0.887) 0770) 0883) 0774) 0676) 0.907) 0843) 0.600) 0973) 0143) 0.860 0.704 0.833 0.753 0.614 0.891 0.622 0.611 0.896 0.114 GanAug50 (0.807- (0.597- (0.790- (0.691- (0.516- (0.852- (0.507- (0.500- (0.857- (0.096- 0904) 0788) 0875) 0810) 0702) 0.926) 0732) 0.718) 0934) 0133) 0.840 0.641 0.842 0.712 0.569 0.866 0.709 0.478 0.945 0.122 GanAug75 (0.781- (0.519- (0.805- (0.654- (0.464- (0.828- (0.580- (0.366- (0.916- (0.104- 0.886) 0.750) 0.880) 0.770) 0.667) 0.906) 0.828) 0.595) 0.971) 0.142) 0.841 0675 0.807 0.742 0585 0.891 0.553 0.625 0858 0124 GanAug100 (0.786- (0.569- (0.761 - (0.682- (0.493- (0851 - (0.449- (0.514 - (0.817- (0.105- 0.886) 0766) 0848) 0797) 0667) 0.926) 0655) 0.729) 0900) 0144) 0.858 0.690 0.836 0.741 0.603 0.883 0.643 0.571 0.911 0.116 GanAug125 (0.808- (0.500- (0.796- (0.681- (0.504- (0.844- (0.525- (0.455- (0.874- (0.098- 0901) 0780) 0878) 0800) 0693) 0.918) 0754) 0.682) 0944) 0135) 0.856 0.693 0.836 0.741 0.603 0.883 0.642 0.572 0.911 0.121 GanAug150 (0.803- (0.573- (0.793- (0.682- (0.504- (0.846- (0.525- (0.458- (0.876- (0.105- 0.900) 0793) 0878) 0.799) 0.691) 0.919) 0.750) 0.676) 0944) 0139) 0.832 0.646 0.810 0.749 0.595 0.894 0.559 0.639 0.859 0.134 GanAug175 (0.778- (0.531- (0.767- (0.687- (0.506- (0.854- (0.451- (0.524- (0.819- (0.115- 0.870) 0742) 0851) 0.805) 0.680) 0.930) 0.659) 0.740) 0899) 0153) 0.841 0.658 0.825 0.738 0.592 0.804 0.604 0.583 0.892 0.129 GanAug200 (0.782- (0.535- (0.781- (0.678- (0.493- (0.844- (0.486- (0.466- (0.852- (0.112-

[0700]

[0701] 0.891) 0.757) 0.866) 0.795) 0.676) 0.920) 0.714) 0.690) 0.929) 0.148)DISCUSSION

[0702] In this international multicenter cohort study, the inventors developed and validated machine learning based molecular classifiers to predict T cell-mediated rejection. They showed that augmenting the original dataset with generative synthetic transcriptomic data can increase the performances of these molecular classifiers and capture additional biologically relevant predictive genes.

[0703] Original data-based molecular classifiers showed steep increase followed by stable performance during the learning curve analysis. In the external validation, the original data-based molecular classifiers showed good performance in multiple metrics including ROCAUC and PRAUC in the external validation cohort for predicting TCMR. Boruta algorithm captured TCMR associated genes as highly important features.

[0704] Molecular classifiers based on the synthetic data augmentation showed stable performance increment during the learning curve analysis. In the external validation, the derived molecular classifiers showed increased performance in the synthetically augmented data in TCMR. In other metrics, the molecular classifiers showed similar performance increase trend in the augmented datasets. Calibration plots showed comparable results across the development and augmented datasets.

[0705] The inventors found that the gene importance varies through the different sets of data by feature selection / importance analysis and learning curve generation. Compared to the original development set, the synthetically augmented sets allowed capturing genes biologically related to B cell activation, T cell activation and macrophage activation, NK-specific activation, and cellular injury and adaptive immune. This newly captured trend in the machine-generated augmented datasets enhances and expands upon previous studies, revealing additional biologically relevant genes for diagnosing kidney rejection23.

[0706] Incorporating GAN-based data augmentation into clinical practice alongside standard monitoring factors presents an opportunity to enhance cost-effectiveness and augment the breadth of data available for analysis24. By satisfying the growing demand for data and contributing to the development of more comprehensive predictive classifiers for TCMR diagnosis, this approach can lead to significant cost savings while simultaneously safeguarding the privacy of kidney allograft recipients. These advancements offer a promising avenue for improving the management of renal transplant patients by refining the clinical indications for biopsies and optimizing overall care protocols.

[0707] This study presents multiple strengths. First, the study cohort comprises a comprehensive set of multimodal datatypes including clinical, biological, histological, donor, recipient, and transplant parameters, in addition to transcriptomic data. This extensive dataset allows for a more holistic analysis and better understanding of kidney allograftrejection. Second, the inventors used multidisciplinary study methods. The study incorporates heuristic selection by kidney transplant experts from the Banff Human Organ Transplant (B-HOT) panel, ensuring the inclusion of the most relevant genes and markers3. Additionally, it integrates specialized expertise from nephrologists and pathologists for multimodal kidney rejection diagnosis. Advanced deep-learning techniques are employed to augment data, aiming to enhance the robustness and accuracy of the classifiers’ performance. Furthermore, multiple machine learning methods are utilized for feature selection, learning curve analysis, hyperparameter optimization, and classifiers’ generation. Finally, external validation cohorts are included to strengthen the reliability and generalizability of the study's findings. Validating classifiers on independent datasets ensures robust performance across different clinical settings.

[0708] In conclusion, the inventors derived and validated machine learning-based molecular classifiers for kidney allograft rejection and enhanced their performance with deep learning-driven synthetic data augmentation. Using synthetic transcriptomic data, they captured additional predictive genes and showed potential for reducing the necessary number of transcriptomic sequencings and costs of time-consuming procedures.References

[0709] 1. Solez, K. et al. International standardization of criteria for the histologic diagnosis of renal allograft rejection: The Banff working classification of kidney transplant pathology. Kidney International 44, 411-422 (1993).

[0710] 2. Schinstock, C. A. et al. Banff survey on antibody-mediated rejection clinical practices in kidney transplantation: Diagnostic misinterpretation has potential therapeutic implications.

[0711] 3. Mengel, M. et al. Banff 2019 Meeting Report: Molecular diagnostics in solid organ transplantation-Consensus for the Banff Human Organ Transplant (B-HOT) gene panel and open source multicenter validation.

[0712] 4. Halloran, P. F., Famulski, K. S. & Reeve, J. Molecular assessment of disease states in kidney transplant biopsy samples. Nature Reviews Nephrology 12, 534-548 (2016). 5. Loupy, A. et al. Molecular Microscope Strategy to Improve Risk...: Journal of the American Society of Nephrology.

[0713] 6. Frid-Adar, M. et al. GAN-based synthetic medical image augmentation for increased CNN performance in liver lesion classification. Neurocomputing 321, 321-331 (2018).

[0714] 7. Ju, L. et al. Leveraging Regular Fundus Images for Training UWF Fundus Diagnosis Models via Adversarial Learning and Pseudo-Labeling. IEEE Transactions on Medical Imaging 40, 2911-2925 (2021).

[0715] 8. Ktena, I. et al. Generative models improve fairness of medical classifiers under distribution shifts. Nature Medicine 30, 1166-1173 (2024).

[0716] 9. Goodfellow, I. et al. Generative Adversarial Nets, in Advances in Neural Information Processing Systems vol. 27 (Curran Associates, Inc., 2014).

[0717] 10. Xu, L. & Veeramachaneni, K. Synthesizing Tabular Data using Generative Adversarial Networks. Preprint at https: / / doi.org / 10.48550 / arXiv.1811.11264 (2018). 11. Xu, L., Skoularidou, M., Cuesta-Infante, A. & Veeramachaneni, K. Modeling Tabular data using Conditional GAN. in Advances in Neural Information Processing Systems vol.

[0718] 32 (Curran Associates, Inc., 2019).

[0719] 12. Elm, E. von et al. The Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) statement: guidelines for reporting observational studies. 13. Heidari, S., Babor, T. F., De Castro, P., Tort, S. & Curno, M. Sex and Gender Equity in Research: rationale for the SAGER guidelines and recommended use. Research Integrity and Peer Review 1, 2 (2016).

[0720] 14. Yoo, D. et al. An automated histological classification system for precision diagnostics of kidney allografts. Nature Medicine 29, 1211-1220 (2023).

[0721] 15. NanoString Technologies. nSolverTM 4.0 Analysis Software. 2018;5-98.. Loupy, A. et al. The Banff 2019 Kidney Meeting Report (I): Updates on and clarification of criteria for T cell- and antibody-mediated rejection.

[0722] . Collins, G. S. et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. (2024) doi: 10.1136 / bmj-2023-078378.

[0723] . Kursa, M. B. & Rudnicki, W. R. Feature Selection with the Boruta Package. Journal of Statistical Software 36, 1-13 (2010).

[0724] . Breiman, L. Random Forests. Machine Learning 45, 5-32 (2001).

[0725] . Patki, N., Wedge, R. & Veeramachaneni, K. The Synthetic Data Vault, in 2016 IEEE International Conference on Data Science and Advanced Analytics (DSAA) 399-410 (2016). doi: 10.1109 / DSAA.2016.49.

[0726] . Massey Jr., F. J. The Kolmogorov-Smirnov Test for Goodness of Fit. Journal of the American Statistical Association 46, 68-78 (1951).

[0727] . Gibbs, A. L. & Su, F. E. On Choosing and Bounding Probability Metrics. International Statistical Review / Revue Internationale de Statistique 70, 419-435 (2002).

[0728] . Halloran, P. F. et al. Real Time Central Assessment of Kidney Transplant Indication Biopsies by Microarrays: The INTERCOMEX Study.

[0729] . Motamed, S., Rogalla, P. & Khalvati, F. Data augmentation using Generative Adversarial Networks (GANs) for GAN-based detection of Pneumonia and COVID-19 in chest X-ray images.

[0730] . K, S. et al. Clinical validation and reproducibility of the Banff schema for renal allograft pathology. PubMed.

[0731] . Schinstock, C. A. et al. Banff survey on antibody-mediated rejection clinical practices in kidney transplantation: Diagnostic misinterpretation has potential therapeutic implications.

Claims

CLAIMS1.- Method for obtaining at least one molecular classifying function (FMC) intended to be used to predict a biological state of a subject, a molecular classifying function (FMC) being adapted to take as input several gene expression level parameters of the subject and provide as output the biological state of the subject, a gene expression level parameter being a parameter representative of the expression level of said gene for said subject, the method for obtaining being computer-implemented, the method comprising the steps of: - obtaining an initial database (IDB), the initial database (IDB) comprising several elements, each element corresponding to a respective subject and providing, for said respective subject, several gene expression level parameters (GELP), other parameters relative to the subject and the biological state (S) of the subject,- training a conditional tabular generative adversarial network on the initial database (IDB) to output elements fulfilling at least one similarity condition with the elements of the initial database (IDB), to obtain a trained conditional tabular generative adversarial network, - using the trained conditional tabular generative adversarial network, to obtain synthetic elements, the set of the initial database (IDB) and the synthetic elements forming an augmented database (ADB),- determining gene expression level parameters (GLEP) of the augmented database (ADB) impacting a prediction of the biological state (S) of the subject by a molecular classifying function, to obtain determined gene expression level parameters (GLEPDET),- selecting gene expression level parameters (GLEP) among the determined gene expression level parameters (GLEPDET), to obtain a set of selected gene expression level parameters (GLEPSEL), the step of selecting comprising iteratively for several determined gene expression level parameters (GLEPDET):- analyzing the impact of adding said determined gene expression level parameters (GLEPDET) in a learning process of a molecular classifying function adapted to take as input at least one already selected gene expression level parameter (GLEPSEL) and said determined gene expression level parameters (GLEPDET) and to provide as output the biological state (S) of the subject, and- adding the determined gene expression level parameter (GLEPDET) in case a selection criterion is fulfilled or discarding said determined gene expression level parameter (GLEPDET) in case said selection criterion is not fulfilled,- generating molecular classifying functions taking as input the set of selected gene expression level parameters (GLEPSEL) and providing as output the biological state (S) of the subject, to obtain molecular classifying function candidates (FMCCAND), and- choosing at least one molecular classifying function (FMC) among the molecular classifying function candidates (FMCCAND) based on a performance criterion, to obtain at least one molecular classifying function (FMC) intended to be used to predict a biological state of a subject.2.- Method for obtaining according to claim 1, wherein the step of determining is carried out by applying a Boruta algorithm on the augmented database (ADB).3.- Method for obtaining according to claim 1 or 2, wherein the number of determined gene expression level parameters (GLEPDET) is comprised between 10 and 50, preferably comprised between 10 and 30.4.- Method for obtaining according to any one of claims 1 to 3, wherein a ratio between a number (N2) of synthetic elements and a number (N1) of elements of the initial database (IDB) is comprised between 0,1 % and 100%, preferably between 0.25% and 0.75% and advantageously equal to 50%.5.- Method for obtaining according to any one of claims 1 to 3, wherein a ratio between a number (N2) of synthetic elements and a number (N1) of elements of the initial database (IDB) is comprised between 100 % and 200%, preferably between 100% and 150% and advantageously equal to 125%.6.- Method for obtaining according to any one of the claims 1 to 5, wherein the molecular classifying function used in the step of determining and in the step of selecting is obtained by an hyperparameter optimization carried out on a same part of the initial database (IDB).7.- Method for obtaining according to any one of the claims 1 to 6, wherein the step of obtaining comprises receiving a cohort (CO) and extracting the initial database (IDB) and a validating database (VDB) from the received cohort (CO), the performance criterion used at the step of choosing being assessed on the validating database (VDB).8.- Method for obtaining according to any one of the claims 1 to 7, wherein the extracting is carried out by splitting the received cohort (CO) into two parts, one part being the initial database (IDB), the other part being the validating database (VDB), said splitting being achieved with a splitting ratio comprised between 7:3 and 9:1.9.- Method for obtaining according to any one of the claims 1 to 8, wherein the other parameters relative to the subject comprise at least one comorbidity parameter (COMP), at least one clinical parameter (CLIP), at least one biological parameter (BIOP) and at least one histological parameter (HISP).10.- Method for obtaining according to claim 9, wherein each parameter of the other parameters relative to the subject is chosen among a comorbidity parameter (COMP), a clinical parameter (CLIP), a biological parameter (BIOP) and an histological parameter (HISP).11.- Method for predicting a biological state (S) of a subject, the method being carrying out by a system for predicting (100), the method for predicting comprising:- providing gene expression level parameters (GELP) of the subject,- applying a molecular classifying function (FMC) obtained by a method for obtaining according to any one of the claims 1 to 10 on the provided gene expression level parameters (GELP), to output a predicted biological state (S) of the subject.12.- Method selected from the group consisting of:- a method for predicting that a subject is at risk of suffering from a disease, the method for predicting comprising at least the steps of:- carrying out the steps of a method for predicting according to claim 11 wherein the step of providing is achieved by receiving the gene expression level parameters (GELP), to obtain a predicted biological state (S) of the subject, and- predicting that the subject is at risk of suffering from the disease based on the predicted biological state (S) of the subject,- a method for diagnosing a disease to a subject, the method for diagnosing comprising at least the steps of:- carrying out the steps of a method for predicting according to claim 11 wherein the step of providing is achieved by receiving the gene expression level parameters (GELP), to obtain a predicted biological state (S) of the subject, and- diagnosing the disease based on the predicted biological state (S) of the subject, - a method for identifying a therapeutic target for preventing and / or treating a disease, the method comprising at least the steps of:- carrying out the steps of a method for predicting for a first subject, to obtain a first predicted biological state (S), wherein the first subject is suffering from the diseaseand the method for predicting is according to claim 11 wherein the step of providing is achieved by receiving the gene expression level parameters (GELP),- carrying out the steps of a method for predicting for a second subject, to obtain a second predicted biological state (S), wherein the second subject is not suffering from the disease and the method for predicting is according to claim 11 wherein the providing is achieved by receiving the gene expression level parameters (GELP), and- selecting a therapeutic target based on the comparison of the first and second predicted biological states (S),- a method for identifying a biomarker for a disease, the biomarker being a diagnosis biomarker of the disease, a susceptibility biomarker of the disease, a prognostic biomarker of the disease or a predictive biomarker in response to the treatment of the disease, the method comprising at least the steps of:- carrying out the steps of a method for predicting for a first subject, to obtain a first predicted biological state (S), wherein the first subject is suffering from the disease and the method for predicting is according to claim 11 wherein the step of providing is achieved by receiving the gene expression level parameters (GELP),- carrying out the steps of a method for predicting for a second subject, to obtain a second predicted biological state (S), wherein the second subject is not suffering from the disease and the method for predicting is according to a claim 11 wherein the providing is achieved by receiving the gene expression level parameters (GELP), and- selecting a biomarker target based on the comparison of the first and second predicted biological states (S),- a method for screening a compound useful as a medicament, the compound having an effect on a known therapeutical target for preventing and / or treating a disease, the method comprising at least the steps of:- carrying out the steps of a method for predicting for a first subject, to obtain a first predicted biological state (S), wherein the first subject is suffering from the disease and has received the compound and the method for predicting is according to claim 11 wherein the step of providing is achieved by receiving the gene expression level parameters (GELP),- carrying out the steps of a method for predicting for a second subject, to obtain a second predicted biological state (S), wherein the second subject is suffering from the disease and has not received the compound and the method for predicting isaccording to claim 11 wherein the step of providing is achieved by receiving the gene expression level parameters (GELP), and- selecting a biomarker target based on the comparison of the first and second predicted biological states (S), and- a method for monitoring patients enrolled in a clinical trial to provide a quantitative measure for the therapeutic efficacy of the therapy which is subject to the clinical trial by carrying out the steps of a method for predicting for said patients, the method for predicting being according to claim 11 wherein the step of providing is achieved by receiving the gene expression level parameters (GELP).13.- Electronic device (50) for obtaining at least one molecular classifying function (FMC) intended to be used to predict a biological state of a subject, a molecular classifying function (FMC) being adapted to take as input parameters representative of the expression level of genes of the subject and provide as output the biological state of the subject, the electronic device (50) comprising:- an obtaining module (52) configured to obtain an initial database (IDB), the initial database (IDB) comprising several elements, each element corresponding to a respective subject and providing, for said respective subject, several gene expression level parameters (GELP), other parameters relative to the subject and the biological state (S) of the subject,- a training module (54) configured to train a conditional tabular generative adversarial network on the initial database (IDB) to output elements fulfilling at least one similarity condition with the elements of the initial database (IDB), to obtain a trained conditional tabular generative adversarial network,- a using module (56) configured to use the trained conditional tabular generative adversarial network, to obtain synthetic elements, the set of the initial database (IDB) and the synthetic elements forming an augmented database (ADB),- a determining module (58) configured to determine gene expression level parameters (GLEP) of the augmented database (ADB) impacting a prediction of the biological state of the subject by a molecular classifying function, to obtain determined gene expression level parameters (GLEPDET),- a selecting module (60) configured to select gene expression level parameters (GLEP) among the determined gene expression level parameters (GLEPDET), to obtain a set of selected gene expression level parameters (GLEPSEL), the selecting module (60) selecting said gene expression level parameters (GLEPSEL) by iteratively for each determined gene expression level parameter (GLEPDET):- analyzing the impact of adding said determined gene expression level parameters (GLEPDET) in a learning process of a molecular classifying function adapted to take as input at least one already selected gene expression level parameter (GLEPSEL) and said determined gene expression level parameters (GLEPDET) and to provide as output the biological state (S) of the subject, and- adding the determined gene expression level parameter (GLEPDET) in case a selection criterion is fulfilled or discarding said determined gene expression level parameter (GLEPDET) in case said selection criterion is not fulfilled,- a generating module (62) configured to generate molecular classifying functions (FMCCAND) taking as input the set of selected gene expression level parameters (GLEPSEL) and providing as output the biological state (S) of the subject, to obtain molecular classifying function candidates (FMCCAND), and- a choosing module (64) configured to choose at least one molecular classifying function (FMC) among the molecular classifying function candidates (FMCCAND) based on a performance criterion, to obtain at least one molecular classifying function (FMC) intended to be used to predict a biological state of a subject.14.- System for predicting (100) a biological state (S) of a subject, the system for predicting (100) comprising:- a providing module (102), the providing module (102) being configured to provide gene expression level parameters (GELP) of the subject, and- an applying module (104), the applying module (104) being configured to apply a molecular classifying function (FMC) on the gene expression level parameters (GELP), which the providing module (102) is configured to provide, to output a predicted biological state (S) of the subject.15.- Computer program product (30) comprising computer program instructions, the computer program instructions being loadable into a data-processing unit (38) and adapted to cause execution of a method according to any one of the claims 1 to 12 when run by the data-processing unit (38).16.- Computer-readable medium (48) comprising computer program instructions which, when executed by a data-processing unit (38), cause execution of a method according to any one of the claims 1 to 12.