Method and device for patient response prediction
The device uses patient-derived organoids and trained models to predict drug response by integrating patient-specific clinical and molecular data, addressing the limitations of current methods by enhancing prediction accuracy through personalized tumor modeling.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-05
- Publication Date
- 2026-03-12
Smart Images

Figure EP2025075355_12032026_PF_FP_ABST
Abstract
Description
METHOD AND DEVICE FOR PATIENT RESPONSE PREDICTIONTECHNICAL FIELD
[0001] The present disclosure relates to the field of computer programs and systems, and more specifically to the field of preclinical testing in drug development and research. Yet more specifically, the present disclosure relates to a method, device and program for predicting a response of a patient or a population of patients to an agent.BACKGROUND
[0002] Drug’s efficacy assessment has developed significantly in the context of personalized medicine, for example precision oncology. There are some existing solutions to predict individual patient responses to a given drug. However, the use of these solutions has not been extended to large cohorts of patient avatars to guide decisionmaking in the drug development context.
[0003] Predicting single patient outcomes in the clinical tests has been centered around two complementary approaches that have been independently applied. The first approach, i.e., the so-called omics-based approach, uses models based on data composed of one or more of transcriptomic, genomic, or clinical data sources to anticipate patient outcomes / responses, e.g., in cancer therapies. On the other hand, the second approach, i.e., the so-called functional approach uses the response of patient-derived organoids, patient-derived xenografts, or patient-derived cell lines to predict individual patient outcomes / responses .
[0004] Examples of known prior art of the functional approach for individual tests may be found in the following articles:Boileve, Alice, et al. "Organoids for functional precision medicine in advanced pancreatic cancer." Gastroenterology (2024).Cartry, Jerome, et al. "Implementing patient derived organoids in functional precision medicine for patients with advanced colorectal cancer." Journal of Experimental & Clinical Cancer Research 42.1 (2023): 1-17.Tan, Tao, et al. "Unified framework for patient-derived, tumor-organoid-based predictive testing of standard-of-care therapies in metastatic colorectal cancer." Cell Reports Medicine 4.12 (2023).Wensink, G. Emerens, et al. "Patient-derived organoids as a predictive biomarker for treatment response in cancer patients." NPJ precision oncology 5.1 (2021): 30.
[0005] Examples of known prior art of the genomics-based approaches may be found in the following articles:Sinha, Sanju, et al. "PERCEPTION predicts patient response and resistance to treatment using single-cell transcriptomics of their tumors." Nature Cancer (2024): 1-15.Dinstag, Gal, et al. "Clinically oriented prediction of patient response to targeted and immunotherapies from the tumor transcriptome." Med 4.1 (2023): 15-30.
[0006] These examples of the omics-based approach have been directed to targeted therapies, for example for predicting the efficacy of immunotherapies, and for the efficacy of chemotherapies (e.g., transcriptomic markers of gemcitabine sensitivity in pancreatic ductal adenocarcinoma (PDAC)).
[0007] Despite the current developments, it is still a challenge to predict the response rate of a patient or a group of patients. This is particularly the case for multi-variate response predictions, i.e., when the prediction is realized using a plurality of variables to forecast possible outcomes. Current methods are each optimised for one or a set of drugs and are generally tumor-type specific. Furthermore, current approaches do not generalise to drugs in development, i.e., to predict responses to drugs that have been never given to patients in the clinic.
[0008] Within this context, there is still a need for an improved solution for predicting a patient response to an agent.SUMMARY
[0009] The present invention relates to a device for predicting a patient’ s response to an agent by using at least one organoid exposed to said agent, said at least one organoidhaving been grown from at least one type of cells previously extracted from the patient, the device comprising: at least one input configured to receive at least: o an avatar of the patient and respective data of organoid response for said patient, said avatar comprising biological data and clinical data of said patient; o a first ensemble of one or more trained learning models each configured to receive as input, with respect to said patient, at least one patient feature and at least one organoid response feature, and to provide a first output vector; at least one processor configured to: o defining said at least one patient feature using said avatar and defining said at least one organoid response feature using at least said respective data of organoid response; o calculate the first output vector by providing said at least one patient feature and at least one organoid response feature as input to said first ensemble of one or more trained learning models; o predict at least a Progression-Free Survival (PFS) and / or an Overall Survival (OS) for the patient using the first output vector; and o obtain a prediction of the patient response to the agent using said obtained PFS and / or OS.
[0010] According to one embodiment, the device for predicting a patient response to an agent, the device comprising: at least one input configured to receive at least: o an avatar of a patient and respective data of organoid response for said patient, said avatar comprising biological data and clinical (i.e., digital data) of said patient, said avatar and said respective data of organoid response defining at least one patient feature and at least one organoid response feature respectively; o a first ensemble of one or more trained learning models each configured to receive as input, with respect to said patient, one or more of said at least one patient feature and one or more of said at least one organoid response feature, and to provide a first output vector; at least one processor configured to:o calculate the first output vector by providing said at least one patient feature and at least one organoid response feature as input to said first ensemble of one or more trained learning models; o obtain at least a Progression-Free Survival (PFS) and an Overall Survival (OS) for the patient using the first output vector; and o obtain a prediction of the patient response to the agent using said obtained PFS and / or OS.
[0011] The device may further comprise one or more of the following: According to other advantageous aspects of the invention, the device comprises one or more of the features described in the following embodiments, taken alone or in any possible combination.
[0012] According to one embodiment, each of the trained learning models of the first ensemble has been previously trained using a training dataset comprising entries for a plurality of subjects, each entry corresponding to a subject and comprising for said subject an avatar of the subject; and respective data of organoid response.
[0013] According to embodiment, each of the trained learning models of the first ensemble has been previously trained using a training dataset comprising entries corresponding to a plurality of subjects, wherein the organoids of the subjects are obtained from the same type of cells as those of the patient; each entry corresponding to a subject and comprising for said subject: an avatar of the subject; respective data of organoid response; and at least one measured Progression-Free Survival (PFS) and / or a measured Overall Survival (OS) of the subject; wherein each of the trained learning models is configured to provide as output a first output vector representative of a Progression-Free Survival (PFS) and / or a measured Overall Survival (OS); wherein each of the trained learning models is obtained by iteratively providing, to a corresponding initial model, the entries of the training dataset one by one untilconvergence, with the progression-free survival (PFS) and / or overall survival (OS) associated with each entry being used as ground truth at each iteration.
[0014] An advantage of training the model with avatars comprising both patient-specific clinical data and the corresponding organoid response data is that the training relies on matched datasets that directly reflect the biological and clinical reality of each patient / subject. Such a configuration preserves the intrinsic link between the patient’s clinical features and the molecular environment of the tumor, including genomic, epigenetic, transcriptomic, and other omics-level influences that shape treatment response. By capturing these patient-specific characteristics, at the inference, each trained model is able to generate predictions that are more accurate and better tailored to the individual patient, thereby improving the reliability of patient’s response prediction. This is in contrast to approaches relying on non-specific or generic cell information (e.g., organoid from non-patient / subject derived cell lineage, such as immortalized cells cancer), which fail to incorporate the fine-grained molecular and clinical context of the patient and thus provide less discriminative and less personalized predictions.
[0015] According to embodiment, the plurality of subjects associated to the training dataset is affected by the same pathology as the patient.
[0016] According to one embodiment, the at least one patient feature is obtained using at least one large language model (LLM) using the clinical data of the patient.
[0017] According to one embodiment, the first ensemble of one or more learning models comprises at least one Cox model and, optionally, one or more models of the following types: regression model, XGBoost, Random Survival Forest, Ridge Cox, and Lasso Cox.
[0018] According to one embodiment, the Progression-Free Survival (PFS) and / or an Overall Survival (OS) is obtained using a pre-trained regression model taking as input the first output vector.
[0019] According to one embodiment, said at least one input is further configured to receive at least one second ensemble of one or more trained learning models each configured to receive as input, with respect to said patient, one or more of said at leastone patient feature and one or more of said at least one organoid response feature, and to provide a second output vector, said second ensemble of one or more trained learning models comprising at least one logistic regression model; and said least one processor is further configured to calculate the second output vector by providing said at least one patient feature and at least one organoid response feature as input to said second ensemble of one or more trained learning models and to obtain an Overall Response (OR) for the patient using the second output vector.
[0020] According to one embodiment, said at least one input is further configured to receive at least one third ensemble of one or more trained learning models each configured to receive as input, with respect to said patient, one or more of said at least one patient feature and one or more of said at least one organoid response feature, and to provide a third output vector, said third ensemble of one or more trained learning models comprising at least one logistic regression model; and said least one processor is further configured to calculate the third output vector by providing said at least one patient feature and at least one organoid response feature as input to said third ensemble of one or more trained learning models and to obtain a prediction of best RECIST response for the patient using the third output vector.
[0021] According to one embodiment, each defined one or more of said at least one patient feature has a statistical significance higher than a first threshold, and each defined one or more of said at least one organoid response feature has a statistical significance higher than a second threshold.
[0022] According to one embodiment, the second ensemble of one or more trained learning models further comprises at least one or more models of the following types: XGBoost, logistic regression, and random forest.
[0023] According to one embodiment, when said second ensemble of one or more trained learning models comprises two or more trained learning models, predictions obtained from each trained learning models are combined according to a voting classifier to obtain the second output vector.
[0024] According to one embodiment, the agent is selected from the group consisting of: a physical treatment (for example radiation therapy), a small molecule, a therapeutic compound, a polypeptide, an antibody or antigen-binding fragment thereof, an aptamer, a nucleic acid (for example a guide RNA), a nanoparticle, an expression vector, a virus, a cell, such as a prokaryotic cell (for example a bacteria) or an eukaryotic cell (for example an immune cell), combinations thereof, and / or compositions thereof.
[0025] According to one embodiment, the organoid response data is selected from the group consisting of a proteomic feature, a genomic feature, an epigenomic feature, a metabolic feature, a microbiome feature, an imaging feature, a histological feature, viability measurements, growth rate, or a combination thereof. Alternatively or additionally, the organoid response data may be resulted from a chemical test, for example an ATP test.
[0026] The present invention further relates to computer-implemented method for predicting a patient response to an agent, the method comprising: receiving an avatar of the patient and respective data of organoid response for said patient, said avatar comprising biological data and clinical data of said patient; receiving a first ensemble of one or more trained learning models each configured to receive as input, with respect to said patient, at least one patient feature and at least one organoid response feature, and to provide a first output vector; defining said at least one patient feature using said avatar and defining said at least one organoid response feature using at least said respective data of organoid response; calculating the first output vector by providing one or more of said at least one patient feature and one or more of at least one organoid response feature as input to said first ensemble of one or more trained learning models; predict at least a Progression-Free Survival (PFS) and / or an Overall Survival (OS) for the patient using the first output vector; andobtaining a prediction of the patient response to the agent using said obtained PFS and / or OS.
[0027] According to one embodiment, the computer-implemented method for predicting a patient response to an agent, the method comprising: receiving an avatar of the patient and respective data of organoid response for said patient, said avatar comprising biological data and clinical data (i.e., digital data) of said patient, said avatar and said respective data of organoid response defining at least one patient feature and at least one organoid response feature respectively; receiving a first ensemble of one or more trained learning models each configured to receive as input, with respect to said patient, one or more of said at least one patient feature and one or more of said at least one organoid response feature, and to provide a first output vector; calculating the first output vector by providing one or more of said at least one patient feature and one or more of at least one organoid response feature as input to said first ensemble of one or more trained learning models; obtaining at least a Progression-Free Survival (PFS) and / or an Overall Survival (OS) for the patient using the first output vector; and obtaining a prediction of the patient response to the agent using said obtained PFS and / or OS.
[0028] It is further provided a method for predicting an effect of an agent on a patient or a population of patients, the method comprising: receiving a test dataset comprising a plurality of entries, each entry corresponding to a test patient and comprising the respective avatar of the test patient and data of organoid response for said test patient; obtaining a simulated population dataset by applying the computer-implement method for predicting a patient response to an agent to the respective avatar of the test patient and organoid response data for each entry, the obtained simulated population dataset comprising at least a PFS and / or an OS for each respective test patient.
[0029] In other words, the method for predicting an effect of an agent on a population of patients comprising: receiving a test dataset comprising a plurality of entries, each entry corresponding to a test patient and comprising the respective avatar of the test patient and data of organoid response for said test patient; receiving a first ensemble of one or more trained learning models each configured to receive as input, with respect to said patient, one or more of said at least one patient feature and one or more of said at least one organoid response feature, and to provide a first output vector; obtained a simulated population dataset comprising at least a PFS and / or an OS for each respective test patient by, for each test patient of the test dataset:■ defining at least one patient feature using the avatar and defining at least one organoid response feature using at least said respective data of organoid response;■ calculate the first output vector by providing said at least one patient feature and at least one organoid response feature as input to said first ensemble of one or more trained learning models;■ obtain at least a Progression-Free Survival (PFS) and / or an Overall Survival (OS) for the patient using the first output vector; and■ obtain a prediction of the patient response to the agent using said obtained PFS and / or OS;■ add said obtained PFS and / or OS to the corresponding entry.
[0030] According to one embodiment, defining at least one patient feature using the avatar comprises: selecting, among the clinical data, a first subset of clinical data relating to the tumor cells; selecting, among the clinical data, a second subset of clinical data relating to the patient; generating a synthetic second subset of clinical data by, for each piece of clinical data comprised into said second subset, comparing it to a corresponding acceptance range of values associated to said piece of clinical data, and wheneverthe piece of clinical data is not comprised in said acceptance range of values generating a corresponding synthetic data value and recording it into the synthetic second subset of data, and otherwise recording the piece of clinical data in the synthetic second subset of data; generating synthetic clinical data for the test patient comprising the first subset of clinical data and the synthetic second subset of clinical data; defining at least one patient feature using said synthetic clinical data.
[0031] According to one embodiment, each acceptance range of values associated to a specific piece of clinical data is defined from information from at least one clinical study, notably eligibility criteria (inclusion / exclusion criteria).
[0032] An advantage of the proposed approach is that the generation of synthetic clinical data preserves, for each test patient, the unaltered information relating to the tumor, thereby maintaining the biological link with the organoids derived from the patient and exposed to the agent under examination. At the same time, the method allows the modification of patient-related information that is not intrinsically linked to the tumor, such as age, sex, or performance status, so that these variables can be adapted to meet the acceptance criteria or eligibility requirements of a clinical study. In this way, a synthetic population dataset can be constructed that remains biologically anchored to the real patient-derived organoids while being aligned with the formal constraints of the clinical trial design. Such a dataset advantageously combines biological fidelity with regulatory and methodological compliance, enabling more robust training, validation, or simulation of therapeutic outcomes.
[0033] The method for predicting an effect of an agent on a population of patients may further comprise one or more of the following: updating said simulated population dataset by applying at least one second ensemble of one or more trained learning models to the respective avatar of the data and the data of organoid response for each entry, the updated simulated population dataset further comprising an Overall Response (OR) for respective test patient, wherein said at least one second ensemble of one or more trained learning models comprises at least one logistic regression model and is configuredto receive as input, with respect to said test patient, one or more of said at least one patient feature and one or more of said at least one organoid response feature, and to provide a second output vector, and wherein the Overall Response (OR) for the test patient is obtained using the second output vector calculated by said at least one second ensemble of one or more trained learning models receiving as input said at least one patient feature and at least one organoid response feature for the test patient; and optionally determining a corrector for the OR of the updated simulated population dataset; wherein the determining the said corrector comprises computing a correction factor of the type:6 := E(p) - p = f(p,NPV, PPV) wherein:• p is an estimator of p,• E(. ) denotes an expectation,denotes a function,• 6 is said correction factor,• NPV is the negative predictive value, and• PPV is the positive predictive value, wherein: p is a ground truth value of the OR of updated population dataset, and PPV and NPV denote a given positive predictive value and a given negative predictive value, respectively of the at least one second ensemble of one or more trained machine learning models.
[0034] According to embodiment, the method for predicting an effect of an agent on a population of patients further comprises predicting an effect of an agent on a population of patients using the simulated population dataset. As the simulated population dataset comprises clinical outcome measures such as progression-free survival (PFS), overall survival (OS) best RECIST response and / or OR (overall response) for each test patient, a survival analysis may then be performed on the simulated population, for instance by constructing Kaplan-Meier curves and estimating median survival times or survival probabilities at predefined time points. The predicted effect of the agent can be quantified by deriving hazard ratios using a Cox proportional hazards model, comparing the survivaldistributions against a reference population or control treatment. In the absence of a control group, the simulated survival outcomes may be evaluated against historical benchmarks or clinical trial eligibility thresholds to determine whether the predicted PFS and / or OS are indicative of a clinically meaningful benefit. In this way, the simulated population dataset enables the estimation of drug efficacy at the population level while preserving the biological anchoring to the organoids derived from the test patients.
[0035] It is further provided a computer-implemented method for training a learning model of the at least first / second / third ensemble of one or more trained learning models discussed above; the method comprising: receiving a training dataset comprising entries, each entry corresponding to a subject and comprising for said subject: o an avatar of the subject; and o data of organoid response data; training the learning model using the training dataset.
[0036] The present invention further relates to a computer-implemented method for training one or more learning models of a first ensemble of one or more learning models to obtain the first ensemble of one or more trained learning models discussed above; the method comprising: receiving the first ensemble of one or more learning models (e.g., untrained) each configured to receive as input, with respect to a patient, one or more of the at least one patient feature and one or more of the at least one organoid response feature, and to provide as output a first output vector representative of a Progression-Free Survival (PFS) and / or a measured Overall Survival (OS); receiving a training dataset comprising entries corresponding to a plurality of subjects, said plurality of subjects being affected by the same pathology as the patient, wherein the organoids of the subjects are obtained from the same type of cells as those of the patient; each entry corresponding to a subject and comprising for said subject: o an avatar of the subject; o respective data of organoid response; ando at least one measured Progression-Free Survival (PFS) and / or a measured Overall Survival (OS) of the subject; obtaining each of the trained learning models by iteratively providing, to the corresponding learning model, the entries of the training dataset one by one until convergence, with the progression-free survival (PFS) and / or overall survival (OS) associated with each entry being used as ground truth at each iteration.
[0037] The present invention further relates to a computer-implemented method for training one or more learning models of a second ensemble of one or more learning models to obtain the second ensemble of one or more trained learning models discussed above; the method comprising: receiving the second ensemble of one or more learning models (e.g., untrained) each configured to receive as input, with respect to a patient, one or more of the at least one patient feature and one or more of the at least one organoid response feature, and to provide as output a second output vector representative of an Overall Response (OR); receiving a training dataset comprising entries corresponding to a plurality of subjects, said plurality of subjects being affected by the same pathology as the patient, wherein the organoids of the subjects are obtained from the same type of cells as those of the patient; each entry corresponding to a subject and comprising for said subject: o an avatar of the subject; o respective data of organoid response; and o at least one measured Overall Response (OR) of the subject; obtaining each of the trained learning models by iteratively providing, to the corresponding learning model, the entries of the training dataset one by one until convergence, with the Overall Response (OR) associated with each entry being used as ground truth at each iteration.
[0038] The present invention further relates to a computer-implemented method for training one or more learning model of a third ensemble of one or more learning modelsto obtain the third ensemble of one or more trained learning models discussed above; the method comprising: receiving the third ensemble of one or more learning models (e.g., untrained) each configured to receive as input, with respect to a patient, one or more of the at least one patient feature and one or more of the at least one organoid response feature, and to provide as output a third output vector representative of a best RECIST response; receiving a training dataset comprising entries corresponding to a plurality of subjects, said plurality of subjects being affected by the same pathology as the patient, wherein the organoids of the subjects are obtained from the same type of cells as those of the patient; each entry corresponding to a subject and comprising for said subject: o an avatar of the subject; o respective data of organoid response; and o at least one measured best RECIST response of the subject; obtaining each of the trained learning models by iteratively providing, to the corresponding learning model, the entries of the training dataset one by one until convergence, with the best RECIST response associated with each entry being used as ground truth at each iteration.
[0039] In addition, the disclosure relates to a computer program comprising software code adapted to perform a method for predicting a patient response to an agent or a method for predicting an effect of an agent on a patient or a population of patients or a method for training, compliant with any of the above execution modes when the program is executed by a processor.
[0040] The present disclosure further pertains to a non-transitory program storage device, readable by a computer, tangibly embodying a program of instructions executable by the computer to perform a method for predicting a patient response to an agent or a method for predicting an effect of an agent on a patient or a population of patients or a method for training, compliant with the present disclosure.
[0041] Such a non-transitory program storage device can be, without limitation, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor device, or any suitable combination of the foregoing. It is to be appreciated that the following, while providing more specific examples, is merely an illustrative and not exhaustive listing as readily appreciated by one of ordinary skill in the art: a portable computer diskette, a hard disk, a ROM, an EPROM (Erasable Programmable ROM) or a Flash memory, a portable CD-ROM (Compact-Disc ROM).
[0042] It is further provided an in vitro method for predicting a patient response to an agent, comprising the steps of: a) providing at least one organoid; b) bringing the at least one organoid in contact with the agent; c) determining an organoid response data from the at least one organoid; d) predicting a patient response to the agent by the method discussed above, or predicting an effect of an agent on a population of patients by the method discussed above.
[0043] The in vitro method may comprise a step of a) providing a plurality of organoids, and b) bringing the plurality of organoids in contact with the agent, said plurality of organoids being derived from all or part of a group consisting of: cerebral organoids, gastrointestinal organoids (for example foregut organoids, midgut organoids, hindgut organoids, intestinal organoids, or gastric organoids), lingual organoids, tooth organoids, thyroid organoids, thymic organoids, testicular organoids, prostate organoids, hepatic organoids, pancreatic organoids, epithelial organoid, lung organoids, kidney organoids, gastruloid organoids, blastoid organoids, endometrial organoids, cardiac organoids, retinal organoids, breast cancer organoids, colorectal cancer organoids, gliobastoma organoids, neuroendocrine tumor organoids, myelin organoids, blood-brain barrier (BBB) organoids.
[0044] It is further provided a method for screening one or more agent(s) for a therapeutic or prophylactic drug or cosmetic, wherein the method comprises: a) providing at least one organoid;b) bringing the at least one organoid in contact with one or more agent(s); c) determining an organoid response data from the at least one organoid; d) predicting a patient response to the agent by the method according to claim 7, or predicting an effect of an agent on a population of patients by the method as discussed above; e) comparing the patient response or the effect of the agent to a reference data; f) identifying the agent as a candidate molecule for a therapeutic or prophylactic drug or cosmetic.DEFINITIONS
[0045] In the present invention, the following terms have the following meanings:
[0046] The terms “adapted” and “configured” are used in the present disclosure as broadly encompassing initial configuration, later adaptation or complementation of the present device, or any combination thereof alike, whether effected through material or software means (including firmware).
[0047] The term “processor” should not be construed to be restricted to hardware capable of executing software, and refers in a general way to a processing device, which can for example include a computer, a microprocessor, an integrated circuit, or a programmable logic device (PLD). The processor may also encompass one or more Graphics Processing Units (GPU), whether exploited for computer graphics and image processing or other functions. Additionally, the instructions and / or data enabling to perform associated and / or resulting functionalities may be stored on any processor- readable medium such as, e.g., an integrated circuit, a hard disk, a CD (Compact Disc), an optical disc such as a DVD (Digital Versatile Disc), a RAM (Random- Access Memory) or a ROM (Read-Only Memory). Instructions may be notably stored in hardware, software, firmware or in any combination thereof.
[0048] “Avatar” designates a representation of a subject / patient. An avatar comprises at least one set of biological data corresponding to said patient and at least one set of clinical data (also called in this text digital data) corresponding to said patient.
[0049] The “biological data from the subject / patient” may comprise or consist of data retrieved from all or part of a biological sample, or fraction thereof, from said patient. Accordingly, a biological sample may be any sample that may be taken from the subject, such as a serum sample, a plasma sample, a urine sample, a blood sample, a lymph sample, or a biopsy, and transformed samples such as patient-derived cells and / or patient-derived biological fluids, and fractions thereof. The biological data from the subject / patient may have a time stamp which defines the time of the data collection. New biological data, for each patient / subject may be created over time with each additionalexperiment conducted on the organoid cell line. The new data may be added using the distinguishing time stamp or replace the existing biological data by updating the existing time stamp. In one example, the biological data further comprise omics data of the one or more organoids previously obtained from the patient. In one example, the biological data comprises also “organoid data” obtained from at least one organoid that was not exposed to an agent. The organoid data (i.e.; organoid non-exposed data) may be selected from the group consisting of: viability (e.g., viability measurements made with CellTiter-Glo (CTG) assay), growth rate, a proteomic feature, a genomic feature, an epigenomic feature, a metabolic feature, a microbiome feature, an imaging feature, a histological feature or a combination thereof. Assessment of cell viability may be achieved by any cell viability assay known to the skilled in the Art, for example an ATP assay.
[0050] The “digital data from the subject / patient” or “clinical data from the subject / patient” may comprise at least one first group of clinical (i.e., digital) data from said patient, such as electronic health records (EHRs). The clinical data may also comprise information such as tumor site, age at diagnosis, number of metastases and the like. Said first group of clinical data may include one of more selected from the group consisting of: a proteomic feature, a genomic feature, an epigenomic feature, a transcriptomic feature, a metabolic feature, a microbiome feature, an imaging feature, a histological feature. Optionally, said clinical data may comprise at least one further group of data, such as a second group of data selected from the group consisting of: a proteomic feature, a genomic feature, an epigenomic feature, a transcriptomic feature, a metabolic feature, a microbiome feature, an imaging feature, a histological feature; with the second (or further) group(s) of data being different from the first group. Yet optionally, said clinical data may comprise patient longitudinal history, including past treatments and their respective responses.
[0051] As used herein, an “organoid” refers to an artificial three-dimensional tissue construct comprising a plurality of cells; which is functionally capable of retaining all or part of an organ tissue, or a portion thereof. Such organoids may comprise at least one, in particular more than one, differentiated cell type(s); and optionally one or more than one undifferentiated cell type(s). The organ tissue which capability is retained by suchorganoids may be a healthy tissue or a tumoral (i.e., cancerous) counterpart. In a non- exhaustive manner, organoid may thus be selected from the group consisting of: cerebral organoids, gastrointestinal organoids (for example foregut organoids, midgut organoids, hindgut organoids, intestinal organoids, or gastric organoids), ovarian organoids, lingual organoids, tooth organoids, thyroid organoids, thymic organoids, testicular organoids, prostate organoids, hepatic organoids, pancreatic organoids, epithelial organoid, lung organoids, kidney organoids, gastruloid organoids, blastoid organoids, endometrial organoids, cardiac organoids, retinal organoids, breast cancer organoids, colorectal cancer organoids, gliobastoma organoids, neuroendocrine tumor organoids, myelin organoids, blood-brain barrier (BBB) organoids. In one example, the organoids are obtained from patient cells that have not been genetically modified, for example by human direct intervention of the genome, notably between the biopsy and the organoid initial creation.
[0052] Accordingly, an organoid which is submitted to a stimulus / agent, for example which is brought in contact with one or more physical, chemical and / or biological agent such as a therapeutic agent, will generate an “organoid response”, which is constitutive of what is referred herein as an “organoid response data” or “data of organoid response”, thus defining an “organoid feature” with respect to the tested agent. The organoid response data may be selected from the group consisting of: viability (e.g., viability measurements made with CellTiter-Glo (CTG) assay), growth rate, a proteomic feature, a genomic feature, an epigenomic feature, a metabolic feature, a microbiome feature, an imaging feature, a histological feature or a combination thereof. Assessment of cell viability may be achieved by any cell viability assay known to the skilled in the Art, for example an ATP assay.
[0053] As used herein, an “agent” may refer to any agent / compound / nucleic acid / polypeptide which is susceptible to be administered to, or brought into contact with, a patient / subject, or an organoid. Thus, the term may comprise or consist of a physical treatment (for example radiation therapy), a small molecule, a therapeutic compound, a therapeutic candidate, a polypeptide, an antibody or antigen-binding fragment thereof, an aptamer, a nucleic acid (for example a guide RNA, a silencing RNA, a micro-RNA, acoding RNA, a non-coding RNA, antisense oligonucleotide - ASO-), a nanoparticle, an expression vector, a virus, a cell, a prokaryotic cell (for example a bacteria), an eukaryotic cell (for example an immune cell), combinations thereof, and / or compositions thereof. Hence, the agent may be a therapeutic agent or non-therapeutic agent, i.e., it does not have the finality of treating a patient. Examples of therapeutic agents which are thus considered under this definition include, for example, any agent which is susceptible to treat or prevent or reduce the likelihood of occurrence or re-occurrence of a given condition. Hence, the term may, in particular, refer to any therapeutic agent, such as those selected from: an antibody or antigen-binding fragment thereof, an immunoconjugate (such as Antibody Drug Conjugate -ADC-), a cytotoxin, a chemotherapeutic agent, a cytokine, an immunosuppressant, an immune stimulator, a lytic peptide, a radioisotope, and antiviral agent, an antiparasitic agent, an antimicrobial agent, a chimeric antigen receptor, a nucleic acid, an expression vector and the like. Examples of cells may, in particular, include live cells and / or immune cells, and engineered forms thereof such as CAR T-cells or CAR-NK cells.
[0054] As used herein, a “proteomic data” or corresponding “proteomic feature” refers to any data which is indicative of the occurrence of a protein, or group of proteins which is expressed or susceptible to be expressed in a given cell or group of cells (e.g., an organoid or a tissue), under a given set of conditions. For example, proteomic data can be retrieved from any known method for determining the occurrence of, or amount of, a given protein / polypeptide or fragment thereof, such as (in a non-exhaustive manner).
[0055] As used herein, a “genomic data” or corresponding “genomic feature” refers to any data which is indicative of the structure (e.g. sequence and / or accessibility) of a corresponding genome, or part thereof, in a given cell or group of cells (e.g., an organoid or a tissue). Hence, the term may refer, in particular, to the detection or quantification of any nucleotide sequence corresponding to all or part of the genome of the cell or group cell.
[0056] As used herein, an “epigenomic data” or corresponding “epigenomic feature” refers to any data which is indicative of one or more epigenetic modification(s) of the said genome or part thereof, in a given cell or group of cells (e.g., an organoid or a tissue); thismay also refer, in particular, to the detection or quantification of any epigenetic modification of the said genome or part thereof, such as any covalent modification of a deoxyribonucleic acid (e.g. DNA methylation) or post-translation modification of the histones.
[0057] As used herein, a “transcriptomic data” or corresponding “transcriptomic feature” refers to any data which is indicative of the occurrence of a nucleic acid (e.g. any ribonucleic acid, such as siRNAs, miRNA, coding and non-coding RNAs, pre-mature messenger RNAs, and mature messenger RNAs) which is transcribed or susceptible to be transcribed in a given cell or group of cells thereof. Accordingly, this term may refer to the detection or quantification of such nucleic acid(s).
[0058] As used herein, a “metabolic data” or corresponding “metabolic feature” refers to any data which is indicative of the occurrence of a small-molecule or compound which is present or produced or susceptible to be present or susceptible to be produced in - or by- a given cell or group of cells (e.g., an organoid or a tissue), under a given set of conditions. Such small-molecules or compounds may, in particular, be selected from the group consisting of sugars, nucleotides, amino acids and lipids.
[0059] As used herein, a “microbiome data” or corresponding “microbiome feature” refers to any data which is indicative of the occurrence of one or more micro-organisms (e.g. bacteria, archaea, fungi, algae, and the like) which are present or susceptible to be present in, or associated to, a given cell or group of cells (e.g., an organoid or a tissue), under a given set of conditions.
[0060] As used herein, an “imaging data” or corresponding “imaging feature” refers to any data which is indicative of the morphology and / or anatomy and / or structure of a given cell or group of cells (e.g., an organoid or a tissue) or part thereof, under a given set of conditions. In a non-exhaustive manner, an imaging data or imaging feature may be retrieved from medical imaging, microscopy (e.g. optic microscopy, electron microscopy, X-ray microscopy) and live-cell imaging methods of the cell or group of cells, such as those selected from the group consisting of phase-contrast microscopy, fluorescent microscopy, quantitative phase contrast microscopy and holo tomography.
[0061] As used herein, a “histological data” or corresponding “histological feature” refers to a subset of imaging data or imaging features, respectively, which is indicative of the of the microscopy morphology and / or anatomy and / or structure of a given cell or group of cells (e.g., an organoid or a tissue) or part thereof.
[0062] For a clinical trial, there are known parameters to assess how well a new treatment works. “Progression-free survival (PFS)” is the length of time during and after the treatment of a disease, such as cancer, that a patient lives with the disease, but it does not get worse. “Overall Survival (OS)” is the length of time from either the date of diagnosis or the start of treatment for a disease, such as cancer, that patients diagnosed with the disease are still alive. “Overall Response Rate (ORR)” is the percentage of people in a study or treatment group who have a partial response or complete response to the treatment within a certain period of time. A partial response is a decrease in the size of a tumor or in the amount of cancer in the body, and a complete response is the disappearance of all signs of cancer in the body.
[0063] The term “Overall Response (OR)” refers to a clinical outcome measure indicating whether a patient’s tumor burden has decreased following treatment, as determined according to standardized response criteria such as the best RECIST response. Overall Response is defined as a complete response, when all target lesions have disappeared, or as a partial response, when a decrease of at least 30% in the sum of diameters of target lesions occurs.
[0064] The term “best RECIST response” refers to the most favorable tumor overall response category achieved by a patient during a defined observation period, as determined in accordance with the Response Evaluation Criteria in Solid Tumors (RECIST). The response is assessed based on tumor measurements and classified as complete response, partial response, stable disease, or progressive disease. The best RECIST response corresponds to the highest level of response observed for the patient from the initiation of treatment until disease progression or study termination.
[0065] “Electronic Health Record (EHR)” designates an electronic version of patients’ medical history, which is maintained by the provider over time. More in detail, an electronic health record (EHR) is a computer-implemented system comprising a digitalrepresentation of a patient’ s medical information, which is created, stored, and managed electronically, and which provides a longitudinal record of patient health. The EHR includes patient-identifying information together with clinical data such as medical history, diagnoses, medications, treatment plans, immunization records, laboratory and diagnostic test results, and physician notes, as well as administrative and billing information. By integrating such data, the EHR enables real-time access to and sharing of patient health information across authorized healthcare providers and institutions.
[0066] “Omics data” refers to data generated from high-throughput technologies used to study the various "omes" of an organism, such as the genome (all the genetic material), transcriptome (all the RNA molecules), proteome (all the proteins), metabolome (all the small molecules), and interactome (all the interactions between biomolecules). Omics data is often used in systems biology and functional genomics to study the relationships between different molecules and how they interact to affect the overall function of cells, tissues, and organisms. Omic data can be complex, high-dimensional, and noisy and requires specialized computational methods and tools for analysis and interpretation.
[0067] As known per se a “cohort” is a group of people with a shared characteristic. A “cohort study” designates a type of epidemiological study in which a group of people with a common characteristic is followed over time to find how many reach a certain health outcome of interest (disease, condition, event, death, or a change in health status or behavior).
[0068] “Machine learning (ML)” designates in a traditional way computer algorithms improving automatically through experience, on the ground of training data enabling to adjust parameters of computer models through gap reductions between expected outputs extracted from the training data and evaluated outputs computed by the computer models.
[0069] A “hyper-parameter” presently means a parameter used to carry out an upstream control of a model construction, such as a remembering-forgetting balance in sample selection or a width of a time window, by contrast with a parameter of a model itself, which depends on specific situations. In ML applications, hyper-parameters are used to control the learning process.
[0070] “Datasets” are collections of data used to build an ML mathematical model, so as to make data-driven predictions or decisions. In “supervised learning” (i.e. inferring functions from known input-output examples in the form of labelled training data), three types of ML datasets (also designated as ML sets) are typically dedicated to three respective kinds of tasks: “training”, i.e. fitting the parameters, “validation”, i.e. tuning ML hyperparameters (which are parameters used to control the learning process), and “testing”, i.e. checking independently of a training dataset exploited for building a mathematical model that the latter model provides satisfying results.
[0071] A “neural network (NN)” designates a category of ML comprising nodes (called “neurons”), and connections between neurons modelled by “weights”. For each neuron, an output is given in function of an input or a set of inputs by an “activation function”. Neurons are generally organized into multiple “layers”, so that neurons of one layer connect only to neurons of the immediately preceding and immediately following layers.
[0072] The above ML definitions are compliant with their usual meaning, and can be completed with numerous associated features and properties, and definitions of related numerical objects, well known to a person skilled in the ML field. Additional terms will be defined, specified or commented wherever useful throughout the following description.BRIEF DESCRIPTION OF THE DRAWINGS
[0073] The present disclosure will be better understood, and other specific features and advantages will emerge upon reading the following description of particular and non-restrictive illustrative embodiments, the description making reference to the annexed drawings wherein:
[0074] Figure 1A presents an example flowchart of the method;
[0075] Figure IB is a block diagram representing schematically a particular mode of a device 1 for predicting a patient’ s response to an agent;
[0076] Figure 2 presents an example schematic of the multivariate prediction workflow;
[0077] Figure 3 presents an example schematic of the training and validation workflow;
[0078] Figure 4 presents an example schematic of the deployment workflow;
[0079] Figure 5 presents an example of biomarker identification using the method;
[0080] Figure 6 shows an example system embodying the device;
[0081] Figures 7a)-c) present an example result of the PFS prediction using the method;
[0082] Figures 8a)-c) present an example result of the OS prediction using the method; and
[0083] Figures 9a)-b) present an example result of the OR prediction using the method.
[0084] Figure 10 presents an example flowchart of the method for predicting an effect of an agent on a population of patients.
[0085] Figure 11 presents an example flowchart of the method for training a learning model.
[0086] Figure 12 presents an example flowchart of the method for predicting a patient response to an agent.
[0087] Figure 13 presents an example flowchart of the method for screening one or more agents for a therapeutic or prophylactic drug or cosmetic.
[0088] On the figures, the drawings are not to scale, and identical or similar elements are designated by the same references.ILLUSTRATIVE EMBODIMENTS
[0089] The present description illustrates the principles of the present disclosure. It will thus be appreciated that those skilled in the art will be able to devise various arrangements that, although not explicitly described or shown herein, embody the principles of the disclosure and are included within its scope.
[0090] All examples and conditional language recited herein are intended for educational purposes to aid the reader in understanding the principles of the disclosureand the concepts contributed by the inventor to furthering the art, and are to be construed as being without limitation to such specifically recited examples and conditions.
[0091] For example, while the embodiments discussed herein are illustrated in particular application to the field of oncology and cancer treatments, it will be appreciated by those skilled in the art that the principles of the disclosure may be applied to any other disease and its respective treatments.
[0092] Moreover, all statements herein reciting principles, aspects, and embodiments of the disclosure, as well as specific examples thereof, are intended to encompass both structural and functional equivalents thereof. Additionally, it is intended that such equivalents include both currently known equivalents as well as equivalents developed in the future, i.e., any elements developed that perform the same function, regardless of structure.
[0093] Thus, for example, it will be appreciated by those skilled in the art that the block diagrams presented herein may represent conceptual views of illustrative circuitry embodying the principles of the disclosure. Similarly, it will be appreciated that any flow charts, flow diagrams, and the like represent various processes which may be substantially represented in computer readable media and so executed by a computer or processor, whether or not such computer or processor is explicitly shown.
[0094] The functions of the various elements shown in the figures may be provided through the use of dedicated hardware as well as hardware capable of executing software in association with appropriate software. When provided by a processor, the functions may be provided by a single dedicated processor, a single shared processor, or a plurality of individual processors, some of which may be shared.
[0095] It should be understood that the elements shown in the figures may be implemented in various forms of hardware, software or combinations thereof. Preferably, these elements are implemented in a combination of hardware and software on one or more appropriately programmed general-purpose devices, which may include a processor, memory and input / output interfaces.
[0096] In reference to Figure 1A, it is provided a computer- implemented method for predicting a patient response (i.e., a response of a patient) to an agent. The methodcomprises, at step S10, receiving at least an avatar of a patient 10 (i.e., the patient for whom the response is to be predicted) and respective data of organoid response 11 for said patient. By an “organoid response of said patient” it is meant a response of an organoid which is obtained from the cells of the patient and therefore matches with said patient. Accordingly, an organoid response to a matching organoid relates to a response from an organoid, for which all or part of the organoid is obtained to match the patient (or a population of patients). Such an organoid response may be, for example, a response to all or part of an organoid matched with said patient (i.e., organoid derived from the patient cells). Such an organoid response may be, for example, a response of one or more organoids obtained from the cells of said patient.
[0097] With respect to the prior art proposing the use of 2D cell cultures (i.e., not organoids, which are 3D cell cultures), patient-matched organoid-based responses offer an advantage over cell-based responses. Organoids maintain a three-dimensional structure that preserves cell-cell interactions and mimics the spatial organization of the original tumor. By contrast, cells grown in two-dimensional monolayers (e.g., in a culture dish) are artificially polarized between their upper and lower surfaces, which alters their molecular profile and causes deviations from the in vivo tumor state. Organoids, by preserving a 3D organization, avoid such artifacts and thus provide a representation that more closely reflects the biological reality of the patient’s tumor.
[0098] Other prior art proposed to predict the response of an individual patient to a therapeutic agent using organoid response data obtained from available generic organoids to the agent as a predictive feature. The generic organoids are not derived from the patient cells, they must instead be selected from other patients’ organoid collections or human cellular lineages (i.e., not coming from a patient at all), typically chosen as they have the same genetic background (e.g., mutation profile). However, the genetic background alone is insufficient to accurately reproduce the molecular characteristics of the patient’s tumor, since differences in the epigenetic landscape and / or cellular state (e.g., transcriptomic profile) remain. As a result, the organoid’ s response to the agent will not accurately reflect the true response of the patient’ s tumor. Advantageously, this limitation is overcome by generating organoids directly from the patient’s own cells, ensuring that the organoidshave the same molecular characteristics as the patient’s tumor and thereby providing a response to the agent that closely mirrors the actual tumor response in vivo.
[0099] In alternative examples, such an organoid response may be, for example, a response to all or part of an organoid which does not itself include any cell or cell line directly obtained from the same patient (or population of patients) for which a patient response must be predicted. Advantageously, and as the organoid response is valid even for patients from whom the cell or cell line of the organoid is not directly obtained, the devices and methods of the disclosure may thus also be applicable to patients for which there is no personalized organoid(s) available, thus overcoming the need for conceiving, or maintaining in culture, unnecessary biological material.
[0100] As reported herein, an avatar of a patient is a representation of said patient. Said avatar and said respective data of organoid response define at least one patient feature and at least one organoid response feature from the organoid exposed to an agent respectively. By a feature, it is meant a characterizing parameter. Thereby, a patient feature may be one or more parameters which define / describe a health status of the patient, such as for example Vital signs (blood pressure, heart rate) Pulmonary function (oxygen saturation), Eastern Cooperative Oncology Group (ECOG) score (scale used in oncology to describe how a patient’ s disease affects their daily living abilities) and the like. Similarly, an organoid response feature may be one or more parameters which define / describe a response of an organoid. An organoid response feature may be equivalently called a “functional feature”.
[0101] The method further comprises, at Step S19, for defining the at least one patient feature using said avatar and defining the at least one organoid response feature using at least said data of organoid response for the same patient.
[0102] More in detail, the organoid response data and / or the avatar biological data are used to define said at least one organoid response feature. In one example, the at least one organoid response feature is defined using the organoid response data, which comprises the organoid response data from one or more organoid (previously) obtained from patient cells that have been exposed to the agent under examination for one or more predefined concentration of the agent. The organoid response feature may be for example aclassification of the organoid as “responder” or “not responder” to the agent to which the organoid was exposed, or a measure of the growth rate of the organoids or a feature derived thereof.
[0103] In one alternative example, at least one organoid response feature is defined using organoid response data obtained from one or more organoids exposed to the agent under examination, together with organoid data from at least one organoid previously derived from patient cells that has not been exposed to the agent. This organoid data may be incorporated into the biological data of the patient. For a given patient, N organoids may be derived from patient cells. Within this set, a number A may be left unexposed to the agent, while the remaining N-A organoids are exposed to the agent under examination at one or more predefined concentrations. The number A and N may be chosen so that the number of unexposed organoids is comprised between 1% and 10% of the total number of organoid N. In this example, the at least one organoid response feature may be a combination of the organoid response data and the biological data relating to the (nonexposed) organoid data, such as the ratio between the viability of the exposed organoid and the viability of the non-exposed organoids, or the ratio between the growth rate of the exposed organoid and the growth rate of the non-exposed organoids and the like.
[0104] The avatar is used to define at least one patient feature notably using said clinical data and / or biological data of the patient.
[0105] In certain embodiments, the patient feature(s) may be derived from heterogeneous clinical data sources, such as electronic health records (EHRs), laboratory test results, imaging data, genomic data, or physician notes. The raw clinical data may be subjected to preprocessing and transformation steps, including normalization, encoding, and harmonization across data types. From this multi-dimensional dataset, a set of patient features can be obtained by applying dimensionality reduction techniques (e.g., principal component analysis, autoencoders, or manifold learning) and / or by selecting the most relevant features through statistical, heuristic, or machine-leaming-based feature selection methods. Such patient features represent a reduced yet informative subset of the original clinical data.
[0106] According to this embodiment, clinical data of a patient, and optionally the biological data of the patient (comprising the data obtained from biological sample(s) such as serum sample, a plasma sample, a urine sample, a blood sample, a lymph sample -e.g., laboratory results- and / or omics data from one or mode organoids of the patient), may be provided to at least one trained large language model (LLM) in order to generate patient relevant embedding (i.e., at least one patient feature). The at least one LLM may be a foundation model or a fine-tuned model that has been further refined with specialized datasets (e.g., scientific texts, clinical and patient data) to adapt it to particular uses. In an example, ChatGPT and LLaMA are foundation models, while ClinicalBERT and BioGPT are fine-tuned models.
[0107] More in detail, the clinical data of a patient are prompted to the LLM. The inputs may take the form of structured tables or free-text clinical reports from physicians. In addition to the patient’s clinical data and optionally patient’s biological data, the at least one LLM is provided with example data from other patients who have been exposed to the same specific agent for which the outcomes are known (e.g., OS, PFS). These examples guide the LLM(s) model’s predictions and help format its response (few-shot prompting).
[0108] In one example, multiple LLMs may be chained, meaning that the output of one model serves as the input for the next, enabling a progressive workflow to achieve the desired outcomes. For example, a first LLM processes the patient’ s clinical data to extract relevant information, such as treatment history or patient description. A second LLM then takes this output and generates a coherent medical report summarizing the patient’s case. A third LLM uses the medical report, along with a given agent, to provide the patient embedding. The patient embedding calculation step may be repeated several times, and the final patient feature may be obtained by averaging the results across runs, providing a more robust estimate. Each LLM in the chain may be configured with a specific temperature, which controls the randomness of its output. In one example, a temperature of 0 produces deterministic and consistent results, while higher temperatures allow for more variability and creative responses.
[0109] Using large language models (LLMs) to obtain a patient embedding provides significant advantages over conventional feature selection methods. Clinical andbiological data from a patient are often highly heterogeneous, comprising structured variables (e.g., laboratory values, diagnoses, medications), semi-structured data (e.g., EHR codes), and unstructured text (e.g., clinical notes, imaging reports). Traditional feature selection approaches rely on identifying a subset of variables considered most relevant, which may lead to loss of contextual or latent information carried by the excluded data. By contrast, an LLM-based embedding enables the transformation of this multi- source, high-dimensional data into a dense representation that concentrates information otherwise diluted across a large number of variables. Such embeddings are capable of capturing semantic relationships and higher-order interactions between data elements that are difficult to extract with conventional dimensionality reduction techniques. As a result, the generated patient feature provides a more faithful and holistic representation of the patient’ s condition, while remaining compact and computationally tractable for downstream tasks such as response prediction.
[0110] The method further comprises, at Step S20, receiving a first ensemble of one or more trained learning models 12. Each trained learning model (which may be called equivalently, a trained model) of the first ensemble is configured to receive as input, with respect to said patient, one or more of said at least one patient feature and one or more of said at least one organoid response feature, and to provide a first output vector. In other words, the first output vector is an encoding of the inputted one or more of said at least one patient / organoid response features.
[0111] The receiving (or equivalently, obtaining), for example in Step S 10 and S20, may comprise accessing to a local storage or a storage on a remote server.
[0112] The method further comprises, at Step S30, calculating the first output vector by providing one or more of said at least one patient feature and one or more of at least one organoid response feature as input to said first ensemble of one or more trained learning models. In one example, at least one of the one or more trained learning models is a regression model.
[0113] The method further comprises, at Step S40, obtaining at least a PFS and / or an OS for the patient using the first output vector. In one example, the first output vector comprises a progression-free survival (PFS) value and / or an overall survival (OS) value.In this example, obtaining at least a Progression-Free Survival (PFS) and / or an Overall Survival (OS) for the patient using the first output vector comprises applying an identity operation to the first output vector to obtain at least a Progression-Free Survival (PFS) and / or an Overall Survival (OS). Alternatively, obtaining at least a PFS and / or an OS for the patient using the first output vector the Progression-Free Survival (PFS) and / or an Overall Survival (OS) is obtained using a pre-trained regression model taking as input the first output vector.
[0114] The method, then at Step S50, obtains (i.e., determines) a prediction of the patient response to the agent using said obtained PFS and / or OS. In examples, such an obtained prediction may be an output comprising said obtained PFS and / or OS.
[0115] Alternatively, the prediction of the patient’s response to the agent may be obtained by providing the PFS and / or OS to a pre-trained classifier model. The pre-trained classifier model may, for example, be a logistic regression model trained to predict whether the patient is a “responder” or a “non-responder” to the agent. Alternatively, to obtain a binary classification of the response to the agent, the pre-trained classifier model may be a random forest. More generally, the pre-trained classifier model may be any suitable machine learning or statistical model configured to perform a binary classification task, including but not limited to decision tree-based models, support vector machines, neural networks, or ensemble methods. The pre-trained classifier model may have been trained using a training dataset comprising training samples obtained from a plurality of subjects that had been exposed to the same agent under evaluation for the patient, each training sample comprising at least the PFS and OS values together with, as ground truth, the actual response of the subject to the agent (i.e., responder or non- responder).
[0116] Such a method for predicting a patient response to an agent may be equivalently referred to as “individual response prediction method”.
[0117] The individual response prediction method according to this disclosure constitutes an improved solution to predict the responses of patients, e.g., cancer patients, in clinical tests. This improvement is achieved through capturing the response from a given patient by the response of the matched organoid, for example a matched patient-derived organoid cell line. The method modulates such a captured organoid response and modulates it using the biological data and the clinical data of the avatar.
[0118] Figure 2 shows an example workflow for prediction according to the method. For an agent 210, which is a therapeutic compound patient’s cell, responses 220 is obtained (e.g., using patient-derived organoids (PDOs), or cell lines). Such an obtained organoid response feature is then provided, in combination with the patient feature 230 (which defines a context about the patient) to the trained ML model 240 and the prediction 250 is obtained.
[0119] It is further provided a device for predicting a patient response to an agent. The device is illustrated figure IB. The device 1 comprises at least one input configured to receive 101 at least an avatar of a patient 10 and respective data of organoid response 11 for said patient. As known per se, an avatar is a representation of said patient. Said avatar and said respective data of organoid response define at least one patient feature and at least one organoid response feature, respectively.
[0120] Said at least one input is further configured to receive 103 a first ensemble of one or more trained learning models 12. Each trained learning model (which may be called equivalently, a trained model) of the first ensemble is configured to receive as input, with respect to said patient, one or more of said at least one patient feature and one or more of said at least one organoid response feature, and to provide a first output vector.
[0121] Said input may be any known hardware or a system of hardware in the field adequate for receiving (or equivalently, obtaining) inputs in a format to be processed by the processor. The receiving may comprise accessing to a local storage or a storage on a remote server.
[0122] The device further comprises at least one processor. Said at least one processor is configured to define 102 the at least one patient feature using said avatar and define the at least one organoid response feature using at least said data of organoid response for the same patient. Said at least one processor is further configured to calculate 103 the first output vector by providing said at least one patient feature and at least one organoid response feature as input to said first ensemble of trained learning models. Said at least one processor is further configured to obtain 104 at least a PFS and an OS for the patientusing the first output vector and optionally, a pre-trained regression model from the first ensemble. In one example, the first output vector is an embedding of the patient feature(s) and organoid response feature(s). Said at least one processor obtain 105 a prediction of the patient response to the agent using said obtained PFS and / or OS, optionally by providing the obtained PFS and / or OS to a pre-trained classifier model.It is further provided a method for designing patients’ avatars for multivariate prediction. To combine multiple variables in order to conduct multivariate predictions, it is necessary to match the different data modalities referring to a single patient. This requires a database with matching genomic, transcriptomic and clinical data from the same patients. Each entry of this dataset is further matched with experimental results on the accompanying cell lines. This method for avatar design may be referred to, hereinafter, as “Avatar Cohort Generation method”.
[0123] The Avatar Cohort Generation method may batch multiple patient avatars together in order to form a cohort of patient avatars. This patient avatar cohort can be generated randomly or in a specific manner in order for it to have distinct properties. For instance, the composition of the cohort can be made to resemble that of a population of interest based on the genomic, demographic or clinical information available. For example, the method may generate a cohort of PDAC avatars, having received a therapeutic agent (such as a chemotherapy agent or a therapeutic antibody) for a given time, in particular a time suitable for the therapeutic agent to have putative therapeutic effect), in a given population or sub-population of patients (e.g. those that have a KRAS G12D mutation).
[0124] In practice, the generation of cohorts is done by selecting a total number of avatars to work on and then by filtering the avatars based on their digital properties of interest (in the above example disease type, past treatment, past treatment duration, KRAS mutation status).
[0125] The avatar cohort generation method, to mimic clinical responses, design patient avatars combining a biological component and a digital component. The biological component of the patient avatars is composed of living patient-derived cells. These cellsare matched with data originating from the same patients. This data is comprised of genomic, transcriptomic data, electronic health records and digital histology.
[0126] It is further provided a device and a computer-implemented method for training (i.e., a training method) a learning model of the at least first / second / third ensemble as discussed above. The method comprises receiving a training dataset comprising entries. Each entry of the dataset corresponds (i.e., is in relation) to a subject (e.g., a patient). Said entry comprises, for the corresponding subject, an avatar, and data of organoid response. In certain embodiments, the entry may further comprise clinical outcome measures, such as progression-free survival (PFS), overall survival (OS), best RECIST response, overall response (OR), or other similar endpoints, which may be used as ground truth during training.
[0127] The method then trains the learning model(s) using the training dataset. The training method may train each of the model(s) of the first / second / third ensemble according to any known method in the field of machine learning. The method may train each model of the first / second ensemble according to a particular training. In particular, the models of the first ensemble may be trained to provide an embedding (first output vector) representative of progression-free survival (PFS) and / or overall survival (OS) from the input patient feature(s) and organoid response feature(s). The models of the second ensemble may be trained to provide an embedding (second output vector) representative of overall response (OR) for the patient from the input patient feature(s) and organoid response feature(s). The models of the third ensemble may be trained to provide an embedding (third output vector) representative of the best RECIST response for the patient from the input patient feature(s) and organoid response feature(s).
[0128] In some embodiments wherein the learning models are regression models, the training comprises minimizing a loss function quantifying the difference between the predicted embedding and the ground truth value for the corresponding clinical outcome measure. Suitable loss functions may include, without limitation, mean squared error (MSE), mean absolute error (MAE), or negative log-likelihood, depending on the type of endpoint. When the ground truth is categorical (e.g., best RECIST response), the regression model may be trained to approximate the ordinal value of the category, thereby capturing both the categorical nature and the relative ordering of responses. Training maybe performed using gradient-based optimization techniques, such as stochastic gradient descent or variants thereof, and may include regularization strategies to prevent overfitting. In one example, at least one of the learning model may be a classifier model.
[0129] According to one embodiment, each of the trained learning models of the first ensemble has been previously trained using a training dataset comprising entries corresponding to a plurality of subjects, each entry corresponding to a subject.
[0130] In one example, each subject of said plurality of subjects is affected by the same pathology as the patient (e.g., breast cancer).
[0131] In one example, the organoids of the subjects are obtained from the same type of cells as those of the patient. If the patient is affected by a cancer, like lung cancer, the organoids of both the patient and the subject are obtained from cancers cells.
[0132] In one example, at least one of the organoids of the subjects has been exposed to the same agent for which the response of the patient has to be predicted.
[0133] More in detail, each entry for one subject may comprise an avatar of the subject, corresponding organoid response data, and at least one measured progression-free survival (PFS) value and / or measured overall survival (OS) value of the subject. By measured progression-free survival (PFS) and / or measured overall survival (OS) is intended the actual values obtained for the subject following treatment for the pathology (i.e., not predicted one).
[0134] As explained above each of the trained learning models is configured to provide as output a first output vector representative of a Progression-Free Survival (PFS) and / or a measured Overall Survival (OS). The training may be performed on initial models in order to obtain the trained learning models. Each of the trained learning models is obtained by iteratively providing, to a corresponding initial model, the entries of the training dataset one by one until convergence, with the progression-free survival (PFS) and / or overall survival (OS) associated with each entry being used as ground truth at each iteration.
[0135] It is further provided a method for predicting an effect of an agent on a population of patients, which is described in figure 10. Such a method for predicting an effect of an agent on a population of patients may be equivalently referred to as a “population-level response prediction method”. The method comprises providing S60 (i.e., receiving or obtaining) a test dataset. The test dataset comprises a plurality of entries. Each entry corresponds (i.e., is in relation to) to a test patient. The entry comprises a respective avatar of the test patient and data of organoid response for said test patient. The method further comprises obtaining S61 a simulated population dataset by applying the method of predicting a patient response to said agent as discussed above. In other words, the method obtains said simulated population dataset by applying the individual response prediction method to the respective avatar and organoid response data for each entry of the test dataset. Thereby, the method obtains a simulated population dataset which comprises at least a PFS and / or an OS for each entry, i.e., for each respective test patient.
[0136] More in details, the population-level response prediction method comprises a step of receiving: (i) a test dataset comprising a plurality of entries, each entry corresponding to a test patient and comprising the respective avatar of the test patient and data of organoid response for said test patient, (ii) a first ensemble of one or more trained learning models each configured to receive as input, with respect to said patient, one or more of said at least one patient feature and one or more of said at least one organoid response feature, and to provide a first output vector and optionally (iii) at least one second ensemble of one or more trained learning models each configured to receive as input, with respect to said patient, one or more of said at least one patient feature and one or more of said at least one organoid response feature, and to provide a second output vector, said second ensemble of one or more trained learning models comprising at least one logistic regression model.
[0137] The population-level response prediction method comprises then obtained a simulated population dataset comprising at least a PFS and / or an OS for each respective test patient by, for each test patient of the test dataset executing a succession of steps as follows: defining at least one patient feature using the avatar and defining at least one organoid response feature using at least said respective data of organoid response;calculate the first output vector by providing said at least one patient feature and at least one organoid response feature as input to said first ensemble of one or more trained learning models; obtain at least a Progression-Free Survival (PFS) and / or an Overall Survival (OS) for the patient using the first output vector; and obtain a prediction of the patient response to the agent using said obtained PFS and / or OS; add said obtained PFS and / or OS to the corresponding entry.
[0138] In one embodiment, the step of defining at least one patient feature using the avatar comprises: selecting, among the clinical data, a first subset of clinical data relating to the tumor cells (e.g., type of tumor, genomic data / features of the tumor, treatment previously received by the test patient for the tumor, and the like); selecting, among the clinical data, a second subset of clinical data relating to the patient (e.g., age, sex, patient clinical status - such as the ECOG score -, number of metastasis, location of metastasis and the like); generating a synthetic second subset of clinical data by, for each piece of clinical data comprised into said second subset, comparing it to a corresponding acceptance range of values associated to said piece of clinical data, and whenever the piece of clinical data is not comprised in said acceptance range of values generating a corresponding synthetic data value and recording it into the synthetic second subset of data (i.e., instead of the original piece of clinical data), and otherwise recording the (original) piece of clinical data in the synthetic second subset of data; generating synthetic clinical data for the test patient comprising the first subset of clinical data and the synthetic second subset of clinical data; defining at least one patient feature using said synthetic clinical data.
[0139] The term “acceptance range of values” may apply not only to numerical values but also to labels or categorical features (e.g., sex = male / female, ECOG score categories, tumor type classes, etc.).
[0140] According to one embodiment, each acceptance range of values associated to a specific piece of clinical data is defined from information from at least one clinical study, notably eligibility criteria (inclusion / exclusion criteria).
[0141] The definition of the at least one patient feature using said synthetic clinical data may be performed using at least one LLM, as described above.
[0142] It is further provided an in vitro method for predicting a patient response to an agent; in particular a therapeutic agent selected from the group consisting of: an antibody or antigen-binding fragment thereof, an immunoconjugate, a cytotoxin, a chemotherapeutic agent, a cytokine, an immunosuppressant, an immune stimulator, a lytic peptide, a radioisotope, and antiviral agent, an antiparasitic agent, an antimicrobial agent, a chimeric antigen receptor, a nucleic acid, an expression vector, a cell, a cell lysate, or fraction thereof, and the like.
[0143] There is no known work designing large cohorts of patient avatars (comprising both patient derived cell lines and matched clinical and sequencing data) which mimic clinical population distributions.
[0144] Examples of the method / device for predicting a patient response to an agent are now discussed.
[0145] In examples, the inputted clinical (i.e., digital) data may comprise genomic data, transcriptomic data, key genomic features (e.g., tumor mutational burden), treatment clinical data (e.g., starting date and finishing date of the treatment). Said clinical data may further comprise functional response data obtained from an organoid. Said functional response data are the scores obtained using a response scoring methodology.
[0146] In examples, the method further comprises receiving at least one second ensemble of one or more trained learning models. Each trained leaning model of the second ensemble is configured to receive as input, with respect to said patient, one or more of said at least one patient feature and one or more of said at least one organoid response feature, and to provide a second output vector.
[0147] In such examples, the method further comprises calculating the second output vector by providing said at least one patient feature and said at least one organoid responsefeature as input to said second ensemble of one or more trained learning models and to obtain an Overall Response (OR) for the patient using the second output vector.
[0148] In such examples, the method further comprises receive at least one third ensemble of one or more trained learning models each configured to receive as input, with respect to said patient, one or more of said at least one patient feature and one or more of said at least one organoid response feature, and to provide a third output vector, said third ensemble of one or more trained learning models comprising at least one logistic regression model.
[0149] In such examples, the method further comprises calculating the third output vector representative of the best RECIST response by providing said at least one patient feature and at least one organoid response feature as input to said third ensemble of one or more trained learning models and to obtain a prediction of best RECIST response for the patient using the third output vector.
[0150] In one example, the method further comprises receive at least one third ensemble of one or more trained learning models each configured to receive as input, with respect to said patient, the at least one patient feature and the at least one organoid response feature, and to provide a third output vector representative of the best RECIST response, said third ensemble of one or more trained learning models comprising at least one logistic regression model. Notably, the third output vector may be the best RECIST response itself, therefore the method is configured to provide as input the at least one patient feature and the at least one organoid response feature to one or more of the trained learning models of the third ensemble in order to obtain one or more best RECIST response predictions.
[0151] Examples of models and feature selection are now discussed.
[0152] In examples, the first ensemble of one or more learning models (i.e., the ensemble which is used to predict the PFS and OS of patients) comprises at least one regression model (such as a Lasso or ElasticNet) or Cox model. The at least one Cox model may be any Cox model known in the field of statistic, see https: / / en.wikipedia.org / wiki / Proportional hazards model. Said Cox model may be according to the lifelines library in Python (see https: / / lifelines.readthedocs.io / en / latest / ).Said first ensemble, in some examples, may consist of (exactly) one Cox model. In some examples, said first ensemble may comprise said at least one Cox model and one or more models of the following types: XGBoost (see https: / / en.wikipedia.org / wiki / XGBoost), Random Survival Forest (Ishwaran et al., “Random survival forests”, Ann. Appl. Stat. 2(3): 841-860, 2008), Ridge Cox, and Lasso Cox.
[0153] In examples where the at least one second ensemble of one or more trained learning models are presented, said second ensemble may comprise at least one logistic regression model. In such examples, the (configured to be) received features by the models of the second ensemble, i.e., the one or more of said at least one patient feature, and the one or more of said at least one organoid response feature may be a statistically significant feature. Being statistically significant may be determined according to any known measure in the field of statistics. In examples, each received one or more of said at least one patient feature has a statistical significance (e.g., of a respective hazard ratio) higher than a first threshold. In examples, each received one or more of said at least one organoid response feature has a statistical significance (e.g., of a respective hazard ratio) higher than a second threshold.
[0154] Alternatively, or additionally, the (configured to be) received features by the models of the second ensemble may have a hazard ratio higher than a third threshold. Optionally, said received features with a hazard ratio higher than said third threshold may not be statistically significant.
[0155] In some examples, the method may determine said (configured to be) received features by the models of the second ensemble via a Principal Component Analysis. Alternatively, said features may be determined using a plurality of univariate analysis to obtain the most statistically significant ones.
[0156] This constitutes an improved solution by limiting the input of the trained model to more explicative features. This enables the device to predict the response more computationally efficient without comprising the accuracy.
[0157] In examples where the at least one second ensemble of one or more trained learning models are present, said second ensemble may further comprise at least one or more models of the following types: XGBoost, logistic regression, and random forest.
[0158] In examples, wherein when said second ensemble of one or more trained learning models comprises two or more trained learning models, predictions obtained from each trained learning models are combined according to a voting classifier to obtain the second output vector.
[0159] Examples of training and validation are now discussed.
[0160] Figure 3 presents an example of the training method.
[0161] In examples , the method may further comprise validating of each trained learning model. The validation of each trained learning model may be according to any known method in the field of machine learning. For example, the validation may by 10-fold cross validation, with an 85%- 15% balance between the training and validation data.
[0162] Examples of the population-level response prediction method are now discussed.
[0163] In examples, the (population-level response prediction) method may update the obtained simulated population dataset by applying at least one second ensemble of one or more trained learning models to the respective avatar data and the data of organoid response each entry. As a result of this application of at least one second ensemble, the updated simulated population dataset further comprises OR for respective test patient.
[0164] In such examples, said at least one second ensemble of one or more trained learning models comprises at least one logistic regression model. Each model of said at least one second ensemble is configured to receive as input, with respect to said test patient, one or more of said at least one patient feature and one or more of said at least one organoid response feature, and to provide a second output vector. The OR for the test patient is obtained using the second output vector calculated by said at least one second ensemble of one or more trained learning models receiving as input said at least one patient feature and at least one organoid response feature for the test patient
[0165] In examples, the (population-level response prediction) method may comprise as an initial step (i.e., prior to the step of receiving) an initialisation biological step. Such an initialisation biological step may comprise one or more of: creating a set of organoids to be screened with said agent; screening the created set of organoids with said agent; andobtaining a respective data of organoid response for said test subject.
[0166] Advantageously, the (population-level response prediction) method may comprise the steps of: a) providing one or more organoid(s); b) bringing into contact the one or more organoid(s) with at least one agent; c) determining at least one data of organoid response to the at least one agent; d) optionally comparing said data of organoid response with a reference data; e) optionally determining at least one organoid response feature.
[0167] According to some embodiments, a reference data may correspond to one or more reference values retrieved from organoid(s) comprising healthy cells and / or organoid(s) comprising cells with a pre-determined condition; for example, retrieved from organoid(s) comprising cells with a tumorous condition. According to some embodiments, a reference data may correspond to one or more reference values retrieved from organoid(s) comprising healthy cells.
[0168] According to some embodiments, a reference data may correspond to one or more reference values retrieved from at least one organoid. According to some embodiments, a reference data may correspond to one or more reference values retrieved from a plurality of organoids.
[0169] In examples, the (population-level response prediction) method may further comprise determining a corrector for the OR of the updated population dataset. The determining of the corrector comprises computing a correction factor of the type:6 := E(p) - p = f(p,NPV, PPV) wherein p is an estimator of p, E(. ) denotes an expectation,denotes a function,8 is said correction factor, NPV is the negative predictive value, and PPV is the positive predictive value. In other words, the corrector / correction factor is a bias-adjustment factor (8) that compensates for the difference between the estimated probability and the true probability, taking into account test performance metrics (NPV and PPV).
[0170] Furthermore, p is a ground truth value of the OR of the updated population dataset, and PPV and NPV denote a given positive predictive value and a given negative predictive value, respectively of the at least one second ensemble of one or more trained machine learning models.
[0171] The correction factor may be used therefore to obtain a corrected Overall Response value. In other words, the correction factor 5 is used to adjust the probability estimates obtained from the predictive model, thereby reducing the bias induced by imperfect prediction metrics (NPV, PPV). The corrected probabilities are then employed in the computation of the odds ratio, such that the corrected odds ratio more accurately reflects the true population-level odds ratio. Notably, the correction factor 5 may be applied as an adjustment to the estimated probabilities from which the odds ratio is derived.
[0172] Figure 4 presents an examples schematic workflow of such a correction. The objective of the example is to predict with the greatest accuracy possible the rate of responders in the clinic to a given agent (e.g., a drug or a therapeutic agent) 410 based on tests conducted on a sample of preclinical models. The method outputs a corrected prediction 420 of the fraction of responders to the agent in the clinic. This fraction can be extended to other clinically relevant read-outs such as the progression free survival (PFS).
[0173] Such a correction constitutes an improved prediction. It is known in the prior art that each test on a PDO is not perfect which results in a prediction error when assessing the patient response. Thereby, combining such tests together generates a biased estimator (in the statistical sense) of the population-level response. Such a correction corrects this bias to improve the quality of predictions in clinical trials.
[0174] Accordingly, the disclosure further relates to a method, which is illustrated in figure 12, for predicting an effect of an agent on a patient or a population of patients, and / or a computer-implemented method for training a learning model, which comprises the steps of: a) providing S80 one or more organoid(s) 14; b) bringing S81 into contact the one or more organoid(s) 14 with at least one agent; c) determining S82 at least one data of organoid response to the at least one agent;d) optionally, comparing said data of organoid response with a reference data; e) optionally, determining at least one organoid response feature; f) predicting S83 a patient response to the agent by the individual response prediction method according to the present disclosure, or predicting an effect of an agent on a population of patients using the population-level response prediction method according to the present disclosure.
[0175] In particular, the methods for predicting a patient response to an agent are susceptible to be applied to a method for screening for a therapeutic or prophylactic drug or cosmetic; for example for identifying one or more candidate molecule(s) causing said response, or being correlated to said response, as a potential drug or cosmetic; for example for identifying the one or more candidate molecule(s) within a library of candidate molecules.
[0176] Hence, according to particular embodiments, the disclosure further relates to a method, which is illustrated in figure 14, for screening one or more agent(s) for a therapeutic or prophylactic drug or cosmetic, wherein the method comprises: a) providing S90 at least one organoid 15; b) bringing S91 the at least one organoid 15 in contact with one or more agent(s); c) determining S92 an organoid response data from the at least one organoid; d) predicting S93 a patient response to the agent by the method according to the disclosure, or predicting an effect of an agent on a population of patients by the method according to the disclosure; e) comparing S94 the patient response or the effect of the agent to a reference data; f) identifying S95 the agent as a candidate molecule for a therapeutic or prophylactic drug or cosmetic.
[0177] Advantageously, the step of comparing the patient response or the effect of the agent to a reference data is indicative of an agent as a candidate molecule.
[0178] Hence, according to particular embodiments, the disclosure further relates to a method for screening one or more agent(s) for a therapeutic or prophylactic drug or cosmetic, wherein the method comprises: a) providing a plurality of organoids; b) bringing the plurality of organoids in contact with one or more agent(s); c) determining an organoid response data from the plurality of organoids; d) predicting a patient response to the agent by the method according to the disclosure, or predicting an effect of an agent on a population of patients by the method according to the disclosure; e) comparing the patient response or the effect of the agent to a reference data; f) identifying the agent as a candidate molecule for a therapeutic or prophylactic drug or cosmetic.
[0179] Advantageously, when the methods disclosed herein require a plurality of organoids brought in contact with one or more agents, each organoid may be brought in contact with a different agent, or combination thereof, or a different amount of said agent(s) or combinations thereof.
[0180] Accordingly, the disclosure further relates to a method for treating or preventing or reducing the likelihood of occurrence or re-occurrence of a condition in a patient or a population of patients, comprising a step of administering an agent to the patient or population thereof; wherein the agent is predicted to have an effect according to the above-mentioned methods and / or computer- implemented methods.
[0181] Figure 6 shows an example system 1000 which embodies a device according to any example device explained above. System 1000 (and therefor said example device) is configured to perform any example method explained above. System 1000 is a client computer system, e.g., a workstation of a user, a laptop, a tablet, a smartphone, or a headmounted display (HMD). The client computer of the example comprises a central processing unit (CPU) or processor 1010 connected to an internal communication BUS 1000, a random-access memory (RAM) 1070 also connected to the BUS. The client computer is further provided with a graphical processing unit (GPU) 1110 which isassociated with a video random access memory 1100 connected to the BUS. Video RAM 1100 is also known in the art as frame buffer. A mass storage device controller 1020 manages accesses to a mass memory device, such as hard drive 1030. Mass memory devices suitable for tangibly embodying computer program instructions and data include all forms of non-volatile memory, including by way of example semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM disks 1040. Any of the foregoing may be supplemented by, or incorporated in, specially designed ASICs (application- specific integrated circuits). A network adapter 1050 manages accesses to a network 1060. The client computer may also include one or several VO (Input / Output) devices, like a haptic device 1090 such as cursor control device, a keyboard or the like. A cursor control device is used in the client computer to permit the user to selectively position a cursor at any desired location on display 1080. In addition, the cursor control device allows the user to select various commands, and input control signals. The cursor control device includes a number of signal generation devices for input control signals to system. Typically, a cursor control device may be a mouse, the button of the mouse being used to generate the signals. Alternatively or additionally, the client computer system may comprise a sensitive pad, and / or a sensitive screen.
[0182] The computer program may comprise instructions executable by a computer, the instructions comprising means for causing the above system to perform the method. The program may be recordable on any data storage medium, including the memory of the system. The program may for example be implemented in digital electronic circuitry, or in computer hardware, firmware, software, or in combinations of them. The program may be implemented as an apparatus, for example a product tangibly embodied in a machine- readable storage device for execution by a programmable processor. Method steps may be performed by a programmable processor executing a program of instructions to perform functions of the method by operating on input data and generating output. The processor may thus be programmable and coupled to receive data and instructions from, and to transmit data and instructions to, a data storage system, at least one input device, and at least one output device. The application program may be implemented in a high- level procedural or object-oriented programming language, or in assembly or machinelanguage if desired. In any case, the language may be a compiled or interpreted language. The program may be a full installation program or an update program. Application of the program on the system results in any case in instructions for performing the method. The computer program may alternatively be stored and executed on a server of a cloud computing environment, the server being in communication across a network with one or more clients. In such a case a processing unit executes the instructions comprised by the program, thereby causing the method to be performed on the cloud computing environment.
[0183] Examples of the methods are now discussed.
[0184] The performance of the Cox model used for these examples is presented using the Concordance Index (C-Index) which is a known measure in the field of statistics. A C-index equals to one corresponds to the best model prediction, while when C-index=0.5 it represents a random prediction by the model. The predictions of these examples were assessed as a function of one or more categories of features. These categories include Clinical, Mutations, Functional, and All (of categories). By “Mutations” it is meant genomic features. A list of different features is presented as Table 1. These examples may further use an algorithm for feature selection.Table 1- List of features of different categories
[0185] Figures 7a)-c) presents an example of the PFS prediction. Figure 7a) shows a box-plot performance of the Cox model for predicting the PFS of patients diagnosed with PDAC and colorectal cancer (CRC) and treated by the known drugs. The PFS is shown as a function of different feature categories. The y-axis shows the performance of the Cox model using the C-Index. Figure 7b) shows the hazard ratios (on x-axis) learned by the Cox model for predicting patient PFS when using a combination of genomic, clinical and functional features (shown on y-axis). Closer a hazard ration is to zero, higher importance the respective variable has. Figure 7c) shows a Kaplan-Meier plot of the PFS of patients according to their risk-score given by the ML model. The ML model accurately predicts the risk score for the patients in the population. Patients classified as high risk have significantly shorter PFS than low-risk patients.
[0186] Figures 8a)-c) presents an example of OS prediction for the PDAC and CRC patients discussed above. Figure 8a) shows a box-plot performance of the Cox model for predicting patient OS for treatment with a given drug as a function of different feature categories. Figure 8b) shows the hazard ratios learned by the Cox model for predicting patient OS when using a combination of genomic, clinical and functional features. Figure 8c) shows a Kaplan-Meier plot of the OS of patients according to their risk- score given by the ML model. The ML model accurately predicts the risk score for the patients in thepopulation. Patients classified as high risk have significantly shorter OS than low-risk patients.
[0187] Figures 9a)-b) present an example of OR prediction for the PDAC and CRC patients discussed above. Figure 9a) presents an example of feature importances learned by the Random Forest model for predicting patient OR when using a combination of genomic, clinical and functional features. The importance is presented on y-axis by a weight. Figure 9b) shows a confusion matrix of the OR prediction model when using a combination of genomic, clinical and functional features.
[0188] These examples show that, for PFS and OR prediction, using at least one organoid response feature (“Functional” in Figures 7-9) overperform the other method.This superiority is not visible in the OS prediction as OS is calculated for a long-term treatment and the patient is given multiple different treatments during this period. Using organoid response feature is more powerful for a same treatment.
Claims
CLAIMS1. A device for predicting a patient’s response to an agent by using at least one organoid exposed to said agent, said at least one organoid having been grown from at least one type of cells previously extracted from the patient, the device comprising: at least one input configured to receive at least: o an avatar of the patient (10) and respective data of organoid response (11) for said patient, said avatar comprising biological data and clinical data of said patient; o a first ensemble of one or more trained learning models (12) each configured to receive as input, with respect to said patient, at least one patient feature and at least one organoid response feature, and to provide a first output vector; at least one processor configured to: o defining said at least one patient feature using said avatar and defining said at least one organoid response feature using at least said respective data of organoid response; o calculate the first output vector by providing said at least one patient feature and at least one organoid response feature as input to said first ensemble of one or more trained learning models; o predict at least a Progression-Free Survival (PFS) and / or an Overall Survival (OS) for the patient using the first output vector; and o obtain a prediction of the patient response to the agent using said obtained PFS and / or OS.
2. The device according to claim 1, wherein each of the trained learning models of the first ensemble has been previously trained using a training dataset comprising entries for a plurality of subjects, each entry corresponding to a subject and comprising for said subject: an avatar of the subject; andrespective data of organoid response.
3. The device according to either one of claims 1 or 2, wherein: said at least one input is further configured to receive at least one second ensemble of one or more trained learning models each configured to receive as input, with respect to said patient, one or more of said at least one patient feature and one or more of said at least one organoid response feature, and to provide a second output vector, said second ensemble of one or more trained learning models comprising at least one logistic regression model; and said least one processor is further configured to calculate the second output vector by providing said at least one patient feature and at least one organoid response feature as input to said second ensemble of one or more trained learning models and to obtain an Overall Response (OR) for the patient using the second output vector.
4. The device according to any one of claim 1 to 3, wherein: said at least one input is further configured to receive at least one third ensemble of one or more trained learning models each configured to receive as input, with respect to said patient, one or more of said at least one patient feature and one or more of said at least one organoid response feature, and to provide a third output vector, said third ensemble of one or more trained learning models comprising at least one regression model; and said least one processor is further configured to calculate the third output vector by providing said at least one patient feature and at least one organoid response feature as input to said third ensemble of one or more trained learning models and to obtain a best RECIST response for the patient using the third output vector.
5. The device according to claim 3, wherein: each defined one or more of said at least one patient feature has a statistical significance higher than a first threshold, andeach defined one or more of said at least one organoid response feature has a statistical significance higher than a second threshold. wherein, preferably, when said second ensemble of one or more trained learning models comprises two or more trained learning models, predictions obtained from each trained learning models are combined according to a voting classifier to obtain the second output vector.
6. The device of any of preceding claims, wherein the agent is selected from the group consisting of: a physical treatment, for example radiation therapy, a small molecule, a therapeutic compound, a polypeptide, an antibody or antigen-binding fragment thereof, an immunoconjugate such as an antibody drug conjugate, an aptamer, a nucleic acid, for example a guide RNA or an antisense oligonucleotide, a nanoparticle, an expression vector, a virus, a prokaryotic cell, for example a bacteria, an eukaryotic cell, for example an immune cell, combinations thereof, and / or compositions thereof.
7. The device of any of the preceding claims, wherein the organoid response data is selected from the group consisting of a proteomic feature, a genomic feature, an epigenomic feature, a metabolic feature, a microbiome feature, an imaging feature, a histological feature, viability measurements, growth rate, or a combination thereof.
8. The device of any of the preceding claims, wherein the Progression-Free Survival (PFS) and / or an Overall Survival (OS) is obtained using a pre-trained regression model taking as input the first output vector.
9. A computer- implemented method for predicting a patient response to an agent, the method comprising: receiving (S10) an avatar of the patient (10) and respective data of organoid response (11) for said patient, said avatar comprising biological data and clinical data of said patient;receiving (S20) a first ensemble of one or more trained learning models (12) each configured to receive as input, with respect to said patient, at least one patient feature and at least one organoid response feature, and to provide a first output vector; defining (SI 9) said at least one patient feature using said avatar and defining said at least one organoid response feature using at least said respective data of organoid response; calculating (S30) the first output vector by providing one or more of said at least one patient feature and one or more of at least one organoid response feature as input to said first ensemble of one or more trained learning models; predict (S40) at least a Progression-Free Survival (PFS) and / or an Overall Survival (OS) for the patient using the first output vector; and obtaining (S50) a prediction of the patient response to the agent using said obtained PFS and / or OS.
10. A method for predicting an effect of an agent on a patient or a population of patients, the method comprising: receiving (S60) a test dataset (13) comprising a plurality of entries, each entry corresponding to a test patient and comprising the respective avatar of the test patient and data of organoid response for said test patient; obtaining (S61) a simulated population dataset by applying the computerimplement method according to claim 9 to the respective avatar of the test patient and organoid response data for each entry, the obtained simulated population dataset comprising at least a PFS and / or an OS for each respective test patient.
11. The method according to claim 10, wherein defining at least one patient feature using the avatar of the computer-implement method according to claim 9 comprises: selecting, among the clinical data, a first subset of clinical data relating to the tumor cells;selecting, among the clinical data, a second subset of clinical data relating to the patient; generating a synthetic second subset of clinical data by, for each piece of clinical data comprised into said second subset, comparing it to a corresponding acceptance range of values associated to said piece of clinical data, and whenever the piece of clinical data is not comprised in said acceptance range of values generating a corresponding synthetic data value and recording it into the synthetic second subset of data, and otherwise recording the piece of clinical data in the synthetic second subset of data; generating synthetic clinical data for the test patient comprising the first subset of clinical data and the synthetic second subset of clinical data; defining at least one patient feature using said synthetic clinical data.
12. The method according to either one of claim 10 or 11, further comprising: updating said simulated population dataset by applying at least one second ensemble of one or more trained learning models to the respective avatar of the of the test patient and the data of organoid response for each entry, the updated simulated population dataset further comprising an Overall Response (OR) for each respective test patient; wherein said at least one second ensemble of one or more trained learning models comprises at least one logistic regression model and is configured to receive as input, with respect to each test patient, one or more of said at least one patient feature and one or more of said at least one organoid response feature, and to provide a second output vector, and wherein the Overall Response (OR) for the test patient is obtained using said second output vector.
13. The method according to claim 12, further comprising: determining a corrector for the Overall Response of the updated simulated population dataset; wherein the determining the said corrector comprises computing a correction factor of the type:6 := E(p) - p = f(p,NPV, PPV)wherein: p is an estimator of p,IE (. ) denotes an expectation, denotes a function,6 is said correction factor,NPV is the negative predictive value, andPPV is the positive predictive value, wherein: p is a ground truth value of the OR of updated population dataset, and PPV and NPV denote a given positive predictive value and a given negative predictive value, respectively of the at least one second ensemble of one or more trained machine learning models.
14. A computer-implemented method for training a learning model of the at least first / second / third ensemble of one or more trained learning models according to any of claims 1 to 8; the method comprising: receiving (S70) a training dataset (14) comprising entries, each entry corresponding to a subject and comprising for said subject at least: o an avatar of the subject; and o data of organoid response; training (S71) the learning model using the training dataset.
15. A computer program product comprising instructions which, when the program is executed by a computer, cause the computer to carry out a method according to any of claims 9 to 14.
16. A computer-readable storage medium comprising instructions which, when executed by a computer, cause the computer to carry out a method according to any of claims 9 to 14.
17. An in vitro method for predicting a patient response to an agent, comprising the steps of: a) providing (S80) at least one organoid (15); b) bringing (S81) the at least one organoid (15) in contact with the agent; c) determining (S82) an organoid response data from the at least one organoid; d) predicting (S83) a patient response to the agent by the method according to claim 9, or predicting an effect of an agent on a population of patients by the method according to any of claims 10 to 13.
18. The in vitro method according to the preceding claim, which comprises a step of a) providing a plurality of organoids, and b) bringing the plurality of organoids in contact with the agent, said plurality of organoids being derived from all or part of a group consisting of: cerebral organoids, gastrointestinal organoids, lingual organoids, tooth organoids, thyroid organoids, thymic organoids, testicular organoids, prostate organoids, hepatic organoids, pancreatic organoids, epithelial organoid, lung organoids, kidney organoids, gastruloid organoids, blastoid organoids, endometrial organoids, cardiac organoids, retinal organoids, breast cancer organoids, colorectal cancer organoids, gliobastoma organoids, neuroendocrine tumor organoids, myelin organoids, blood-brain barrier (BBB) organoids.
19. A method for screening one or more agent(s) for a therapeutic or prophylactic drug or cosmetic, wherein the method comprises:• providing (S90) at least one organoid (15);• bringing (S91) the at least one organoid (15) in contact with one or more agent(s);• determining (S92) an organoid response data from the at least one organoid;• predicting (S93) a patient response to the agent by the method according to claim 9, or predicting an effect of an agent on a population of patients by the method according to any of claims 10 to 13;• comparing (S94) the patient response or the effect of the agent to a reference data;• identifying (S95) the agent as a candidate molecule for a therapeutic or prophylactic drug or cosmetic.
Citation Information
Patent Citations
Predicting disease outcomes using machine learned models
US20210366577A1
Large scale organoid analysis
US20220341914A1
Dynamic sampling for tumor features and metabolites
US20240087676A1