Machine learning computer system and method for predicting the efficient of chemical and biological agents

By developing the COLOXIS model, using machine learning to predict the response of GI cancer patients to chemical and biological agents, the problem of lack of effective biomarkers in the prior art was solved, and more accurate treatment options were achieved, improving treatment effect and safety.

CN119998886APending Publication Date: 2025-05-13DEEP RX INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202380062226.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-06-27
Filing Date
2023-06-20
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the treatment of gastrointestinal cancer, the prior art lacks effective biomarkers to predict the efficacy of chemical and biological agents, resulting in patients suffering from unnecessary adverse reactions and poor efficacy.

Method used

A machine learning computer system and method, called the COLOXIS model, was developed to predict the efficacy of GI cancer patients with specific chemical and biological regimens (such as FOLFOX) or a single drug (such as oxaliplatin or bevacizumab). This model uses causal discovery methods to identify cancer driver genes and their target differentially expressed genes, constructing metagenes to reflect the state of the cellular signaling system, thereby predicting drug sensitivity.

Benefits of technology

By accurately predicting the efficacy of drugs, the COLOXIS model helps clinicians personalize treatment plans, improve treatment effects, and reduce the occurrence of adverse reactions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119998886A_ABST
    Figure CN119998886A_ABST
Patent Text Reader

Abstract

The machine learning system employs a causal discovery method to identify genes that cause colorectal cancer when affected by genome alteration. The co-expression pattern between their target differentially expressed genes (DEGs) is found to construct a set of "meta-genes" such that their expression values reflect the status of the cellular signaling system. The tumor is represented using meta-genes as features, and a classification model is trained to predict whether tumor cells of a patient are sensitive to chemotherapy and biological drugs.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Priority declaration

[0002] This application claims priority to U.S. Provisional Patent No. 63 / 355,725, filed on June 27, 2022, entitled "Machine Learning Computer Systems and Methods for Predicting the Efficacy of Chemical and Biological Agents for Treating Diseases Such as Gastrointestinal Cancer." Background Art

[0003] The global incidence of colorectal cancer (CRC) is estimated to be 1.9 million cases in 2020. CRC is the second leading cause of cancer death, accounting for 9.4% of cancer deaths worldwide. More than 60% of CRC patients receive chemotherapy or biologic agents in one or more settings at different stages of the disease course: preoperative neoadjuvant therapy, postoperative adjuvant therapy, and palliative chemotherapy for metastatic patients.

[0004] Commonly used chemotherapeutic and biologic agents for the treatment of CRC and other gastrointestinal (GI) cancers include fluorouracil plus leucovorin (FULV), oxaliplatin, irinotecan, and bevacizumab (Bev). Different combinations of these agents are used in clinical practice: FULV; FULV plus oxaliplatin (FOLFOX); FULV plus irinotecan (FOLFIRI); FOLFOX plus Bev; and FOLFIRI plus Bev. For neoadjuvant therapy, FOLFOX is the most used regimen; for adjuvant therapy, FOLFOX is the standard of care; for metastatic patients, FOLFOX + / - Bev and FOLFIRI + / - Bev are used.

[0005] Due to the high incidence of adverse effects associated with each of the above agents, patients treated with chemotherapy combinations are likely to be treated with drugs that provide no benefit to the patient but cause significant adverse effects. In current clinical practice, there are no biomarkers that can predict the efficacy of the above agents, either alone or in combination. Therefore, the precise selection of effective drugs while avoiding non-beneficial agents is a key issue in the care of cancer patients. Summary of the invention

[0006] In one general aspect, the present invention relates to computer systems and methods for training models through machine learning to predict the efficacy of biologics such as regimens (such as FOLFOX) or single drugs (such as oxaliplatin or bevacizumab) in treating patients with GI cancers such as esophageal cancer, gastric cancer, and colorectal cancer. The machine learning model, referred to herein as the "COLOXIS" model, can also be used to predict the sensitivity of GI cancer patients to regimens / drugs (where the regimens / drugs are commonly used). This can enable clinicians to form individualized treatment plans by selecting effective drugs to treat GI cancer patients. These and other benefits that can be achieved by embodiments of the present invention will be apparent from the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Various embodiments of the present invention are described herein by way of example with reference to the following drawings.

[0008] Figure 1 is a flowchart describing a method for training a machine learning model for predicting the efficacy of chemical and biological agents in treating gastrointestinal cancer according to various embodiments of the present invention.

[0009] Figure 2 is a graph showing how a model trained according to an embodiment of the present invention can predict the outcome of CRC patients treated with FOLFOX. Figure 2 Kaplan–Meir curves for predicted responders and non-responders in TCGA CRC patients treated with FOLFOX are shown.

[0010] Figure 3 Included is a graph showing how COLOXIS characteristics correlate with outcomes in patients treated with FULV alone. Figure 3 Figure 2. Kaplan–Meir curves for COLOXIS+ and COLOXIS- patients who received FULV alone in the C-07 trial.

[0011] Figure 4A Relapse-free survival (RFS) was compared between patients treated with FULV and FOLFOX in the COLOXIS+ design. Cox proportional hazards p-values ​​and interaction p-values ​​for the COLOXIS signature are shown. Figure 4B The RFS of patients treated with FULV in the group designed as COLOXIS- was compared with that of patients treated with FOLFOX. Figure 4C Multivariate Cox proportional hazards analysis in the group designed as COLOXIS+ is shown.

[0012] Figure 5ARFS was compared between patients treated with FLOX / FOLFOX (Bev=0) and FOLFOX+bevacizumab (Bev=1) in the group designed as COLOXIS+. Figure 5B Comparison of RFS between patients treated with FLOX / FOLFOX and patients treated with FOLFOX+Bev in the group designed as COLOXIS- is shown. Figure 5C Multivariate Cox proportional hazards analysis of the effect of Bev is shown.

[0013] Figure 6 is a flow chart of a method according to various embodiments of the present invention for using Figure 1 A machine learning model was developed in order to predict drug response in patients with GI cancer.

[0014] Figure 7 is a diagram of a computer system according to various embodiments of the present invention. Specific implementation plan

[0015] Commonly used chemical and biological agents for the treatment of colorectal cancer (CRC) and other gastrointestinal (GI) cancers include fluorouracil, folinic acid, oxaliplatin and bevacizumab. Different combinations of these drugs are widely used as different regimens to treat CRC patients, but there are no perfect biomarkers or decision support systems to enable clinicians to select the most effective regimen for a given patient among multiple candidate regimens. This patent application describes an artificial intelligence system, the COLOXIS system or model, to support the selection of effective drugs to form the best regimen for treating CRC patients or other GI cancers (as the case may be). The machine learning system uses a causal discovery method to identify cancer driver genes and their target differentially expressed genes (DEGs) involved in the disease development process of CRC (or other GI cancers, as the case may be). It was found that the co-expression patterns between these target DEGs constructed a set of "metagenes" so that their expression values ​​reflect the state of the cell signaling system. Using metagenes as features representing tumors, a classification model (COLOXIS) is trained to predict whether tumor cells are sensitive to, for example, oxaliplatin and bevacizumab. A validation study using data from a large-scale phase III clinical trial showed that the COLOXIS model can predict patients' responses to single drugs (oxaliplatin and bevacizumab) and combination regimens (such as FOLFOX, a chemotherapy regimen consisting of folinic acid (leucovorin, FOL), fluorouracil (5-FU, F) and oxaliplatin (oxaliplatin, OX)). Accurate prediction of the efficacy of these drugs can help clinicians and CRC patients make decisions.

[0016] Figure 1is a flowchart depicting a method 10 for training a machine learning model 12 according to various embodiments of the present invention, wherein the machine learning model 12 is used to predict the efficacy of chemical and biological agents in treating CRC. The development of a model (e.g., the "COLOXIS model" described above) can consist of four main stages, which, in various embodiments, are: (I) modeling disease mechanisms that affect heterogeneous responses to drugs, (II) discovering transcriptome patterns that reflect the disease mechanisms of cancer cells, (III) training a model for predicting drug sensitivity based on the disease mechanisms of cancer cells, and finally (IV) validating the predictive model. Method 10 can be performed in part or in whole using a computer system, for example in combination with the following Figure 7 A computer system 100 is described.

[0017] Phase I involves modeling the heterogeneous disease mechanisms of CRC using causal discovery approaches. For every drug used to treat CRC, less than 40% of patients respond. The heterogeneous response of cancer cells to drugs is due to differences in disease mechanisms. That is, tumors have different disease mechanisms because different driver somatic genomic alterations (SGAs) in individual tumors interfere with signaling pathways, resulting in different responses to drugs. Understanding the disease mechanisms of cancer cells in tumors will enable prediction of drug responses in tumor cells.

[0018] In order to study the common disease mechanisms of CRC, a tumor-specific causal inference (TCI) algorithm 14 can be applied to CRC genomic data 16 and CRC transcriptome data 18. TCI is a Bayesian causal discovery algorithm invented by Dr. Xinghua Lu, Dr. Gregory Cooper, et al. from the University of Pittsburgh, and described in U.S. Patent Application No. 16 / 349,192, with publication number 2019 / 0287651A1, the entire contents of which are incorporated herein by reference. In various embodiments, 290 CRC tumors depicted by the Cancer Genome Atlas (TCGA) can be used as CRC genomic and transcriptome data 16, 18 to identify the driving SGA 20 and their target differentially expressed genes (DEGs) 22 in individual tumors. TCI searches for the SGA that is most likely to cause the molecular phenotype (e.g., DEG event) observed in the tumor in a specific tumor. In the experiment, the driving SGA in the individual tumor identified by the TCI algorithm was compiled. Among more than 10,000 SGA interference genes, 37 genes were designed as the main driving factors of the CRC cohort. The discovery of driver factors significantly narrowed the number of candidate driver genes. In addition, TCI analysis identified 2,691 genes regulated by these driver SGAs (i.e., target DEGs). The identification of driver SGAs 20 and their target DEGs 22 enabled one to infer the disease mechanisms (status of signaling systems) of cancer cells based on genomic and transcriptomic data from tumors.

[0019] Phase II involves discovering transcriptome patterns that reflect the disease mechanisms of cancer cells. The expression state of a gene reflects the state of the signaling pathway that regulates its expression, which can be used to infer the state of the cell signaling pathway. However, the expression values ​​of a single gene in different tumors are highly variable. Therefore, single gene expression is an unreliable marker for inferring the state of a signaling pathway. Since signaling pathways typically regulate a group of genes (gene modules) in cells, the expression state of gene modules is a better biomarker for inferring the state of signaling pathways. Identifying the gene expression modules in DEG 22 found in Phase I will be able to infer the state of the main signaling pathways that are disturbed in CRC, which can be further used to predict drug response.

[0020] In various embodiments, a database 24 of CRC transcriptome data is used to extract the target DEG in step 26. In various embodiments, the Gene Expression Omnibus (GEO) with transcriptome data can be used as database 24. The inventor's experiment collected 4,199 CRC tumors from the GEO database. In step 26, the expression values ​​of 2,691 target DEGs (see box 22) in these 4,199 CRC tumors are extracted in box 28. In various embodiments, the extraction step 26 may involve a series of consensus clustering analyses to determine a set (e.g., 10 to 50, including 10 and 50, preferably about 15) of co-expression modules (metagenes) that exhibit clear co-expression patterns and provide strong signals related to drug response. Consistency clustering is a method of aggregating (possibly conflicting) results from multiple clustering algorithms. It refers to a situation where multiple different (input) clusters have been obtained for a particular data set, and it is desired to find a single (consistent) cluster, such as a target DEG, that is more suitable in some sense than the existing clusters. Examples of suitable extraction techniques are described in Monti S., Tamyo P, Mesirov J, Golub T., "Consensus Clustering: A Resampling-Based Method for Class Discovery and Visualization of Gene Expression Microarray Data," Machine Learning 2003; 52: 91-118, which is incorporated herein by reference. The Monti consensus clustering algorithm described in this article is used to determine the number of clusters K. Given a data set with a total number of points to be clustered N, the algorithm works by resampling and clustering the data, and for each K, an N×N consistency matrix is ​​calculated, where each element represents the number of times two samples are clustered together. A completely stable matrix will consist entirely of 0s and 1s, representing all pairs of samples that are always clustered together or not clustered together in all resampling iterations. The relative stability of the consistency matrix can be used to infer the optimal K.

[0021] In step 30, a gene set variation analysis (GSVA) ​​method can be used to estimate the expression values ​​of the metagenes in each tumor. The expression values ​​(box 32) can then be used as features representing the cell state of cells in the tumor. GSVA calculates sample gene set enrichment scores based on functions of genes in and out of the gene set, similar to competitive gene set testing. In addition, it estimates the changes in gene set enrichment over samples independently of any class labels. Conceptually, this method can be understood as a change in the coordinate system of gene expression data from genes to gene sets. This transformation facilitates the post-hoc construction of pathway-centric models, such as differential pathway activity identification or survival prediction. An example of a suitable GSVA method is described in Hanzelmann S, Castelo R, Guinney J., "GSVA: gene set variation analysis for microarray and RNA-seq data," BMC Bioinformatics 2013; 14: 7, which is incorporated herein by reference.

[0022] Phase III includes training for predicting the classification model of drug sensitivity. From the collection of data sets, such as GEO data sets (GSE19860, GSE28702, GSE72970, GSE69657, GSE104645), the cohort of mCRC patients can be identified as the discovery data set at frame 36 in step 34, wherein the patient in the cohort has been treated with, for example, FOLFOX regimen. The therapeutic response is measured with RECIST (solid tumor response evaluation criteria) score. In various embodiments, for each patient, a feature vector composed of the GSVA scores of a relatively small amount of metagenes (e.g., 10-50 metagenes) can be constructed. In various embodiments, 15 metagenes can be used. Based on RECIST scores, class labels can be assigned to each patient in the training data set as "responders" (CR and PR) or "non-responders" (SD or PD). Based on the dimensionality reduction training data 36, ​​COLOXIS model 12 can be trained in step 38. In various embodiments, COLOXIS model 12 is a regularized logistic regression model. A regularized logistic regression model (or other machine learning classifiers as appropriate) can be trained in step 38 to predict the response to FOLFOX in these mCRC patients. As further explained below, the performance of model 12 can be evaluated by cross-validation and external validation experiments. In other embodiments, in addition to the regularized logistic regression model, other machine learning models can also be trained. In other embodiments, the machine learning model 12 trained from the data set 36 in step 38 can be, for example, an integration of a deep learning artificial neural network, a support vector machine and / or a decision tree or a machine learning model.

[0023] Phase IV includes validating the predictive accuracy of the COLOXIS Model 12 at step 40. In various embodiments, the validation is performed in the adjuvant treatment of CRC. The clinical utility of the COLOXIS Model 12 can be validated using data from TCGA and from two Phase III clinical trials.

[0024] The inventors evaluated the COLOXIS model 12 to predict the response of CRC patients to FOLFOX. In the TCGA study, genomic and clinical data of 87 patients who were treated with the FOLFOX regimen and whose overall survival results were known were collected. The gene expression data of the patients were converted using the GSVA algorithm to project the patients into a 15-element gene space, and the COLOXIS model 12 was applied to predict whether the patients responded to FOLFOX (referred to as the COLOXIS+ and COLOXIS- groups). The survival of the two groups of patients was compared. Figure 2 As shown, the results showed that patients assigned to the COLOXIS+ group had significantly better overall survival.

[0025] The inventors also evaluated the prognostic value of the COLOXIS signature. The COLOXIS model 12 was applied to a cohort of 1,285 patients with colon cancer who received postoperative adjuvant therapy and was studied in two Phase III clinical trials (the C-07 and C-08 trials) conducted by the National Surgical Adjuvant Breast and Bowel Project (NSABP), a nonprofit clinical trial management organization. The C-07 trial compared the efficacy (in terms of preventing recurrence) of two chemotherapy combinations: (1) fluorouracil plus folinic acid (FULV); and (2) FULV plus oxaliplatin (FOLFOX). The C-08 trial compared the efficacy (in terms of preventing recurrence) of two chemotherapy combinations: (1) FOLFOX; and (2) FOLFOX+Bev.

[0026] The prognostic value of the COLOXIS Model 12 was evaluated by comparing recurrence-free survival (RFS) in the COLOXIS+ group with that in the COLOXIS- group in patients who received FULV alone. Figure 3 As shown, the COLOXIS signature was significantly associated with the patient's prognosis (HR: 1.52, 95% CI = 1.07 to 2.15, p = 0.017).

[0027] The inventors also evaluated the role of the COLOXIS model in predicting the benefit of oxaliplatin. Among 1,065 patients treated with fluorouracil plus leucovorin (FULV) (N=421) and FOLFOX (N=644), 526 patients were predicted to benefit from the oxaliplatin-containing regimen (called COLOXIS+), and 539 patients were predicted not to benefit from the oxaliplatin-containing regimen (COLOXIS-). The predictive value of the COLOXIS model was tested by comparing the response of the COLOXIS+ group to oxaliplatin treatment with the COLOXIS- group. Figure 4A -C shows that COLOXIS+ patients benefited from oxaliplatin (HR=0.65, 95%CI=0.48-0.89, P=0.0065, int P=0.03), but COLOXIS- patients did not benefit (COLOXIS-HR=1.08, 95%CI=0.77-1.52, P=0.65). Therefore, the COLOXIS signature can predict the benefit of oxaliplatin.

[0028] The inventors also evaluated the COLOXIS model in predicting the benefit of FOLFOX+Bev. Of the 644 patients treated with FOLFOX and the 219 patients treated with FOLFOX+Bev, 491 patients were assigned to the COLOXIS+ group and 372 patients were assigned to the COLOXIS- group. The predictive value of the COLOXIS model was tested by comparing the response of the COLOXIS+ group to the COLOXIS- group when bevacizumab was added to the treatment. Figure 5C As shown, the COLOXIS+ group benefited significantly after the addition of bevacizumab to FOLFOX (HR=0.58, 95%CI=0.36-0.94, p=0.025, int p=0.101), but the COLOXIS- group did not (HR=1.02, 95%CI=0.64-1.63, p=0.94). Therefore, the COLOXIS signature can predict the benefit of bevacizumab.

[0029] Once the model 12 is trained and validated, it can be used as a diagnostic tool to predict a patient's response to an agent. For example, if trained as described above, the COLOXIS model can be trained to predict whether a CRC patient will benefit from FOLFOX, oxaliplatin, or bevacizumab. Figure 66 is a flow chart of a method for using a model as a decision support tool according to various embodiments of the present invention. In step 60, a patient visits a medical provider and is diagnosed with GI cancer, which results in a need to make a decision about whether the patient's tumor cells are sensitive to oxaliplatin, bevacizumab or FOLFOX regimens. In step 62, a tumor tissue sample from the patient is collected. For example, the sample can be collected by biopsy or surgery. Next, in step 64, a transcriptome analysis (RNA detection and quantification) of the sample collected in step 62 is performed. In step 64, any suitable technology / platform designed to analyze the transcriptome of the sample, such as gene expression arrays or next generation sequencing, can be used.

[0030] In step 66, extract the Figure 1 Quantification of expression of genes involved in tumor formation identified in Phase I of the method shown and described (see Figure 1 In step 14), data obtained using different platforms can be converted. Finally, in step 68, the transcriptome of the tumor cell can be mapped to a reduced-dimensional metagene space. For example, as described above, the metagene space can be a 15-metagene space.

[0031] At step 70, the metagene expression of the tumor is used as input to the COLOXIS model, which calculates a probability or binary call to indicate whether tumor cells from the patient will respond to FOLFOX, oxaliplatin and / or bevacizumab. At step 72, the predictions of the COLOXIS model 12 can be used by clinicians to make treatment decisions for the patient.

[0032] Figure 7 1 is a diagram of a computer system 100 that can be used to implement the above-described embodiments. The illustrated computer system 100 includes one or more processors 102 and one or more memory units 104 that communicate via a data bus and / or an electronic data network. For simplicity, Figure 7 Only one processor 102 and one memory unit 104 are shown. The memory 104 may store various software modules 106, 108, 110, and 112, which include software or computer instructions executed by the processor 102. For example, the TCI module 106 may include a program for executing Figure 1 DEG extraction module 108 may include software for performing TCI analysis in step 14; Figure 1 The software for target DEG extraction in step 26; the metagene learning and GSVA module 110 may include software for Figure 1 Step 30 of performing metagene learning and GSVA software; and the machine learning training module 112 may include software for Figure 1Step 38 trains the software of the model 12 .

[0033] according to Figure 6 , Figure 7 The computer system 100 can also be used to make patient predictions. The memory 104 can store software for TCI analysis in step 66; for mapping the transcriptome of tumor cells to a reduced-dimensional metagene space in step 68; and for inputting the reduced-dimensional metagene space into a trained, validated COLOXIS model in step 70 to obtain drug response predictions.

[0034] Processor 102 may include one or more CPU cores, GPU cores, and / or AI accelerator cores. Memory 104 may include primary computer memory, such as read-only memory (ROM) and / or random access memory (e.g., RAM). Memory 104 may also include secondary memory, such as, for example, a disk or optical drive or flash memory. Software modules 106, 108, 110, and 112 may be implemented in computer software using any suitable computer programming language, such as .NET, C, C++, or Python, and using traditional, functional, or object-oriented techniques. The programming language for instructions for computer software and other computer implementations may be translated into machine language by a compiler or assembler prior to execution, and / or may be translated directly by an interpreter at runtime. Examples of assembly languages ​​include ARM, MIPS, and x86; examples of high-level languages ​​include Ada, BASIC, C, C++, C#, Python, R, COBOL, Fortran, Java, Lisp, Pascal, ObjectPascal, Haskell, ML; and examples of scripting languages ​​include Bourne shell script, JavaScript, Python, Ruby, Lua, PHP, and Perl. In Figure 1 The various data used in the methods may be stored in primary, secondary, tertiary and / or offline (e.g., cloud) storage.

[0035] Computer system 100 may be of various types, including a workstation, a server, a computing cluster, a blade server, a server farm, or any other data processing system or computing device. Computer system 100 may be, for example, a server on a cloud network. Due to the ever-changing nature of computers and networks, Figure 7 The description of the computer system 100 depicted in the drawings is intended only as a specific example for illustrating some embodiments. Many other configurations of the computer system 100 may have more Figure 7 The computer systems shown may have additional or fewer components.

[0036] Therefore, in various embodiments, the present invention provides a new causal analysis method for searching for genes involved in the disease development of CRC, which is used as a more informative feature for detecting heterogeneous responses to drugs. The present invention also provides a new feature construction method (discovery of metagenes), which lays the foundation for building a stable prediction model. The COLOXIS model can predict the prognosis (outcome) of patients who only receive FULV treatment. This can be used to predict who has better results if treated with FULV. The COLOXIS prediction can also be a prognostic result for patients treated with FOLFOX. This can be used to predict the results of patients receiving FOLFOX treatment.

[0037] The COLOXIS model also predicted the benefit of oxaliplatin treatment for CRC patients. This could enable clinicians to determine whether CRC patients should be treated with oxaliplatin. Accurate decision-making would improve treatment efficacy and prevent overtreatment with oxaliplatin, which would only lead to adverse reactions.

[0038] The COLOXIS model also predicts the benefit of FOLFOX+Bev in the adjuvant setting of CRC. This could enable clinicians to include bevacizumab in the adjuvant setting of a subset of patients to increase the therapeutic effect.

[0039] The COLOXIS model can also be used to predict the sensitivity of gastrointestinal (GI) cancers such as esophageal cancer, gastric cancer, and colorectal cancer, where the regimens and drugs are commonly used, to the above regimens (FOLFOX) or single drugs (oxaliplatin or bevacizumab). This will enable clinicians to form personalized treatment plans by selecting effective drugs to treat GI cancer patients.

[0040] Thus, in one general aspect, the present invention is a computer-implemented system and method for training machine learning to predict the efficacy of biologics for gastrointestinal cancer, such that the output of the model can be used by clinicians to determine treatment for new patients with gastrointestinal cancer. The method includes calculating, by a computer system 100, expression values ​​of a set of metagenes from target differentially expressed genes (DEGs) for gastrointestinal cancer, wherein the set of metagenes exhibits a co-expression pattern associated with the response of gastrointestinal cancer patients to biologics, wherein the computer system 100 includes one or more processors 102 that execute instructions stored in a computer memory 104. The method also includes training a model by machine learning by the computer system to predict the efficacy of biologics for gastrointestinal cancer, such that the output of the model can be used by clinicians to determine treatment for new gastrointestinal cancer patients.

[0041] In various embodiments, the method also includes identifying the target DEG of gastrointestinal cancer by a computer system. The identification can be performed using a tumor-specific causal inference algorithm applied to genomic data and transcriptome data. The method can further include extracting the metagene by using cluster analysis to identify the reduced metagene showing a clean co-expression pattern and a strong drug response signal by the computer system. The reduced metagene includes ten to twenty (including ten and twenty) metagenes, preferably about 15 metagenes. Gene set variation analysis (GSVA) ​​can be used to calculate expression values.

[0042] In various embodiments, the model includes a classifier, such as a logistic regression model.

[0043] In various embodiments, the gastrointestinal cancer comprises colorectal cancer.The biologic may comprise, for example, FOLFOX, oxaliplatin, and / or bevacizumab.

[0044] In various embodiments, the GSVA scores of the set of genes for the patient cohort are feature vectors used to train the model by machine learning. In this case, each patient in the cohort can be labeled as a responder or non-responder to a biologic.

[0045] In various embodiments, the method further includes, after training the model: collecting tumor tissue samples from new patients diagnosed with gastrointestinal cancer; analyzing the transcriptome of the sample; mapping the transcriptome to the metagene; and using the model to classify the efficacy of the biological agent on the new patient based on the mapping of the new patient's transcriptome to the metagene.

[0046] The examples given here are intended to illustrate potential and specific embodiments of the present invention. It is understood that these embodiments are primarily intended to illustrate the present invention to those of ordinary skill in the art. The specific aspects of the embodiments do not necessarily limit the scope of the present invention. In addition, it should be understood that the drawings and descriptions of the present invention have been simplified to illustrate relevant elements in order to clearly understand the present invention, while other elements have been omitted for clarity. Although various embodiments have been described herein, it should be clear that various modifications, changes and adjustments to these embodiments can be conceived by those skilled in the art, and at least some advantages are obtained. Therefore, the disclosed embodiments are intended to include all such modifications, changes and adjustments without departing from the scope of the embodiments set forth herein.

Claims

1. A computer system comprising: one or more processor cores; and a computer memory in communication with the one or more processor cores, wherein the computer memory stores instructions that, when executed by the one or more processor cores, cause the one or more processor cores to: calculating expression values ​​of a set of metagenes from target differentially expressed genes (DEGs) of gastrointestinal cancer, wherein the set of metagenes exhibits a co-expression pattern associated with the response of the gastrointestinal cancer patient to a biologic agent; and The model is trained by machine learning to predict the efficacy of the biologic on the gastrointestinal cancer, such that the output of the model can be used by clinicians to determine treatment for new patients with the gastrointestinal cancer. 2 . The computer system of claim 1 , the computer memory storing instructions that, when executed by the one or more processor cores, cause the one or more processor cores to identify a target DEG for the gastrointestinal cancer. 3 .

3. A computer system according to claim 2, wherein the computer memory stores instructions that, when executed by the one or more processor cores, cause the one or more processor cores to use a tumor-specific causal inference algorithm applied to genomic data and transcriptomic data to identify the target DEG of the gastrointestinal cancer.

4. The computer system of claim 2, the computer memory storing instructions that, when executed by the one or more processor cores, cause the one or more processor cores to extract the set of metagenes by identifying a set of metagenes that exhibit clean co-expression patterns and a reduction in strong drug response signals using cluster analysis.

5. The computer system according to claim 4, wherein: The reduced set of metagenes includes ten to twenty metagenes, inclusive.

6. The computer system of claim 4, the computer memory storing instructions that, when executed by the one or more processor cores, cause the one or more processor cores to calculate the expression values ​​using gene set variation analysis (GSVA).

7. The computer system according to claim 1, wherein: The model includes a classifier.

8. The computer system according to claim 7, wherein: The classifier includes a logistic regression model.

9. The computer system according to claim 1, wherein: The gastrointestinal cancer includes colorectal cancer.

10. The computer system according to claim 9, wherein: The biologics include FOLFOX.

11. The computer system according to claim 9, wherein: The biologics include oxaliplatin.

12. The computer system according to claim 9, wherein: The biologics include bevacizumab.

13. The computer system according to claim 6, wherein: The GSVA scores of the set of metagenes for the patient cohort are the feature vectors used to train the model by machine learning.

14. The computer system of claim 13, wherein: Each patient in the cohort was labeled as a responder or non-responder to the biologic.

15. A method comprising: calculating, by a computer system comprising one or more processors, expression values ​​of a set of metagenes from target differentially expressed genes (DEGs) of gastrointestinal cancer, wherein the set of metagenes exhibits a co-expression pattern associated with a response of a gastrointestinal cancer patient to a biologic agent; and The computer system trains a model through machine learning to predict the efficacy of the biological agent on the gastrointestinal cancer, so that the output of the model can be used by clinicians to determine the treatment of new patients with the gastrointestinal cancer. 16 . The method of claim 15 , further comprising identifying, by the computer system, a target DEG for the gastrointestinal cancer.

17. The method according to claim 16, wherein: Identification of the target DEGs involved the use of tumor-specific causal inference algorithms applied to genomic and transcriptomic data.

18. The method of claim 16, further comprising extracting, by the computer system, the set of metagenes by identifying a set of metagenes that exhibit clean co-expression patterns and a reduction in strong drug response signals using cluster analysis.

19. The method according to claim 18, wherein: The reduced set of metagenes includes ten to twenty metagenes, inclusive.

20. The method according to claim 18, wherein: Calculating the expression values ​​includes using Gene Set Variant Analysis (GSVA).

21. The method according to claim 15, wherein: The model includes a classifier.

22. The method according to claim 21, wherein: The classifier includes a logistic regression model.

23. The method according to claim 22, wherein: The gastrointestinal cancer includes colorectal cancer.

24. The method according to claim 23, wherein: The biologics include FOLFOX.

25. The method according to claim 23, wherein: The biologics include oxaliplatin.

26. The method of claim 23, wherein: The biologics include bevacizumab.

27. The method according to claim 20, wherein: The GSVA scores of the set of metagenes for the patient cohort are the feature vectors used to train the model by machine learning.

28. The method according to claim 27, wherein: Each patient in the cohort was labeled as a responder or non-responder to the biologic.

29. The method according to claim 15, further comprising, after training the model: Collect tumor tissue samples from patients newly diagnosed with gastrointestinal cancer; analyzing the transcriptome of the sample; mapping the transcriptome to the set of metagenes; as well as Based on the mapping of the transcriptome of the new patient to the set of metagenes, the model is used to classify the efficacy of the biologic for the new patient.

Citation Information

Patent Citations

  • Identification of instance-specific somatic genome alterations with functional impact

    US11990209B2

  • Identification of instance-specific somatic genome alterations with functional impact

    US20190287651A1