Machine learning computer systems and methods for predicting the efficacy of chemical and biological agents for treating diseases such as gastrointestinal cancer

The COLOXIS model addresses the lack of biomarkers by predicting drug efficacy in gastrointestinal cancers, allowing for personalized treatment decisions that improve outcomes and reduce side effects.

JP2025531946APending Publication Date: 2025-09-25DEEP RX INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025536420
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-06-27
Filing Date
2023-06-20
Publication Date
2025-09-25

AI Technical Summary

Technical Problem

Current clinical practices lack biomarkers to predict the effectiveness of chemical and biological agents for treating gastrointestinal cancers, leading to patients receiving drugs that cause significant side effects rather than benefits.

Method used

A machine learning model, COLOXIS, is developed to predict the efficacy of drug regimens and individual drugs like FOLFOX, oxaliplatin, and bevacizumab by identifying cancer driver genes and their target differentially expressed genes, constructing metagenes, and training a classification model to predict drug sensitivity.

Benefits of technology

The COLOXIS model accurately predicts patient response to drug regimens, enabling personalized treatment decisions that enhance treatment efficacy and reduce side effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025531946000001_ABST
    Figure 2025531946000001_ABST
Patent Text Reader

Abstract

The machine learning system employs a causal search approach to identify genes that, when affected by genomic alterations, contribute to colorectal cancer. Co-expression patterns among these target differentially expressed genes (DEGs) are discovered to construct a set of "metagenes," whose expression values ​​reflect the state of cell signaling systems. Using the metagenes as tumor signatures, a classification model is trained to predict whether a patient's tumor cells are sensitive to chemotherapy and biologic drugs.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] <Priority Claim> This application claims priority to U.S. Provisional Patent Application No. 63 / 355,725, filed June 27, 2022, entitled "MACHINE-LEARNING COMPUTER SYSTEMS AND METHODS FOR PREDICTING EFFICACY OF CHEMICAL AND BIOLOGICAL AGENTS FOR TREATING DISEASES, SUCH AS GASTROINTESTINAL CANCERS."

[0002] There are an estimated 1.9 million incident cases of colorectal cancer (CRC) worldwide in 2020. CRC is the second leading cause of cancer death, accounting for 9.4% of cancer deaths worldwide. At various stages of the disease course, more than 60% of CRC patients receive chemical or biologic agents in treatment in one or more settings, such as neoadjuvant therapy before surgery, adjuvant therapy after surgery, and palliative chemotherapy for metastatic patients. [Background technology]

[0003] Common chemical and biological agents for treating not only CRC but also other gastrointestinal (GI) cancers include fluorouracil plus leucovorin (FULV), oxaliplatin, irinotecan, and bevacizumab (Bev). Various combinations of the above drugs are used in clinical practice: FULV, FULV plus oxaliplatin (FOLFOX), FULV plus irinotecan (FOLFIRI), FOLFOX plus Bev, and FOLFIRI plus Bev. In neoadjuvant therapy, FOLFOX is the most commonly used regimen. In adjuvant therapy, FOLFOX is the standard of care. In metastatic patients, FOLFOX plus / with Bev and FOLFIRI plus / with Bev are used.

[0004] Due to the high incidence of side effects associated with each of the above drugs, patients treated with chemotherapy combinations are likely to be treated with drugs that cause significant side effects rather than benefit the patient. In current clinical practice, there are no biomarkers that can predict the effectiveness of these drugs, either individually or in combination. Thus, accurately selecting effective drugs while avoiding non-beneficial drugs is an essential issue in the treatment of cancer patients. Summary of the Invention

[0005] In one general aspect, the present invention is directed to a computer system and method for training a model to predict, through machine learning, the efficacy of a regimen (such as FOLFOX) or a biologic agent, such as a single drug (such as oxaliplatin or bevacizumab), in treating patients with GI cancers, including esophageal, gastric, and colorectal cancers. Furthermore, a machine learning model, referred to herein as the "COLOXIS" model, can be used to predict the sensitivity of GI cancer patients to commonly used regimens / drugs. This may enable one or more clinicians to create personalized regimens by selecting effective drugs to treat GI cancer patients. These and other benefits that can be realized through embodiments of the present invention will become apparent from the description below. [Brief explanation of the drawings]

[0006] Various embodiments of the present invention are herein described by way of example in conjunction with the following figures:

[0007] [Figure 1] FIG. 1 is a flowchart illustrating a process for training a machine learning model to predict the efficacy of chemical and biological agents for treating gastrointestinal cancers, according to various embodiments of the present invention.

[0008] [Figure 2]Figure 2 is a graph showing how a model trained according to an embodiment of the present invention can predict the outcome of CRC patients treated with FOLFOX. Kaplan-Meier curves of predicted responders or non-responders among TCGA CRC patients treated with FOLFOX are shown.

[0009] [Figure 3] Figure 3 contains a graph showing how the COLOXIS signature correlates with prognosis in patients treated with FULV alone. Figure 3 includes Kaplan-Meier curves for COLOXIS+ and COLOXIS- patients treated with FULV alone in the C-07 trial.

[0010] [Figure 4A] Figure 4A compares recurrence-free survival (RFS) between patients treated with FULV and those treated with FOLFOX in the group designated COLOXIS+. Cox proportional hazards p-values ​​and interaction p-values ​​for the COLOXIS signature are shown. [Figure 4B] Figure 4B compares the RFS of patients treated with FULV versus those treated with FOLFOX in the group designated COLOXIS-. [Figure 4C] Figure 4C shows the multivariate Cox proportional hazards analysis in the group designated COLOXIS+.

[0011] [Figure 5A] Figure 5A compares RFS between patients treated with FLOX / FOLFOX (Bev=0) and FOLFOX+bevacizumab (Bev=1) in the group designated COLOXIS+. [Figure 5B] Figure 5B shows a comparison of RFS between patients treated with FLOX / FOLFOX and those treated with FOLFOX+Bev in the group designated as COLOXIS-. [Figure 5C] Figure 5C shows a multivariate Cox proportional hazards analysis of the effect of Bev.

[0012] [Figure 6] FIG. 6 is a flowchart of a process according to various embodiments of the present invention for predicting drug response by GI cancer patients using the machine learning model developed in FIG.

[0013] [Figure 7] FIG. 7 is a schematic diagram of a computer system according to various embodiments of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0014] Common chemical and biological agents for treating colorectal cancer (CRC) and other gastrointestinal (GI) cancers include fluorouracil, leucovorin, oxaliplatin, and bevacizumab. While various combinations of these drugs are widely used as individualized regimens in the treatment of CRC patients, no well-established biomarkers or decision-support systems exist to help clinicians select from the many candidate regimens that are most effective for a given patient. This application describes an artificial intelligence system, the COLOXIS system or model, that assists in selecting effective drugs to create optimal regimens for treating CRC patients or other GI cancers, as the case may be. The machine learning system employs causal discovery methods to identify cancer driver genes involved in the disease development process of CRC (or possibly other GI cancers) and their target differentially expressed genes (DEGs). Co-expression patterns among these target DEGs are discovered to construct a set of "metagenes," whose expression levels reflect the state of cell signaling systems. Metagenes are used as tumor signatures to train a classification model (COLOXIS) to predict whether tumor cells are sensitive to, for example, oxaliplatin and bevacizumab. Validation studies using data from large-scale phase III clinical trials have demonstrated that the COLOXIS model can predict patient response not only to individual drugs (oxaliplatin and bevacizumab) but also to combination regimens such as FOLFOX, a chemotherapy regimen consisting of the drugs folinic acid (leucovorin, FOL), fluorouracil (5-FU, F), and oxaliplatin (ELOXATIN, OX). Accurate prediction of the efficacy of these drugs will facilitate decision-making by clinicians and CRC patients.

[0015] 1 is a flowchart illustrating a process 10 for training a machine learning model 12 for predicting the efficacy of chemical and biological agents for treating CRC, according to various embodiments of the present invention. Model development (e.g., the aforementioned "COLOXIS model") can consist of four major stages, which in various embodiments may be organized in the following order: (I) modeling disease mechanisms that influence heterogeneous drug responses; (II) discovering transcriptome patterns that reflect cancer cell disease mechanisms; (III) training a model to predict drug sensitivity based on cancer cell disease mechanisms; and finally, (IV) validating the predictive model. Process 10 may be performed in part or in whole using a computer system, such as computer system 100 described below in connection with FIG. 7.

[0016] Stage I involves modeling the heterogeneous disease mechanisms of CRC using causal discovery methods. For each drug used to treat CRC, fewer than 40% of patients respond. The heterogeneous response of cancer cells to drugs is due to differences in disease mechanisms. That is, individual tumors have different driver somatic genome alterations (SGAs) that disrupt signaling pathways and respond differently to drugs, resulting in different tumor disease mechanisms. Understanding the disease mechanisms of cancer cells within a tumor will enable us to predict drug responses in tumor cells.

[0017] To investigate common disease mechanisms in CRC, a tumor-specific causal inference (TCI) algorithm 14 can be applied to CRC genomic data 16 and CRC transcriptomic data 18. TCI is a Bayesian causal discovery algorithm invented by Drs. Xinghua Lu, Gregory Cooper, and others at the University of Pittsburgh, and is described in U.S. Patent Application Publication No. 16 / 349,192, published as U.S. Patent Application Publication No. 2019 / 0287651 A1, which is incorporated herein by reference in its entirety. In various embodiments, 290 CRC tumors characterized by The Cancer Genome Atlas (TCGA) are used as the CRC genomic data 16 and CRC transcriptomic data 18 to identify driver SGAs 20 and their target differentially expressed genes (DEGs) 22 in individual tumors. TCI searches for SGAs in specific tumors that are most likely to cause the molecular phenotype (e.g., DEG events) observed in the tumor. In this experiment, driver SGAs identified by the TCI algorithm were collected for each tumor. Among over 10,000 SGA-perturbed genes, 37 genes were identified as major drivers in the CRC cohort. Driver discovery significantly narrowed the number of candidate driver genes. Furthermore, TCI analysis identified 2,691 genes (i.e., target DEGs) regulated by these driver SGAs. The identification of 20 driver SGAs and 22 target DEGs of the driver SGAs allows us to infer the disease pathogenesis (the state of the signaling system) of cancer cells based on genomic and transcriptomic data obtained from tumors.

[0018] Stage II involves discovering transcriptome patterns that reflect disease pathogenesis in cancer cells. Gene expression states reflect the state of signaling pathways that control gene expression, which can be used to infer the state of cell signaling pathways. However, the expression levels of individual genes in different tumors vary greatly. Thus, single-gene expression is an unreliable marker for inferring signaling pathways. Because signaling pathways typically regulate a set of genes (gene modules) in cells, the expression state of gene modules is a better biomarker for inferring the state of signaling pathways. Identifying gene expression modules among the 22 DEGs discovered in Stage I allows us to infer the state of key signaling pathways disrupted in CRC, which can then be used to infer drug response.

[0019] In various embodiments, a database 24 of CRC transcriptome data is used to extract target DEGs in step 26. In various embodiments, a Gene Expression Omnibus (GEO) database with transcriptome data may be used as database 24. In our experiments, we collected 4,199 CRC tumors from the GEO database. In these 4,199 CRC tumors, the expression values ​​of the 2,691 target DEGs (see block 22) in block 28 were extracted in step 26. In various embodiments, extraction step 26 may include a series of consensus clustering analyses to identify a reduced set (e.g., 10 to 50, preferably approximately 15) of co-expression modules (metagenes) that showed clear co-expression patterns and provided strong signals for drug response. Consensus clustering is a method for aggregating (potentially contradictory) results obtained from multiple clustering algorithms. It is used when several different (input) clusterings are obtained for a particular dataset, and a single (consensus) clustering is desired, for example, to find target DEGs that fit better in some sense than the existing clusterings. An example of a suitable clustering technique is described in Monti S., Tamyo P, Mesirov J, Golub T., "Consensus Clustering: A Resampling-Based Method for Class Discovery and Visualization of Gene Expression Microarray Data," Machine Learning 2003;52:91-118, which is incorporated herein by reference. The Monti consensus clustering algorithm described in this paper is used to determine the number of clusters K. Given a dataset with a total number of points to be clustered, N, the algorithm resamples and clusters the data, and calculates an N x N consensus matrix for each K. Each element represents the proportion of times two samples are clustered together.A perfect stability matrix consists of all 0s and 1s, representing that all pairs of samples always cluster together or do not cluster together across all resampling iterations. The relative stability of the consensus matrices can be used to estimate the optimal K.

[0020] In step 30, gene set variation analysis (GSVA) ​​methods can be used to estimate meta-gene expression values ​​for each tumor. The expression values ​​(block 32) are then used as features representing the cellular state of cells within the tumor. Similar to competitive gene set testing, GSVA calculates a gene set enrichment score for each sample as a function of genes inside and outside the gene set. Furthermore, GSVA estimates the variation in gene set enrichment across samples, independent of any class label. Conceptually, this methodology can be understood as changing the coordinate system of gene expression data from genes to gene sets. This transformation facilitates the subsequent construction of pathway-centric models, such as identifying differential pathway activity or predicting survival. An example of a suitable GSVA method is described in Hanzelmann S, Castelo R, Guinney J., "GSVA: gene set variation analysis for microarray and RNA-seq data," BMC Bioinformatics 2013;14:7, which is incorporated herein by reference.

[0021] Stage III involves training a classification model to predict drug sensitivity. From a collection of datasets, such as the GEO dataset (GSE19860, GSE28702, GSE72970, GSE69657, and GSE104645), a cohort of mCRC patients is identified in step 34 as a discovery dataset in block 36. The patients in the cohort were treated, for example, with a FOLFOX regimen. Treatment response was measured as a Response Evaluation Criteria in Solid Tumors (RECIST) score. In various embodiments, a feature vector consisting of GSVA scores of a relatively small number of meta-genes (e.g., 10-50 meta-genes) may be constructed for each patient. In various embodiments, 15 meta-genes may be used. Based on the RECIST score, a class label may be assigned to each patient in the training dataset as a "responder" (CR and PR) or a "non-responder" (SD or PD). Based on this dimensionality-reduced training data 36, ​​a COLOXIS model 12 is trained in step 38. In various embodiments, the COLOXIS model 12 is a regularized logistic regression model. In step 38, a regularized logistic regression model (or other machine learning classifier, as appropriate) can be trained to predict response to FOLFOX in these mCRC patients. As described further below, the performance of the model 12 can be evaluated through cross-validation and external validation experiments. In other embodiments, other machine learning models may be trained in addition to the regularized logistic regression model. In other embodiments, the machine learning model 12 trained from the dataset 36 in step 38 can be, for example, a deep learning artificial neural network, a support vector machine and / or a decision tree, or an ensemble of machine learning models.

[0022] Stage IV involves validation of the predictive accuracy of COLOXIS model 12, at step 40. In various embodiments, validation was performed in the adjuvant setting of CRC. The clinical utility of COLOXIS model 12 may be validated using data obtained from TCGA and data obtained from two Phase III clinical trials.

[0023] We evaluated the COLOXIS model 12 to predict response to FOLFOX in patients with CRC. Genomic and clinical data were collected from 87 patients treated with the FOLFOX regimen and for whom overall survival outcomes were known from the TCGA study. Patient gene expression data were transformed using the GSVA algorithm to project patients into a 15-gene meta-space. The COLOXIS model 12 was then applied to predict whether patients would respond to FOLFOX (referred to as the COLOXIS+ and COLOXIS- groups). The survival of patients in the two groups was compared. The results, shown in Figure 2, indicate that patients assigned to the COLOXIS+ group had significantly better overall survival.

[0024] We also evaluated the prognostic value of the COLOXIS signature. The COLOXIS model 12 was applied to a cohort of 1,285 colon cancer patients who received adjuvant therapy after surgery. These patients were studied in two Phase III clinical trials (C-07 and C-08) conducted by the National Surgery Adjuvant Breast and Bowel Project (NSABP), a nonprofit clinical trial management organization. The C-07 trial compared the efficacy (in terms of preventing recurrence) of two chemotherapy combinations: (1) fluorouracil plus leucovorin (FULV) and (2) FULV plus oxaliplatin (FOLFOX). The C-08 trial compared the efficacy (in terms of preventing recurrence) of two chemotherapy combinations: (1) FOLFOX and (2) FOLFOX plus Bev.

[0025] We evaluated the prognostic value of the COLOXIS model by comparing recurrence-free survival (RFS) between the COLOXIS+ group and the COLOXIS- group among patients treated with FULV alone. As shown in Figure 3, the COLOXIS signature was statistically significantly associated with patient outcome (HR: 1.52, 95% CI: 1.07-2.15, P = 0.017).

[0026] We also evaluated the COLOXIS model for predicting the efficacy of oxaliplatin. Among 1,065 patients treated with fluorouracil plus leucovorin (FULV) (N=421) and FOLFOX (N=644), 526 were predicted to benefit from the oxaliplatin-containing regimen (referred to as COLOXIS+) and 539 were predicted not to benefit (COLOXIS-). The predictive value of the COLOXIS model was examined by comparing the response to oxaliplatin treatment in the COLOXIS+ group versus the COLOXIS- group. As shown in Figures 4A-C, COLOXIS+ patients benefited from oxaliplatin (HR = 0.65, 95% CI = 0.48-0.89, P = 0.0065, P for interaction = 0.03), whereas COLOXIS- patients did not (COLOXIS- HR = 1.08, 95% CI = 0.77-1.52, P = 0.65). Thus, the COLOXIS signature can predict the efficacy of oxaliplatin.

[0027] We also evaluated the COLOXIS model for predicting the efficacy of FOLFOX plus Bev. Of 644 patients treated with FOLFOX and 219 patients treated with FOLFOX plus Bev, 491 were assigned to the COLOXIS+ group and 372 to the COLOXIS- group. We examined the predictive value of the COLOXIS model by comparing the response to the addition of bevacizumab to treatment in the COLOXIS+ group versus the COLOXIS- group. As shown in Figure 5C, the COLOXIS+ group significantly benefited from the addition of bevacizumab to FOLFOX (HR = 0.58, 95% CI = 0.36-0.94, P = 0.025, interaction P = 0.101), whereas the COLOXIS- group did not (HR = 1.02, 95% CI = 0.64-1.63, P = 0.94). Thus, the COLOXIS signature can predict the efficacy of bevacizumab.

[0028] Once model 21 is trained and evaluated, it can be used as a diagnostic tool to predict how a patient will respond to a drug. For example, once trained as described above, a COLOXIS model can be trained to predict whether a CRC patient will respond to FOLFOX, oxaliplatin, or bevacizumab. FIG. 6 is a flowchart of a process for using the COLOXIS model as a decision support tool, according to various embodiments of the present invention. At step 60, a patient visits a healthcare provider and is diagnosed with GI cancer, thereby creating a need to make a determination as to whether the patient's tumor cells are sensitive to the oxaliplatin, bevacizumab, or FOLFOX regimen. At step 62, a tumor tissue sample is collected from the patient. The sample may be collected, for example, via biopsy or surgery. Next, at step 64, transcriptome profiling (RNA detection and quantification) of the sample collected at step 62 is performed. At step 64, any suitable technology / platform designed to characterize the transcriptome of a sample is used, such as a gene expression array or next-generation sequencing.

[0029] In step 66, expression quantification of genes involved in tumorigenesis identified in stage I of the process shown and described in connection with FIG. 1 is extracted (see step 14 of FIG. 1). Data obtained using various platforms can be transformed. Finally, in step 68, the tumor cell transcriptome can be mapped to a reduced-dimensional meta-gene space. For example, as described above, the meta-gene space can be a 15-dimensional meta-gene space.

[0030] Using the tumor meta-gene representation as input for the COLOXIS model, the model calculates probabilities or binary calls indicating whether the patient's tumor cells will respond to FOLFOX, oxaliplatin, and / or bevacizumab, at step 70. At step 72, one or more clinicians use the predictions from the COLOXIS model 12 to make treatment decisions for the patient.

[0031] FIG. 7 is a schematic diagram of a computer system 100 that can be used to implement the above-described embodiments. The illustrated computer system 100 includes one or more processors 102 and one or more memory units 104 in communication via a data bus and / or electronic data network. For simplicity, only one processor 102 and one memory unit are shown in FIG. 7. The memory 104 may store various software modules 106, 108, 110, and 112 comprising software or computer instructions executed by the processor 102. For example, the TCI module 106 may include software for performing the TCI analysis of step 14 of FIG. 1; the DEG extraction module 108 may include software for performing the target DEG extraction of step 26 of FIG. 1; the meta-gene learning and GSVA module 110 may include software for performing the meta-gene learning and GSVA of step 30 of FIG. 1; and the machine learning training module 112 may include software for training the model 12 of step 38 of FIG. 1.

[0032] The computer system 100 of Figure 7 may also be used to perform patient prediction according to Figure 6. The memory 104 may store software for TCI analysis in step 66, software for mapping the tumor cell transcriptome to a reduced dimensionality meta-gene space in step 68, and software for inputting the reduced dimensionality meta-gene space into a trained and validated COLOXIS model in step 70 to obtain a drug response prediction.

[0033] The one or more processors 102 may include one or more CPU cores, GPU cores, and / or AI accelerator cores. Memory 104 may comprise primary computer memory, such as read-only memory (ROM) and / or random access memory (e.g., RAM). Memory 104 may also comprise secondary memory, such as a magnetic or optical disk drive or flash memory. Software modules 106, 108, 110, and 112 may be implemented in computer software using any suitable computer programming language, such as .NET, C, C++, or Python, and using conventional, functional, or object-oriented techniques. The programming language for the computer software and other computer-implemented instructions may be translated into machine language by a compiler or assembler before execution and / or directly translated at runtime by an interpreter. Examples of assembly languages ​​include ARM, MIPS, and x86, examples of high-level languages ​​include Ada, BASIC, C, C++, C#, Python, R, COBOL, Fortran, Java, Lisp, Pascal, Object Pascal, Haskell, ML, and examples of scripting languages ​​include Bourne shell script, JavaScript, Python, Ruby, Lua, PHP, and Perl. Various data used in the processing of Figure 1 may be stored in primary, secondary, tertiary, and / or offline (e.g., cloud) storage.

[0034] Computer system 100 can be of various types, including a workstation, a server, a computing cluster, a blade server, a server farm, or any other data processing system or computing device. Computer system 100 could be, for example, a server on a cloud network. Because the nature of computers and networks is constantly changing, the description of computer system 100 depicted in FIG. 7 only represents a specific example for purposes of illustrating some implementations. Many other configurations of computer system 100 may have more or fewer components than the computer system depicted in FIG. 7.

[0035] Thus, in various embodiments, the present invention provides a novel causal method for discovering genes involved in CRC disease pathogenesis, which serve as informative features for detecting heterogeneous responses to drugs. The present invention also provides a novel feature-building method (metagene discovery) that lays the foundation for building robust predictive models. The COLOXIS model can predict the prognosis (outcome) of patients treated with FULV alone. The COLOXIS model can be used to predict who will have a better outcome if treated with FULV. The COLOXIS prediction can also show the prognostic course of outcome for patients treated with FOLFOX. The COLOXIS prediction can be used to predict patient outcome when treated with FOLFOX.

[0036] The COLOXIS model also predicts the effectiveness of oxaliplatin in treating patients with CRC, potentially enabling clinicians to decide whether or not to treat CRC patients with oxaliplatin. Accurate judgment not only improves the effectiveness of treatment but also prevents overtreatment with oxaliplatin, which only causes side effects.

[0037] The COLOXIS model also predicts the efficacy of FOLFOX+Bev in the adjuvant setting of CRC. The COLOXIS model allows clinicians to include bevacizumab in the adjuvant therapy for a subset of patients to enhance the efficacy of treatment.

[0038] Furthermore, the COLOXIS model can be used to predict sensitivity to the above regimen (FOLFOX) or single drug (oxaliplatin or bevacizumab) in gastrointestinal (GI) cancers, such as esophageal, gastric, and colorectal cancers, which commonly use regimens and drugs, allowing clinicians to create personalized regimens by selecting effective drugs to treat patients with GI cancer.

[0039] Thus, in one general aspect, the present invention provides a computer-implemented system and method for training a machine learning model to predict the efficacy of a biologic agent for gastrointestinal cancer, where the model output can be used by a clinician to determine treatment for a new gastrointestinal cancer patient. The method includes calculating, using a computer system 100 with one or more processors 102 executing instructions stored in a computer memory 104, expression values ​​for a set of meta-genes derived from target differentially expressed genes (DEGs) for gastrointestinal cancer, the set of meta-genes exhibiting co-expression patterns related to response to the biologic agent for a gastrointestinal cancer patient. The method also includes training, by the computer system, a model to predict the efficacy of the biologic agent for gastrointestinal cancer through machine learning, where the model output can be used by a clinician to determine treatment for a new gastrointestinal cancer patient.

[0040] In various embodiments, the method of the present invention further comprises identifying target DEGs for gastrointestinal cancer using a computer system. This identification may be performed using a tumor-specific causal inference algorithm applied to genomic and transcriptomic data. The method of the present invention may further comprise extracting a set of metagenes by using clustering analysis to identify a reduced set of metagenes that show distinct co-expression patterns and strong drug response signals. The reduced set of metagenes comprises 10 to 20 metagenes, preferably approximately 15 metagenes. Expression values ​​may be calculated using gene set variation analysis (GSVA).

[0041] In various implementations, the model comprises a classifier such as a logistic regression model.

[0042] In various implementations, the gastrointestinal cancer includes colon cancer. The biologic agent may include, for example, FOLFOX, oxaliplatin, and / or bevacizumab.

[0043] In various implementations, the GSVA scores for the set of meta-genes for a cohort of patients are feature vectors for training a model through machine learning, in which case each patient in the cohort may be labeled as a responder or non-responder to the biologic agent.

[0044] In various implementations, the methods of the present invention further include, after training the model, collecting tumor tissue samples from new patients diagnosed with gastrointestinal cancer, characterizing the transcriptomes of the samples, mapping the transcriptomes to the set of metagenes, and using the model to classify the efficacy of a biologic agent for the new patient based on mapping the transcriptome for the new patient to the set of metagenes.

[0045] The examples presented herein are intended to illustrate possible and specific implementations of the present invention. It will be understood that these examples are primarily intended to explain the present invention to those skilled in the art. Any particular aspect or aspects of the examples are not necessarily intended to limit the scope of the present invention. Furthermore, it will be appreciated that the drawings and descriptions of the present invention have been simplified to illustrate elements relevant to a clear understanding of the present invention, while excluding other elements for the sake of clarity. While various embodiments have been described herein, it will be apparent that those skilled in the art may devise various modifications, variations, and adaptations to these embodiments that achieve at least some of the advantages. Accordingly, the disclosed embodiments are intended to include all such modifications, variations, and adaptations without departing from the scope of the embodiments as set forth herein.

Claims

1. 1. A computer system comprising: one or more processor cores; computer memory in communication with the one or more processor cores; It is equipped with The computer memory includes: When executed by the one or more processor cores, the one or more processor cores calculating expression values ​​for a set of metagenes derived from target differentially expressed genes (DEGs) for gastrointestinal cancer, the set of metagenes exhibiting co-expression patterns related to response to a biological agent for patients with said gastrointestinal cancer; training a model to predict the efficacy of the biologic agent for the gastrointestinal cancer through machine learning; The output of the model can be used by medical practitioners to determine treatment for new gastrointestinal cancer patients. Computer system.

2. 2. The computer system of claim 1, wherein the computer memory stores instructions that, when executed by the one or more processor cores, cause the one or more processor cores to identify the target DEGs related to the gastrointestinal cancer.

3. 3. The computer system of claim 2, wherein the computer memory stores instructions that, when executed by the one or more processor cores, cause the one or more processor cores to identify the target DEGs related to the gastrointestinal cancer using a tumor-specific causal inference algorithm applied to genomic and transcriptomic data.

4. 3. The computer system of claim 2, wherein the computer memory stores instructions that, when executed by the one or more processor cores, cause the one or more processor cores to extract the set of metagenes by using clustering analysis to identify a reduced set of metagenes that exhibit distinct co-expression patterns and strong drug response signals.

5. The computer system of claim 4 , wherein the reduced set of metagenes comprises between 10 and 20 metagenes.

6. 5. The computer system of claim 4, wherein the computer memory stores instructions that, when executed by the one or more processor cores, cause the one or more processor cores to calculate the expression values ​​using Gene Set Variation Analysis (GSVA).

7. The computer system of claim 1 , wherein the model comprises a classifier.

8. The computer system of claim 7 , wherein the classifier comprises a logistic regression model.

9. The computer system of claim 1 , wherein the gastrointestinal cancer comprises colon cancer.

10. The computer system of claim 9 , wherein the biologic agent comprises FOLFOX.

11. 10. The computer system of claim 9, wherein the biologic agent comprises oxaliplatin.

12. The computer system of claim 9 , wherein the biologic agent comprises bevacizumab.

13. 7. The computer system of claim 6, wherein the GSVA scores of the set of meta-genes for a cohort of patients are feature vectors for training the model through the machine learning.

14. 14. The computer system of claim 13, wherein each patient in the cohort is labeled as a responder or non-responder with respect to the biologic agent.

15. calculating, by a computer system comprising one or more processors, expression values ​​for a set of meta-genes obtained from target differentially expressed genes (DEGs) for gastrointestinal cancer, the set of meta-genes exhibiting co-expression patterns related to response to a biologic agent for a patient with the gastrointestinal cancer; The computer system trains a model to predict the efficacy of the biologic agent for the gastrointestinal cancer through machine learning, and the output of the model can be used by a medical practitioner to determine treatment for new gastrointestinal cancer patients; A method comprising:

16. The method of claim 15, further comprising identifying the target DEGs related to the gastrointestinal cancer by the computer system.

17. 17. The method of claim 16, wherein identifying the target DEGs comprises using a tumor-specific causal inference algorithm applied to genomic and transcriptomic data.

18. 17. The method of claim 16, further comprising extracting, by the computer system, the set of metagenes by using clustering analysis to identify a reduced set of metagenes that exhibit distinct co-expression patterns and strong drug response signals.

19. 20. The method of claim 18, wherein the reduced set of metagenes comprises between 10 and 20 metagenes.

20. 20. The method of claim 18, wherein calculating the expression values ​​comprises using gene set variation analysis (GSVA).

21. The method of claim 15 , wherein the model comprises a classifier.

22. 22. The method of claim 21, wherein the classifier comprises a logistic regression model.

23. 23. The method of claim 22, wherein the gastrointestinal cancer comprises colon cancer.

24. 24. The method of claim 23, wherein the biologic agent comprises FOLFOX.

25. 24. The method of claim 23, wherein the biologic agent comprises oxaliplatin.

26. 24. The method of claim 23, wherein the biologic agent comprises bevacizumab.

27. 21. The method of claim 20, wherein the GSVA scores of the set of meta-genes for a cohort of patients are feature vectors for training the model through the machine learning.

28. 28. The method of claim 27, wherein each patient in the cohort is labeled as a responder or non-responder with respect to the biologic agent.

29. After training the model, collecting a tumor tissue sample from a new patient diagnosed with said gastrointestinal cancer; characterizing the transcriptome of said sample; mapping the transcriptome to the set of metagenes; using the model to classify the efficacy of the biologic agent for the new patient based on mapping the transcriptome for the new patient to the set of meta-genes; 16. The method of claim 15, further comprising: