Prognostic gene expression signature for prediction of treatment resistance of patients with pancreatic carcinoma
A 14-gene expression signature for pancreatic adenocarcinoma predicts treatment resistance and survival, addressing the limitations of existing methods by providing precise gene analysis and enabling personalized treatment strategies.
Patent Information
- Application Number
- PCT/US2025/017979
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-01
- Filing Date
- 2025-02-28
- Publication Date
- 2025-09-04
AI Technical Summary
Current methods for predicting treatment resistance in pancreatic adenocarcinoma (PAAD) lack precision and accuracy due to the absence of clear statistical significance thresholds for gene expression analysis, leading to incomplete interpretation of gene significance and limited personalized treatment options.
A 14-gene expression signature is developed using differential gene expression data from untreated tumor tissue, calculated through regularized Cox regression, to predict treatment resistance and survival, enabling personalized treatment strategies.
The 14-gene signature accurately identifies treatment-resistant patients, allowing for tailored therapeutic interventions, improving patient survival and treatment outcomes.
Smart Images

Figure US2025017979_04092025_PF_FP_ABST
Abstract
Description
PROGNOSTIC GENE EXPRESSION SIGNATURE FOR PREDICTION OF TREATMENT RESISTANCE OF PATIENTS WITH PANCREATIC CARCINOMA
[0001] This application claims priority to copending US provisional patent application with the serial number 63 / 560,543, which was filed March 1, 2024, and which is incorporated by reference herein.Statement of Government Interest
[0002] This invention was made with government support under the following grants: “Breast Cancer Center of Excellence”, grant numbers HU0001-16-2-0004 Sub #3406 (FYI 8); HU0001-19-20058 Sub #5171 (FY19); HU0001-21-20001 Sub #5686 (FY20); HU00012220005 Sub #5892 (FY 21); HU00012320022 Sub #6256 (FY 22) by the Uniformed Services University of the Health Sciences through the Henry M. Jackson Foundation for the Advancement of Military Medicine. The government may have certain rights in the invention.Field of the Invention
[0003] The field of the invention is diagnostic and prognostic system and methods, especially as it relates to gene expression-based prognostic tests to identify treatment resistance in patients with pancreatic carcinoma.Background of the Invention
[0004] The background description includes information that may be useful in understanding the present invention. It is not an admission that any of the information provided herein is prior art or relevant to the presently claimed invention, or that any publication specifically or implicitly referenced is prior art.
[0005] All publications and patent applications herein are incorporated by reference to the same extent as if each individual publication or patent application were specifically and individually indicated to be incorporated by reference. Where a definition or use of a term in an incorporated reference is inconsistent or contrary to the definition of that term provided herein, the definition of that term provided herein applies and the definition of that term in the reference does not apply.
[0006] Pancreatic adenocarcinoma (PAAD) is highly aggressive and difficult to treat. The 5- year survival rate is only about 12.5%. The Cancer Genome Atlas (TCGA) data has shown about 51% of PAAD patients didn’t achieve complete response after first course of treatment (CRAFCT) and had a very poor survival compared to patients achieved complete response. This is due in part to the fact that the predictability of response and treatment resistance has been elusive.
[0007] Various methods have been developed to attempt to address this difficulty in prognosis. For example, US20240263243A1 attempts to improve the selection of effective first-line systemic treatments for patients with pancreatic ductal adenocarcinoma (PDAC). This reference discusses analyzing the RNA transcript profile of PDAC tumor samples, and performing independent component analysis (ICA) to identify key gene expression correlated with drug sensitivity. However, this process of ICA and correlation analysis merely provides a holistic approach wherein certain genes are classified as having positive or negative correlations with drug-sensitivity without providing a method of accounting for granularity and particularization with regard to each specific gene and the degree to which they are relevant. Further, the absence of clear, pre-specified statistical significance thresholds used to determine the inclusion of genes in the sensitivity signatures makes it more difficult to fully interpret the biological significance of the genes comprising the signatures, as it is unclear which ones are truly the most important predictors of sensitivity. Further, without clear significance criteria, the patent's method may be overly stringent in selecting only the most strongly correlated genes for the sensitivity signatures, while excluding genes with more moderate, but still potentially meaningful, associations with drug response.
[0008] In another example of an attempt to optimize the prognosis of PAAD, US20220170109A1 describes methods for determining the subtype of a pancreatic tumor as either basal -like or classical based on some information about the expression of 16 specified genes in tumor cells, and then using this subtype information to "treat the patient with gemcitabine, optionally in combination with nab-paclitaxel, if the assigned subtype is a basal- like subtype and treating the patient with FOLFIRINOX if the assigned subtype is classical". This patent fails to analyze each gene’s expression level in response to various treatments. Additionally, as before, this patent does not assign any statistical significance or weighting values to the individual genes used in the subtyping classifier. The subtyping is based solely on the relative expression of gene pairs, with a simple binary classification of basal-like orclassical based on the overall calculated "Raw Score." Thus, the methods disclosed in this patent are unable to account for more subtle relationships between gene expression patterns and treatment response with an optimal level of accuracy. Additionally, the lack of statistical thresholds for the individual genes means the method cannot readily identify genes that may be more or less relevant drivers of the subtype distinctions and treatment sensitivities. As a result, the patent's treatment recommendations are limited to the two broad subtype-based regimens, without the ability to consider a more graduated spectrum of gene expression profiles that could inform a wider range of personalized treatment options for patients. This binary classification therefore lacks the nuanced, quantitative assessment of the underlying gene expression data which may be necessary be improved success and survival of patients.
[0009] Another reference, US20120264639A1, identifies a 6-gene expression signature to stratify PDAC patients into high-risk or low-risk groups with different median overall survival times. This patent states that it uses Distance Weighted Discrimination (DWD) and Significance Analysis of Microarrays (SAM) to identify the 6 differentially expressed genes between non-metastatic and metastatic PDAC tumor samples. However, at merely 6 genes identified as being relevant, the methods of this patent fail to account for other genes which may be moderately relevant or slightly relevant, yet still impact the prognosis. Further, the data used to identify the 6-gene signature consisted of data from 15 non-metastatic primary PDAC tumors and 15 metastatic primary PDAC tumors for a total of 30 samples. This volume of data is relatively small for the type of genomic analysis necessary for optimal accuracy in identifying relevant genes. With a 30-sample dataset, there is limited statistical power to detect all truly differentially expressed genes, especially those with smaller effect sizes, increased risk of overfitting the gene signature to the specific training samples, and difficulty in validating the robustness of the signature across all patient cohorts and variants of tumors.
[0010] Therefore, although there are methods in the art which attempt to enhance prognosis of pancreatic cancer, there remains a need for tests that predict treatment resistance in patients with PAAD with higher precision and accuracy.Brief Description of the Drawings
[0011] Figure 1 depicts a schematic flow-chart for derivation and validation of a 55-gene signature.
[0012] Figure 2 depicts data about a PAAD signature of 225 genes where the cutoff was 0.9 and the accuracy was 0.56. The figure includes K-M graphs, boxplots and Wilcoxon tests, and a plot.
[0013] Figure 3 depicts data about a PAAD signature of 55 genes, wherein FC=2, the cutoff was 0.3, and the accuracy was 0.657. The figure includes K-M graphs, boxplots and Wilcoxon tests, and a plot.
[0014] Figure 4 depicts logistic regression data and K-M plots for various datasets using the 55 -gene signature.
[0015] Figure 5 depicts data from the RandomForest analysis conducted to validate the information obtained about the 55 genes, as well as other relevant information. The RandomForest accuracy was measured to be 0.646. A mean decrease gini (MDG) table is also depicted.
[0016] Figure 6 depicts results from the Leave-one-out-cross-validation (LOOCV) analysis conducted to further validate the information obtained about the 55 genes. The LOOCV accuracy was measured to be 0.647. An MDG table is further shown.
[0017] Figure 7 depicts results from using a public dataset and signature score and logistic regression methods to validate the 55-gene signature.
[0018] Figure 8 depicts a schematic flow-chart of a derivation and validation process for a 14- gene signature.
[0019] Figure 9 depicts a table of clinical characteristics observed when gathering data from the training data (patients) and the testing data for the 14-gene signature analysis (also for the 55-gene signature analysis).
[0020] Figure 10 depicts a Sankey plot of treatment information for available data.
[0021] Figure 11 depicts a table with the drugs used versus the treatment in training data.
[0022] Figure 12 depicts a volcano plot wherein, which identifies genes that exhibit resistance versus genes that exhibit sensitivity, and the accuracy of such results are additionally measured (-log (adjusted P value)).
[0023] Figure 13 provides various results about the derivation and validation for the 14- treatment survival gene signature, including K-M plots, boxplots and Wilcoxon tests, etc.
[0024] Figure 14 depicts an expression heatmap and coefficient for each of the 14 identified genes, and risk score plot.
[0025] Figure 15 depicts logistic regression results for the 14 identified genes.
[0026] Figure 16 depicts results from a Random Forest analysis for the 14 identified genes. Further, a table is depicted which shows the mean decrease gini (MDG) for 14 genes, plus AJCC stage, gender, and race, listed in order from highest MDG to lowest MDG. Further depicted is an ROC curve for Random Forest having an area under the curve of 0.64. The Random Forest accuracy was measured to be 0.64.
[0027] Figure 17 depicts predicted and actual results from a Random Forest analysis measuring the correlation between overall survival and the training dataset and testing dataset for the 14 identified genes wherein the cutoff of probability is 0.45. Further, a table is depicted which shows the mean decrease gini (MDG) for 14 genes, listed in order from highest MDG to lowest MDG. Further depicted is an ROC curve for Random Forest having an area under the curve of 0.66. The Random Forest accuracy was measured to be 0.66.
[0028] Figure 18 depicts predicted and actual results from a Leave-one-out-cross-validation (LOOCV) analysis measuring the predicted and actual correlation between overall survival and the training and testing data sets for the 14 identified genes. The LOOCV accuracy was measured to be 0.647. An MDG table is further shown.
[0029] Figure 19 depicts results from using a public dataset and signature score and logistic regression methods to validate the signature of the 14 identified genes, compared to the validation of the 55-gene signature.
[0030] Figure 20 depicts further results from using another public dataset and signature score and logistic regression methods to validate the signature of the 14 identified genes.
[0031] Figure 21 depicts tables with the 304 identified treatment associated genes, and the 55 and 14 identified survival and treatment associated genes.Summary of The Invention
[0032] The inventive subject matter is directed to various diagnostic / prognostic systems and methods that allow for stratification of patients into likely treatment responders versus treatment resistant patients. In preferred embodiments, the stratification is based on a gene expression signature that is associated with treatment response and survival. Preferably, such signature is derived from clinical data (the outcome of the first course treatment) in which complete response was characterized as Complete Remission / Response (Yes) vs Progressive Disease (No). Contemplated systems and methods will allow screening PAAD patients with high-risk of treatment resistance directly from an analysis of gene expression of the untreated tumor and may further provide novel therapeutic strategies to customize treatment for those patients to improve patient survival.
[0033] Therefore, in one aspect of the inventive subject matter, the inventors contemplate a method of identifying treatment resistance in a patient diagnosed with PAAD (pancreatic adenocarcinoma) that includes a step of obtaining gene expression data from a plurality of genes of a patient tumor, wherein the genes are previously identified as genes that are differentially expressed in response to a treatment.
[0034] In some embodiments, the previously identified genes are further selected in a survival analysis. For example, contemplated gene expression data can be obtained from one or more (e.g., at least 10, or at least 30, or all of the) genes selected from the group consisting of GRIA2, KCNJ3, CTSV, FAM83A, BSN, CYP2C8, CDHR1, RFX6, SERPINA10, HMGA2, ANKS1B, KCNB1, CACNB2, BCAM, FGF17, AQP4, LYPD2, SLC29A4, TAT, DRAIC, KCNH6, FAM83A-AS1, AC015819.1, AMER3, LINGO4, CACNA1B, AGBL4, ANKRD36BP2, CALB1, CARMIL3, LINC02384, MYH16, KRT16, LINC01146, SPTB, MYT1, C8orf31, AMPH, KCNMB2, UCN3, BRINP1, ATP6V0E2-AS1, LINC00683, FGL1, CCDC188, SNAP91, PAEP, MSMB, EFR3B, GNAZ, AL031595.3, RAB39B, KCNC1, CELF3, and DGKB. A signature score is then calculated fromstattX genet, wherein the stat is from DESeq2 analysis results, and / or the threshold score is 0.3, the patient is treated with one or more chemotherapeutic drugs when the signature score is below a threshold score.
[0035] In another example, contemplated gene expression data can be obtained from one or more (e.g., at least 5, or at least 10, or all of the) genes selected from the group consisting of MYH16, FGL1, CTSV, CYP2C8, CXCL9, FAM83A, CDHR1, C8orf31, IGHM, AL031777.1,ANKRD36BP2, IGHJ3, LINC02384, and MSMB. In such embodiments, the signature score can be calculated from i4coefficientt X gene=0.14252487 x MYH16 + (-0.05441616) x FGL1 + 0.07047860 x CTSV + (-0.04119509) x CYP2C8 + 0.67476731 x CXCL9 + 0.19773183 x FAM83A + (-0.34206776) x CDHR1 + 0.14538827 x C8orfi l + (-0.03638636) x IGHM + (- 0.26946263) x AL031777.1 + (-0.19391202) x ANKRD36BP2 + (-0.16645176) x IGHJ3 + (-0.19189528) x LINC02384 + 0.28444088 x MSMB, and / or the threshold score is -0.3. As will be readily appreciated, the treatment may comprise administration of at least one of gemcitabine, capecitabine, oxaliplatin, irinotecan, fluorouracil, and leucovorin.
[0036] In additional embodiments, the inventors contemplate a method of treating a patient diagnosed with PAAD where the method involves the steps of obtaining gene expression data from a plurality of genes of a patient tumor, wherein the genes are selected from the group consisting of MYH16, FGL1, CTSV, CYP2C8, CXCL9, FAM83A, CDHR1, C8orf31, IGHM, AL031777.1, ANKRD36BP2, IGHJ3, LINC02384, and MSMB, and further calculating a signature score from the obtained expression data; and treating the patient with an anti-metabolite when the signature score is below a threshold score, or treating the patient with alternative regimens such as combination with radiation when the signature score is above a threshold score.
[0037] As will be readily appreciated, in another embodiment, the threshold score is -0.3.
[0038] In a different embodiment, the inventors contemplate that the signature score is calculated fromcoefficient^ X genet=0.14252487 x MYH16 + (-0.05441616) x FGL1 + 0.07047860 x CTSV + (-0.04119509) x CYP2C8 + 0.67476731 x CXCL9 + 0.19773183 x FAM83A + (-0.34206776) x CDHR1 + 0.14538827 x C8orf31 + (- 0.03638636) x IGHM + (- 0.26946263) x AL031777.1 + (-0.19391202) x ANKRD36BP2 + (-0.16645176) x IGHJ3 + (-0.19189528) x LINC02384 + 0.28444088 x MSMB.
[0039] The inventors also contemplate that the patient may have undergone prior treatment for PAAD.
[0040] In another aspect of the inventive subject matter, the tumor is a primary tumor
[0041] In yet another aspect, the tumor is a metastatic tumor
[0042] The inventors contemplate that the expression data, in some instances, was obtained from a biopsy sample of the tumor.
[0043] However, in other embodiments, the expression data was obtained from a blood sample containing a nucleic acid of the tumor.
[0044] The anti-metabolite may be, in some aspects of the disclosed subject matter, gemcitabine.
[0045] In a further embodiment, the inventors contemplate a biological classification system for selection of treatment options in PAAD comprising a database storing gene expression data from a plurality of genes of a patient tumor, wherein the genes are selected from the group consisting of MYH16, FGL1, CTSV, CYP2C8, CXCL9, FAM83A, CDHR1, C8orf31, IGHM, AL031777.1, ANKRD36BP2, IGHJ3, LINC02384, and MSMB, and further comprising at least one processor coupled to the database and configured to execute an algorithm on the gene expression data to derive digital descriptors, calculate a signature score using the digital descriptors, and cause the patient to be treated with an anti-metabolite when the signature score is below a threshold score, or cause the patient to be treated with alternative regimens such as combination with radiation when the signature score is above a threshold score.
[0046] Regarding this system, the inventors contemplate that in some embodiments, the expression data was obtained from a biopsy sample of the tumor.
[0047] In a further embodiment, the expression data was obtained from a blood sample containing a nucleic acid of the tumor.
[0048] It is additionally contemplated that the plurality of genes may be selected based on the differential expression of the genes in response to exposure to the anti-metabolite.
[0049] It is yet further contemplated that the processor can be further configured to recommend a specific anti-metabolite based on the signature score.
[0050] In some embodiments, the anti-metabolite is gemcitabine.
[0051] In other embodiments, the threshold score is -0.3.
[0052] The inventors also contemplate that the signature score may be calculated from l4coefficient^ X genet= 0.14252487 x MYH16 + (-0.05441616) x FGL1 + 0.07047860 x CTSV + (-0.04119509) x CYP2C8 + 0.67476731 x CXCL9 + 0.19773183 x FAM83 A + (- 0.34206776) x CDHR1 + 0.14538827 x C8orf31 + (-0.03638636) x IGHM + (- 0.26946263) x AL031777.1 + (-0.19391202) x ANKRD36BP2 + (-0.16645176) x IGHJ3 + (-0.19189528) x LINC02384 + 0.28444088 x MSMB.
[0053] Various objects, features, aspects, and advantages of the inventive subject matter will become more apparent from the following detailed description of preferred embodiments.Detailed Description
[0054] The inventors have discovered systems and methods for identification of treatment resistant PAAD / PDAC patients where identification uses gene expression data from untreated tumor tissue, and where the so obtained gene expression data are used to calculate a signature score that is indicative of treatment resistance. Accordingly, where a signature score is below a threshold score, patients can be treated using conventional chemotherapeutical treatment options whereas patients can be treated with alternate treatments (e.g., non-conventional chemotherapeutical treatment options) where the signature score is above a threshold value.
[0055] In one typical embodiment, the inventors used publicly available TCGA data to develop a survival-treatment associated gene signature to screen patients with high-risk of treatment resistance, who may need to be administered novel therapeutic strategies to customize patient treatment. Differential gene expression in a tumor tissue expression was measured across >10,000 genes to so arrive at treatment relevant list of genes (here: about 304 genes) that are significantly differentially expressed between tumors from patients who achieved CRAFCT and patients who did not achieve CRAFCT (complete response after first course of treatment)^ These genes were then subjected to further machine learning and / or different statistical methods to so arrive at models with significantly reduced complexity. As is described in more detail below, the inventors derived two exemplary models, with one model using 55 gene expression data and with another model using gene expression data of 14 genes. Thus, it should be appreciated that multi-gene signatures (e.g., 55 or 14 genes) can be derived from previously known data and can be directly used to predict treatment resistance and patient survival, leading to improved patient outcomes and potential guidance to alternate therapy options. In one exemplary method, RNAseq data of 177 PAAD patients and their clinical data weredownloaded from TCGA data portal and processed. The dataset was randomly split into training and testing data using stratification method. RNAseq data analysis and survival analyses were performed on the training data to identify the survival -treatment associated genes, and a survival -treatment gene signature was derived and independently tested on the test data using signature score by cutoff, logistic regression, and random forest. The signature was then validated by an independent dataset of 96 cases. Among the 177 patients, 137 patients had the information of CRAFCT, and 122 patients had specific treatment information. Among those 122 patients, 99% received chemotherapy, 36% received radiation therapy. About half of the PAAD patients (51%) did not achieve CRAFCT.
[0056] In the above method, RNAseq data analysis revealed 304 significantly expressed genes associated with the status of complete response Yes vs No, and survival analyses further selected 55 genes associated with both treatment and survival as a signature. This signature, using z-scaled scores and a cut-off of 0.3, was applied to the 99 patients of training data. In this model, the accuracy was 0.66, and patients predicted to be Yes had a significantly better (p=0.005, HR=0.42) overall survival than those predicted No. To further test the signature, the inventors applied the signature to the test data of 78 patients. A random forest model showed patients screened to be Yes had a significantly better (p=0.016, HR=0.43) overall survival than those screened No. The signature (here: using 40 of the 55 genes) was applied to an independent dataset of 96 cases to significantly screen the patients to be Yes vs No (resistance to treatment) with p=0.04 and HR=0.57, showing the 55-gene signature was valid and robust. In this embodiment, the direction of the score is opposite to the final direction as in the 14-gene example, so here a score higher than the threshold of 0.3 means the patient will benefit from the conventional therapy.
[0057] More specifically, and as discussed in more detail below, the inventors used machine learning method to split the dataset to be training (50 complete response and 49 progression of the disease cases) and testing (78 cases) datasets, and used DESeq2 analysis to get 304 genes associated with treatment response. Further survival analysis was performed on the genes to obtain signature genes associated with both treatment and survival. These were then further validated by Random Forest (RF) and Leave-one-out-cross-validation (LOOCV) methods, and the signature was additionally validated by signature score and logistic regression methods using public dataset. In an alternate approach, “regularized Cox regression” was applied to the 304 treatment associated genes to systematically select signature genes which were associatedwith both treatment and survival, resulting in 14 genes suitable for a signature score. This 14- gene signature was further evaluated and validated using signature score method, logistic regression, random forest, and leave one out cross validation method, and the signature was also validated using an external independent dataset by signature score and logistic regression methods. Notably, the results were even better than the 55-gene signature. Accordingly, a 14- gene signature may also be used where a simplified test with better predictability is desired.
[0058] In particular, and as shown in more detail in the Experimental Data and Figures, an exemplary signature test can therefore include 14 genes in a gene panel. As will be readily appreciated, the 14 genes can be assessed by sequencing or other methods such as PCR from an untreated tumor sample to get RNA expression values. These values can then be normalized (up-quartile) values, log2(X+l) transformed, z-scaled, and then weighted by a coefficient value for each corresponding gene from regularized Cox regression on training data. Summation of all weighted data will so generate a signature score. For example, a score following this method can be calculated by: Signature score = ^coefficienti xgenei = 0.14252487 x MYH16 + (- 0.05441616) x FGL1 + 0.07047860 x CTSV + (-0.04119509) x CYP2C8 + 0.67476731 x CXCL9 + 0.19773183 x FAM83A + (-0.34206776) x CDHR1 + 0.14538827 x C8orf31 + (- 0.03638636) x IGHM + (- 0.26946263) x AL031777.1 + (-0.19391202) x ANKRD36BP2 + (- 0.16645176) x IGHJ3 + (-0.19189528) x LINC02384 + 0.28444088 x MSMB, which can be further z- scaled (using the same mean (- 2.798x l0'17) and SD (0.97406) for z-scaling the signature of training data). If the signature score is > - 0.3, the patient is classified high-risk of treatment resistance or treatment complete response of No group. Or the z-scaled score(s) can be input into a logistic model built from the training data, and when the predicted probability of the patient >0.44, the patient is classified high-risk of treatment resistance or treatment complete response of No group. Alternatively, a trained random forest (RF) or RF-LOOCV model can be used to screen the patient(s) to be yes (low-risk) or No (high-risk). For those high-risk of treatment resistant patients, a new or more complex treatment strategy may be appropriate. For the patients who are classified low-risk of treatment resistance (e.g., signature score is <= -0.3 or predicted probability of the patient <= 0.44), one can avoid over treatment.
[0059] Accordingly, the inventors contemplate methods and systems of identifying treatment resistance in a patient diagnosed with PAAD using data from the 14 gene panel discussed above. For example, and in one embodiment, the inventors contemplate a biological identification system, comprising a database storing gene expression data from a plurality ofgenes of a patient tumor, wherein the genes are selected from the group consisting of MYH16, FGL1, CTSV, CYP2C8, CXCL9, FAM83A, CDHR1, C8orf31, IGHM, AL031777.1, ANKRD36BP2, IGHJ3, LINC02384, and MSMB. The system typically further includes at least one processor coupled to the database and configured to: execute an algorithm on the gene expression data to derive digital descriptors; calculate a signature score using the digital descriptors; and cause the patient to be treated with an anti -metabolite when the signature score is below a threshold score; or cause the patient to be treated with alternative regimens such as combination with radiation when the signature score is above a threshold score. In another embodiment, the processor may be further configured to recommend a specific anti -metabolite, such as gemcitabine or any other anti-metabolite, based on the signature score calculated by the processor or otherwise.
[0060] It is further noted that gene expression data in any of the subject matter disclosed herein may be obtained from a biopsy sample of a tumor, a blood sample containing a nucleic acid of the tumor, obtaining fragments of the tumor and implanting the fragments into immunodeficient mice and generating PDX models from the samples, producing primary cell cultures and collecting data from the cultures, generating a pancreatic organoid culture directly from a biopsy and collecting data from the culture, and / or utilizing publicly available gene expression data from commercially available cell lines or other sources to validate any of the foregoing options.
[0061] In further embodiments, out of the 14 genes, the inventors contemplate that less than 14 genes may be selected for analysis, and specific genes may be selected based on the differential expression of the genes in response to exposure to various treatments, such as exposure to an anti-metabolite. For example, only genes having the strongest correlation between expression and treatment may be selected as desired.
[0062] The inventors contemplate that a significant advantage of the disclosed subject matter over the prior art is unique method of selection and validation of genes, which in turn results in a more graded and accurate signature score. That is, the method of selecting and validating genes described throughout this disclosure is not reproduced elsewhere, and this method results in a highly accurate signature score which is more likely than methods in the prior art to result in successful treatment. Further, the disclosed methods and systems are more likely to enable successful particularized treatment. For example, the threshold score may be adjusted based on specific patient tumors, or the coefficients may be freely changed based on specific patientgroup or assay methods. As such, it should be recognized that the coefficients and thresholds are shown as examples, which can be method and / or cohort-dependent (but can be adjusted into a refined model). In addition, this disclosure can measure the likelihood of genes having treatment resistance and the rather than merely measuring which genes are sensitive to treatments, which allows targeted therapies designed to overcome resistance mechanisms, enables more personalized treatment, and improved prognosis over alternative methods.
[0063] It is noted that while the methods taught herein are applicable to PAAD / PDAC, it contemplated that similar results would be available for other types of pancreatic tumors, such as those seen in patients having acinar cell carcinoma, adenosquamous carcinoma, cystadenocarcinoma, and giant cell carcinoma, all considered subtypes of pancreatic adenocarcinoma depending on the specific cell type involved.Experimental Data
[0064] The discussion below represents exemplary data gathered through nonlimiting embodiments of the disclosed subject matter. As such, the data and examples discussed below and in the figures should be understood as illustrative rather than restrictive, and they are not intended to limit the scope of the disclosed subject matter. Variations, modifications, and equivalents of these examples may be apparent to those skilled in the art and are considered within the scope of the invention as defined by the claims.
[0065] When conducting the experiment, the inventors intended to address various problem with treating PAAD and similar cancers, which are some of the deadliest solid malignancies. The recurrence rates remain extremely high after first course of treatment. Based on TCGA data, patients with complete response show 34% recurrence, where patients with non-complete response show 89% recurrence. Thus, the aim of the experiments was to develop a gene expression prognostic signature to screen patients before treatment, and further to identify the treatment complete responders for the most effective therapeutic strategies to improve personalized oncology care, and to identify the patients with treatment resistance to guide oncologists to apply new therapeutic strategies.
[0066] As mentioned above and throughout this disclosure, the inventors used DESeq2, regularized Cox regression, signature score method, logistic regression, random Forest (RF), and RF-Leave one out cross validation (RF-LOOCV). The training and testing data comprisedTCGA (expression data from tissues before treatment). The two validation datasets comprised pancreatic adenocarcinoma were obtained from public sources.
[0067] Figure 1 depicts a schematic flow-chart for derivation and validation of a 55-gene signature. Figure 2 depicts data about a PAAD signature of 225 genes where the cutoff was 0.9 and the accuracy was found to be 0.56. In particular, Figure 2 depicts predicted results and actual results from a DESeq analysis for the correlation between overall survival and a training dataset and a testing dataset. Further depicted is a boxplot and Wilcoxon test of the signature score comparison between treatment response Yes vs No for training data. Additionally, a signature score plot for predicted results in PAAD is depicted.
[0068] Once derivation and validation were completed, a 55-gene signature was developed. First, RNAseq data analysis (DESeq2) identified 304 differentially expressed genes (FC of 2 and Padj <0.05) associated with the status of treatment complete response Yes vs No from training data. Then, survival analysis of the expression high vs low further selected 55 genes associated with both treatment and survival as a signature. Signature score was calculated by Signature score =stati X gene=3.4567 x GRIA2 + 3.2408 x KCNJ3 + (-4.5676) x CTSV + (-3.5793) x FAM83A + 4.2062 x BSN + 3.2438 x CYP2C8 + 3.7051 x CDHR1 + 3.3121 x RFX6 + 3.6747 x SERPINA10 + (-3.3455) x HMGA2 + 4.0279 x ANKS1B + 5.0776 x KCNB1 + 5.0221 x CACNB2 + 4.6127 x BCAM + 3.4896 x FGF17 + 3.2464 x AQP4 + (- 3.5479 x LYPD2) + 4.2457 x SLC29A4 + 4.935 x TAT + 4.4373 x DRAIC + 3.5514 x KCNH6 + (-3.4317) x FAM83A-AS1 + 4.2064 x AC015819.1 + 4.1213 x AMER3 + 3.9207 x LINGO4 + 4.0741 x CACNA1B + 3.7985 x AGBL4 + 4.1143 x ANKRD36BP2 + 4.6366 x CALB1 + 4.7514 x CARMIL3 + 3.7838 x LINC02384 + (-4.1154) x MYH16 + (-4.521) x KRT16 + 3.8584 x LINC01146 + 4.0174 x SPTB + 3.7231 x MYT1 + (-3.1571) x C8orf31 + 4.2327 x AMPH + 3.5019 x KCNMB2 + 3.4371 x UCN3 + 4.8941 x BRINP1 + 4.98 x ATP6V0E2- AS1 + 3.695 x LINC00683 + 3.9148 x FGL1 + 4.6313 x CCDC188 + 3.9941 x SNAP91 + (- 3.3991) x PAEP + 3.8539 x MSMB + 4.6541 x EFR3B + 4.4243 x GNAZ + 4.2415 x AL031595.3 + 4.033 x RAB39B + 3.7296 x KCNC1 + 4.8262 x CELF3 + 3.4288 x DGKB, and wherein stat is from DESeq 2 analysis results. Input data: gene values, Log2(up-quartile normalized expression values +1).
[0069] The signature, using z-scaled scores and cut off of 0.3, was applied to the training data for evaluation, to the test data (using the same mean (720.391) and SD (268.256) for z-scaling the signature of training data), and external data for validation. The signature was furtherevaluated by the random forest (RF) model and leave one out cross validation RF (LOOCV- RF) model on training data and validated from the test data.
[0070] Figure 3 depicts the data about the PAAD signature of these 55 identified genes, wherein FC=2, the cutoff was 0.3, and the accuracy was 0.657. In particular, Figure 3 depicts predicted results for the correlation between survival and the training and testing datasets. Further, the actual results for the correlation between training dataset and survival is depicted. Additionally, a boxplot and Wilcoxon test of the signature score comparison between treatment response Yes vs No for training data is depicted. Finally, a signature score plot for predicted results in PAAD is depicted.
[0071] Figure 4 depicts logistic regression data and K-M plots for various datasets related to these 55 genes. In particular, Figure 4 depicts an ROC curve used to evaluate a logistic regression wherein the area under the curve is 0.697551. Further, the predicted survival graphs are provided with ‘Yes’ and ‘No’ lines for both the training dataset and testing dataset. In addition, the actual survival results are depicted for the testing dataset.
[0072] The inventors then performed various analyses to further validate the previously obtained information about the differentially expressed genes. Figure 5 depicts data from the RandomForest analysis conducted to validate the information obtained about the 55 genes, as well as other relevant information. For example, Figure 5 depicts predicted and actual results from a Random Forest analysis measuring the correlation between overall survival and the training dataset and testing dataset. Further, a table is depicted which shows the mean decrease gini (MDG) for 55 genes, listed in order from highest MDG to lowest MDG. Further depicted is an ROC curve for Random Forest having an area under the curve of 0.69. The Random Forest accuracy was measured to be 0.646.
[0073] Further, the inventors validated the data using the LOOCV analysis, as depicted in Figure 6. In particular, Figure 6 depicts predicated and actual results from the LOOCV analysis measuring the predicted and actual correlation between overall survival and the training and testing data sets. The LOOCV accuracy was measured to be 0.647. An MDG table is further shown.
[0074] Once the data was validated by RandomForest and LOOCV, the inventors further validated the data by comparing it to a public dataset, as depicted in Figure 7. In particular, Figure 7 depicts results from this public dataset and signature score and logistic regressionmethods used to validate the signature. The public dataset included 96 cases having survival data. This data was downloaded from cbioportal: Pancreatic Adenocarcinoma (QCMG, Nature 2016).
[0075] At the conclusion of the analysis of the 55 gene signature, the results were as follows: Signature score method: cutoff= 0.3, Training: acc.=0.66, testing (all testing cases) p-value= 0.049; testing (Yes cases) p-value=0.033. Logistic regression training accuracy = 0.62, testing (all testing cases) p-value= 0.038; AUC=0.70. Random forest method: Training: acc.=0.65, testing (all testing cases) p-value= 0.018; testing (NA + partial / stable cases) p-value= 0.023; testing (NA cases) p-value=0.048; AUC=0.69. LOOCV Random Forest method: Training: acc.=0.65, testing (all testing cases) p-value= 0.015; testing (NA + partial / stable cases) p- value= 0.045; testing (NA cases) p-value=0.156.
[0076] The three methods relatively agree with one another. The signature worked especially well for testing data. Further, the public data validated the signature significantly.
[0077] Based on the above analysis, 14 genes were identified as being especially relevant. The inventors therefore repeated similar analyses as the above with a focus on the 14 gene signature. Figure 8 depicts a schematic flow-chart of a derivation and validation process for a 14-gene signature. Figure 9 depicts a table of clinical characteristics observed when gathering data from the training data (patients) and the testing data.
[0078] Figure 10 depicts a Sankey plot of treatment for available data. As depicted, 78.7% of the patient population underwent the Whipple procedure, and 86% of the population were administered gemcitabine (by itself or in combination with other drugs). Figure 11 depicts a table with the drugs used versus the treatment response in training data.
[0079] Figure 12 depicts a volcano plot wherein, out of 304 genes identified as differentially expressed genes associated with treatment (DEGs), 57 of the genes are correlated with resistance and 247 genes are correlated with sensitivity to treatment. In other words, 247 genes are associated with an increase in ‘Yes’ response and 57 genes are ‘associated with a decrease in ‘Yes’ response.
[0080] A regularized Cox regression model was then used to select the signature genes. The results were as follows:
[0081] DESeq 2 analysis on complete response Yes vs No to 304 genes (FC of 2 and Padj<0.05)
[0082] Input data: gene values, Log2(up-quartile normalized expression values +1), then z- scaled the values across samples, and survival data
[0083] Regularized Cox regression model using 10-fold cross-validation to get minimized lambda to select significant genes (here 14 genes)
[0084] The 14 genes and their corresponding coefficients are displayed in the signature score formula below:
[0085] Signature score = i4coefficient^ X gene^ = 0.14252487 x MYH16 + (-0.05441616) x FGL1 + 0.07047860 x CTSV + (-0.04119509) x CYP2C8 + 0.67476731 x CXCL9 + 0.19773183 x FAM83A + (-0.34206776) x CDHR1 + 0.14538827 x C8orfi l + (-0.03638636) x IGHM + (- 0.26946263) x AL031777.1 + (-0.19391202) x ANKRD36BP2 + (-0.16645176) x IGHJ3 + (-0.19189528) x LINC02384 + 0.28444088 x MSMB
[0086] Training data: Z-Scaled the signature scores to perform the following analysis: boxplot comparison, score analysis, logistic regression. Matrix value (scaled across patients) x coefficients from regularized Cox model, and input into RF, RF-LOOCV.
[0087] Testing data: Z-Scaled the signature scores (using the same mean and SD for z-scaling the signature of training data) to perform the following analysis: score method, logistic regression. Matrix value (scaled across patients) x coefficients from regularized Cox model, and input into RF, RF-LOOCV.
[0088] Independent data (already z-scored data): Z-Scaled the signature scores to perform the following analysis: score method, logistic regression. Matrix value x coefficient from regularized Cox model, and input into RF, RF-LOOCV.
[0089] Figure 13 provides various results about the derivation and validation for the 14- treatment survival gene signature. For example, a boxplot and Wilcoxon test is displayed of the score comparison between treatment response Yes v. No for training data. Further, in the top right, a graph is provided with the plot of Yes (dot) and No (cross) of signature scores with cutoff line a -0.3 for training data, wherein the majority of dots below the cutoff line are ‘Yes’ and the majority of crosses above the cutoff line are ’No’. In the bottom left, a K-M plot of treatment response for training using score cutoff at -0.3 is provided. In the bottom middle, theprediction of the treatment response for the test data using the same cutoff is provided. In the bottom right, a graph of predicted survival of the 20 Yes testing data points is depicted.
[0090] Figure 14 depicts an expression heatmap and coefficient for each of the 14 identified genes and risk score plot. As depicted, each of the 14 genes has a unique Z-score associated with its level of differential expression. Upon confirming the 14 gene signature, the inventors validated the data. For example, Figure 15 depicts logistic regression data for the 14 identified genes.
[0091] Further, Figure 16 depicts predicted and actual results from a Random Forest analysis measuring the correlation between overall survival and the training dataset and testing dataset for the 14 identified genes. Further, a table is depicted which shows the mean decrease gini (MDG) for 14 genes, listed in order from highest MDG to lowest MDG. Three important clinical features such as AJCC stage, gender, and race were used and contributed less in the model, showing the signature can work effectively and independently to screen the patients. Further depicted is an ROC curve for Random Forest having an area under the curve of 0.64. The Random Forest accuracy was measured to be 0.64. Figure 17 depicts predicted and actual results from another Random Forest analysis measuring the correlation between overall survival and the training dataset and testing dataset for the 14 identified genes wherein the cutoff of the probability is 0.45 and without the clinical features. Further, a table is depicted which shows the mean decrease gini (MDG) for 14 genes, listed in order from highest MDG to lowest MDG. Further depicted is an ROC curve for Random Forest having an area under the curve of 0.66. The Random Forest accuracy was measured to be 0.67.
[0092] Figure 18 depicts predicted and actual results from a Leave-one-out-cross-validation (LOOCV) analysis measuring the predicted and actual correlation between overall survival and the training and testing data sets for the 14 identified genes. The LOOCV accuracy was measured to be 0.647. An MDG table is further shown. Figure 19 depicts results from using a public dataset and signature score and logistic regression methods to validate the signature of the 14 identified genes.
[0093] Figure 20 depicts further results from using another public dataset and signature score and logistic regression methods to validate the signature of the 14 identified genes.
[0094] Figure 21 depicts tables with the 304 identified treatment associated genes, and the 55 and 14 identified survival and treatment associated genes.
[0095] In some embodiments, the numbers expressing quantities of ingredients, properties such as concentration, reaction conditions, and so forth, used to describe and claim certain embodiments of the invention are to be understood as being modified in some instances by the term “about.” As used herein, the terms "about" and "approximately", when referring to a specified, measurable value (such as a parameter, an amount, a temporal duration, and the like), is meant to encompass the specified value and variations of and from the specified value, such as variations of + / -20% or less, alternatively variations of + / -10% or less, alternatively + / -5% or less, alternatively + / -1% or less, alternatively + / -0.1% or less of and from the specified value, insofar as such variations are appropriate to perform in the disclosed embodiments. Thus, the value to which the modifier "about" or "approximately" refers is itself also specifically disclosed. The recitation of ranges of values herein is merely intended to serve as a shorthand method of referring individually to each separate value falling within the range. Unless otherwise indicated herein, each individual value is incorporated into the specification as if it were individually recited herein.
[0096] All methods described herein can be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by context. The use of any and all examples, or exemplary language (e.g., “such as”) provided with respect to certain embodiments herein is intended merely to better illuminate the invention and does not pose a limitation on the scope of the invention otherwise claimed. No language in the specification should be construed as indicating any non-claimed element essential to the practice of the invention.
[0097] As used in the description herein and throughout the claims that follow, the meaning of “a,” “an,” and “the” includes plural reference unless the context clearly dictates otherwise. Also, as used in the description herein, the meaning of “in” includes “in” and “on” unless the context clearly dictates otherwise. As also used herein, and unless the context dictates otherwise, the term "coupled to" is intended to include both direct coupling (in which two elements that are coupled to each other contact each other) and indirect coupling (in which at least one additional element is located between the two elements). Therefore, the terms "coupled to" and "coupled with" are used synonymously.
[0098] It should be apparent to those skilled in the art that many more modifications besides those already described are possible without departing from the inventive concepts herein. The inventive subject matter, therefore, is not to be restricted except in the scope of the appended claims. Moreover, in interpreting both the specification and the claims, all terms should beinterpreted in the broadest possible manner consistent with the context. In particular, the terms “comprises” and “comprising” should be interpreted as referring to elements, components, or steps in a non-exclusive manner, indicating that the referenced elements, components, or steps may be present, or utilized, or combined with other elements, components, or steps that are not expressly referenced. Where the specification or claims refer to at least one of something selected from the group consisting of A, B, C . . . . and N, the text should be interpreted as requiring only one element from the group, not A plus N, or B plus N, etc.
Claims
CLAIMSWhat is claimed is:
1. A method of identifying treatment resistance in a patient diagnosed with PAAD (pancreatic adenocarcinoma), comprising: obtaining gene expression data from a plurality of genes of a patient tumor, wherein the genes are previously identified as genes that are differentially expressed in response to a treatment; calculating a signature score from the obtained expression data; wherein the gene expression data are obtained from genes selected from the group consisting of GRIA2, KCNJ3, CTSV, FAM83A, BSN, CYP2C8, CDHR1, RFX6, SERPINA10, HMGA2, ANKS1B, KCNB1, CACNB2, BCAM, FGF17, AQP4, LYPD2, SLC29A4, TAT, DRAIC, KCNH6, FAM83A-AS1, AC015819.1, AMER3, LING04, CACNA1B, AGBL4, ANKRD36BP2, CALB1, CARMIL3, LINC02384, MYH16, KRT16, LINC01146, SPTB, MYT1, C8orf31, AMPH, KCNMB2, UCN3, BRINP1, ATP6V0E2-AS1, LINC00683, FGL1, CCDC188, SNAP91, PAEP, MSMB, EFR3B, GNAZ, AL031595.3, RAB39B, KCNC1, CELF3, and DGKB; and treating the patient with an antimetabolite when the signature score is above a threshold value.
2. The method of claim 1, wherein the previously identified genes are further selected in a survival analysis.
3. The method of claim 1, wherein the gene expression data are obtained from at least 30 genes of the group of genes.
4. The method of claim 1, wherein the gene expression data are obtained from all genes of the group of genes.
5. The method of claim 1, wherein the signature score is calculated fromstattX genet = 3.4567 x GRIA2 + 3.2408 x KCNJ3 + (-4.5676) x CTSV + (-3.5793) x FAM83A + 4.2062 x BSN + 3.2438 x CYP2C8 + 3.7051 x CDHR1 + 3.3121 x RFX6 + 3.6747 x SERPINA10 + (-3.3455) x HMGA2 + 4.0279 x ANKS1B + 5.0776 x KCNB1 + 5.0221 x CACNB2 + 4.6127 x BCAM + 3.4896 x FGF17 + 3.2464 x AQP4 + (-3.5479 x LYPD2) + 4.2457 x SLC29A4 + 4.935 x TAT + 4.4373 x DRAIC + 3.5514 x KCNH6 + (-3.4317)x FAM83A-AS1 + 4.2064 x AC015819.1 + 4.1213 x AMER3 + 3.9207 x LINGO4 + 4.0741 x CACNA1B + 3.7985 x AGBL4 + 4.1143 x ANKRD36BP2 + 4.6366 x CALB1 + 4.7514 x CARMIL3 + 3.7838 x LINC02384 + (-4.1154) x MYH16 + (-4.521) x KRT16 + 3.8584 x LINC01146 + 4.0174 x SPTB + 3.7231 x MYT1 + (-3.1571) x C8orf31 + 4.2327 x AMPH + 3.5019 x KCNMB2 + 3.4371 x UCN3 + 4.8941 x BRINP1 + 4.98 x ATP6V0E2-AS1 + 3.695 x LINC00683 + 3.9148 x FGL1 + 4.6313 x CCDC188 + 3.9941 x SNAP91 + (-3.3991) x PAEP + 3.8539 x MSMB + 4.6541 x EFR3B + 4.4243 x GNAZ + 4.2415 x AL031595.3 + 4.033 x RAB39B + 3.7296 x KCNC1 + 4.8262 x CELF3 + 3.4288 x DGKB, and wherein stat is from DESeq 2 analysis results.
6. The method of claim 1, comprising a step of treating the patient with one or more chemotherapeutic drug other than the antimetabolite, and optionally radiation therapy when the signature score is below the threshold value.
7. The method of claim 6, wherein the chemotherapeutic drug is selected from the group consisting of fluorouracil, oxaliplatin, capecitabine, leucovorin, fluorouracil, irinotecan, and oxaliplatin.
8. The method of claim 6, wherein the one or more chemotherapeutic drug comprises a combination of fluorouracil and oxaliplatin.
9. The method of claim 6, wherein the one or more chemotherapeutic drug comprises a combination of leucovorin, fluorouracil, irinotecan, and oxaliplatin.
10. The method of claim 6, wherein step of treating comprises radiation therapy.
11. The method of claim 1, wherein the threshold score is 0.3.
12. A method of treating a patient diagnosed with PAAD (pancreatic adenocarcinoma), comprising: obtaining gene expression data from a plurality of genes of a patient tumor, wherein the genes are selected from the group consisting of MYH16, FGL1, CTSV, CYP2C8, CXCL9, FAM83A, CDHR1, C8orf31, IGHM, AL031777.1, ANKRD36BP2, IGHJ3, LINC02384, and MSMB; calculating a signature score from the obtained expression data; andtreating the patient with an antimetabolite when the signature score is below a threshold score.
13. The method of claim 12, wherein the threshold score is -0.3.
14. The method of claim 12, wherein the signature score is calculated from l4coefficient^ X genet= 0.14252487 x MYH16 + (-0.05441616) x FGL1 + 0.07047860 x CTSV + (-0.04119509) x CYP2C8 + 0.67476731 x CXCL9 + 0.19773183 x FAM83A + (-0.34206776) x CDHR1 + 0.14538827 x C8orf31 + (-0.03638636) x IGHM + (- 0.26946263) x AL031777.1 + (-0.19391202) x ANKRD36BP2 + (-0.16645176) x IGHJ3 + (-0.19189528) x LINC02384 + 0.28444088 x MSMB.
15. The method of claim 12, wherein the antimetabolite is gemcitabine.
16. The method of claim 12, comprising a step of treating the patient with one or more chemotherapeutic drug other than the antimetabolite, and optionally radiation therapy when the signature score is above the threshold value.
17. The method of claim 16, wherein the chemotherapeutic drug is selected from the group consisting of fluorouracil, oxaliplatin, capecitabine, leucovorin, fluorouracil, irinotecan, and oxaliplatin.
18. The method of claim 12, wherein patient has undergone prior treatment for PAAD.
19. The method of claim 12, wherein the tumor is a primary tumor.
20. The method of claim 12, wherein the tumor is a metastatic tumor.
21. The method of claim 12, wherein the expression data was obtained from a biopsy sample of the tumor.
22. The method of claim 12, wherein the expression data was obtained from a blood sample containing a nucleic acid of the tumor.
23. A biological classification system for selection of treatment options in PAAD (pancreatic adenocarcinoma), comprising: a database storing gene expression data from a plurality of genes of a patient tumor, wherein the genes are selected from the group consisting of MYH16, FGL1,CTSV, CYP2C8, CXCL9, FAM83A, CDHR1, C8orf31, IGHM, AL031777.1, ANKRD36BP2, IGHJ3, LINC02384, and MSMB; at least one processor coupled to the database and configured to: execute an algorithm on the gene expression data to derive digital descriptors; calculate a signature score using the digital descriptors; and cause the patient to be treated with an antimetabolite when the signature score is below a threshold score.
24. The system of claim 23, wherein the expression data was obtained from a biopsy sample of the tumor.
25. The system of claim 23, wherein the expression data was obtained from a blood sample containing a nucleic acid of the tumor.
26. The system of claim 23, wherein the plurality of genes are selected based on the differential expression of the genes in response to exposure to the anti-metabolite.
27. The system of claim 23, wherein the processor is further configured to recommend a specific anti-metabolite based on the signature score.
28. The system of claim 23, wherein the anti-metabolite is gemcitabine.
29. The system of claim 23, wherein the threshold score is -0.3.
30. The system of claim 23, wherein the signature score is calculated from i4coefficient^ X genet = 0.14252487 x MYH16 + (-0.05441616) x FGL1 + 0.07047860 x CTSV + (- 0.04119509) x CYP2C8 + 0.67476731 x CXCL9 + 0.19773183 x FAM83A + (- 0.34206776) x CDHR1 + 0.14538827 x C8orf31 + (-0.03638636) x IGHM + (- 0.26946263) x AL031777.1 + (-0.19391202) x ANKRD36BP2 + (-0.16645176) x IGHJ3 + (-0.19189528) x LINC02384 + 0.28444088 x MSMB.
Citation Information
Patent Citations
Gene expression classifier for predicting pancreatic cancer metastasis risk and in-vitro diagnostic kit
CN112626218A