Methods, devices, and models for predicting prognosis in patients with multiple myeloma (MM)
The detection of 97 genes in MM patients using a Bayesian classifier classifies them into MCL1-M subtypes, predicting prognosis and treatment response, addressing the limitations of current classification schemes and enhancing personalized treatment for MM.
Patent Information
- Application Number
- JP2020560444
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-04-28
- Filing Date
- 2019-04-25
- Publication Date
- 2025-11-27
- Estimated Expiration
- 2039-04-25
AI Technical Summary
Current molecular classification schemes for multiple myeloma (MM) do not predict treatment response and fail to correlate with the pathogenesis of the disease, hindering personalized treatment development.
A method involving the detection of expression levels of 97 genes in MM patients using a Bayesian classifier to classify patients into MCL1-M high or low subtypes, which predicts prognosis and response to bortezomib treatment.
Enables accurate prognosis prediction and tailored treatment decisions for MM patients, improving survival outcomes and reducing unnecessary treatment costs and side effects.
Smart Images

Figure 0007776848000005 
Figure 0007776848000006 
Figure 0007776848000007
Abstract
Description
[Technical Field]
[0001] The present invention is in the field of biotechnology, and in particular relates to the molecular classification and application of multiple myeloma. [Background technology]
[0002] Multiple myeloma (MM) is the second most common hematologic malignancy caused by the abnormal proliferation of plasma cells. In China, the incidence of MM is estimated to be 1–2 cases per 100,000 people. The majority of MM occurs in the elderly population aged 60 years or older. Within China's aging population, the incidence of MM will increase over time, posing a serious health risk to the elderly population. MM is typically manifested by the excessive proliferation of abnormal plasma cells and the secretion of abnormal immunoglobulin proteins or immunoglobulin protein fragments called M proteins. M protein concentration is an important diagnostic indicator of MM.
[0003] The development of the proteasome inhibitor bortezomib and immunomodulatory agents such as lenadoline and thalidomide has significantly improved the survival of patients with MM. However, MM remains incurable. MM exhibits wide heterogeneity in biological and clinical characteristics. As a result, response to multidrug combination therapy and survival improvement vary substantially among MM patients. The underlying mechanisms remain unclear, hindering the development of personalized treatments. To further our understanding of MM biology and facilitate treatment decisions, it is important to develop a simple and reliable molecular classification method for MM. Several molecular classification schemes have been proposed. For example, Bergsagel et al. proposed a classification scheme using eight MM subtypes based on differential cyclin D expression and chromosomal translocations. Based on unbiased transcriptome analysis, Zhan et al. and Broyl et al. proposed seven to ten molecular subtypes of MM. These subtypes were further simplified into high-risk and low-risk groups based on patient survival. Additionally, prognosis-related gene expression profiles, such as UAMS-70, UAMS-17, UAMS-80, IFM-15, Millennium-100, EMC-92, gene amplification index GPI-5, MRC-IX-6, and centrosome duplication index, have also been proposed.
[0004] However, the molecular classification schemes and gene expression profiles described above do not predict treatment response, and no correlation between molecular classification and plasma cell development has been made.Furthermore, no attempt has been made to correlate genetic classification with the pathogenesis of MM. Summary of the Invention
[0005] In order to better clarify the cytological origin of multiple myeloma and provide targeted treatment for multiple myeloma, the present invention provides the following technical solutions: The object of the present invention is to provide a use for obtaining or detecting the expression of 97 genes in multiple myeloma patients to be tested.
[0006] The present invention provides the application of materials for obtaining or detecting the expression of 97 genes in a tested multiple myeloma patient in preparing a product for predicting the prognosis of the tested multiple myeloma patient.
[0007] Survival outcomes include survival rate, survival duration, and degree of survival risk.
[0008] Survival rates include overall survival and progression-free survival.
[0009] The present invention further provides the following functions a to c of a substance for obtaining or detecting the expression of 97 genes in a multiple myeloma patient to be tested: a) To detect the effect of bortezomib or bortezomib-containing treatment in patients with MM b) To detect sensitivity to bortezomib or bortezomib-containing treatment in MM patients; c) providing an application for preparing a product having at least one of the following: bortezomib or a bortezomib-containing treatment for an MM patient.
[0010] Another object of the present invention is to provide a material for obtaining or detecting the expression of 97 genes in a test multiple myeloma (MM) patient or the use of a device for implementing a Bayesian classifier for multiple myeloma.
[0011] The present invention provides the application of materials for obtaining or detecting the expression of 97 genes in a test multiple myeloma (MM) patient, or a device for implementing a Bayesian classifier for multiple myeloma, in the preparation of a product for predicting the prognosis of a test multiple myeloma patient.
[0012] Prognosis is reflected in the prognosis of survival, survival time, or degree of survival risk.
[0013] The present invention further provides a material for acquiring or detecting the expression of 97 genes in a test multiple myeloma (MM) patient, or an apparatus for implementing a Bayesian classifier for multiple myeloma, having the following functions a to c: a.Detecting the efficacy of bortezomib or bortezomib-containing treatment in MM patients; b. Detection of sensitivity to bortezomib or bortezomib-containing therapy in MM patients; c. Instructions for administration of bortezomib or a treatment comprising bortezomib in MM patients; The 97 genes are ACBD3, ADAR, ADSS, ALDH2, ANP32E, ANXA2, ATF3, ATP8B2, CACYBP, CAPN2, CCND1, CCT3, CDC42SE1, CERS2, CHSY3, CLIC1, CLMN, COPA, CSNK1G3, DAP3, DENND1B, ENSA, EPRS, EPST I1, EVL, FAM13A, FAM49A, FLAD1, FRZB, GLRX2, HAX1, HDGF, HLA-A, HLA-B, HLA-C, HLA-F, HLA- G, IL6R, ISG20L2, JTB, KLF2, LAMTOR2, LDHA, MCL1, MOXD1, MRPL24, MRPL9, MVP, MYL6, NDUFS2, including NOP58, NOTCH2NL, NTAN1, PAK1, PI4KB, PIEZO1, PIK3AP1, PIM2, PIP5K1B, PMVK, POGZ, PPIA, PRCC, PRKCA, PRRC2C, PSMB4, PSMD4, RAB29, RCBTB2, SCAMP3, SCAPER, SDHC, SEL1L3, SELPLG, SHC1, SIDT1, SSR2, STAP1, TAP1, TIMM17A, TLR10, TMCO1, TOR1AIP2, TOR3A, TP53INP1, TPM3, TRANK1, TROVE2, UAP1, UBE2Q1, UBQLN4, UHMK1, VPS45, YY1AP1, ZC3H11A, ZFP36, and ZNF593; Bayesian classifier for multiple myeloma 1) obtaining expression data for 97 classifier genes in n MM samples; The 97 gene expression data of MM samples is obtained from an existing database or is 97 gene expression data of multiple myeloma samples constructed using more than 100 samples; n is 100 or greater, The expression levels of the 97 genes are the expression levels of the 97 genes in multiple myeloma cells; 2) assigning MM samples to MCL1-M high or MCL1-M low subtypes by consensus clustering; 3) constructing a Bayesian classifier using the naive Bayes method based on the two subtypes of step 2), the 97 gene expression data of the n multiple myeloma samples of step 1), and the prognosis survival data of the n multiple myeloma samples, wherein the Bayesian classifier is constructed using the naive Bayes method.
[0014] In step 3 above, first, n multiple myeloma samples were randomly divided into a training set and a validation set according to a sample number ratio of greater than 1:1. Then, the expression data of 97 genes were used in training, and a consensus clustering algorithm was used to obtain high and low MCL1-M subtype tags for each sample. Next, a Bayesian classifier for multiple myeloma was constructed using the naive Bayes algorithm of klaR, an R machine learning package, to predict the high and low MCL1-M subtypes of a single patient.
[0015] The above-mentioned method of obtaining the expression data of the 97 genes of each multiple myeloma sample is to detect the expression of the 97 genes of the multiple myeloma sample or to obtain the expression of the 97 genes of the multiple myeloma sample from a database.
[0016] A third object of the present invention is to provide a product.
[0017] The articles of manufacture provided by the present invention include a device (the device can be a CD or a computer, etc.) that obtains or detects the expression of 97 genes in multiple myeloma patients and runs a multiple myeloma Bayesian classifier.
[0018] With respect to the above product, the product has at least one of the following features: The product has the following functions 1) to 4), namely: 1) predict the prognosis of the multiple myeloma patients tested; 2) To detect sensitivity to bortezomib or bortezomib-containing drugs in patients with multiple myeloma. 3) To detect the efficacy of bortezomib or bortezomib-containing drugs in the tested multiple myeloma patients; 4) Instructing the test subject multiple myeloma patient to receive bortezomib or a bortezomib-containing drug.
[0019] The article of manufacture further includes a carrier for recording the detection method.
[0020] The detection method includes the steps of obtaining or detecting the expression of 97 genes in a test multiple myeloma patient to obtain expression data of the 97 genes in the test multiple myeloma patient; and then classifying the expression data of the 97 genes in the test multiple myeloma patient using a Bayesian classifier for multiple myeloma, wherein the prognosis of multiple myeloma patients belonging to the high MCL1-M subtype is significantly poorer than the prognosis of multiple myeloma patients belonging to the low MCL1-M subtype; Alternatively, the detection method includes the steps of obtaining or detecting the expression of 97 genes in a test multiple myeloma patient to obtain expression data of the 97 genes in the test multiple myeloma patient; and then classifying the expression data of the 97 genes in the test multiple myeloma patient using a Bayesian classifier for multiple myeloma, wherein the efficacy of bortezomib or a bortezomib-containing drug in multiple myeloma patients belonging to the high MCL1-M subtype is superior to the efficacy in multiple myeloma patients belonging to the low MCL1-M subtype; Alternatively, the detection method includes the steps of obtaining or detecting the expression of 97 genes in a test multiple myeloma patient to obtain expression data of the 97 genes in the test multiple myeloma patient; and then classifying the expression data of the 97 genes in the test multiple myeloma patient using a Bayesian classifier for multiple myeloma, wherein if the test multiple myeloma patient belongs to a high MCL1-M subtype, bortezomib or a bortezomib-containing drug is used for treatment, and if the test multiple myeloma patient belongs to a low MCL1-M subtype, bortezomib or a bortezomib-containing drug is not used for treatment.
[0021] In the above-mentioned products, the multiple myeloma patient being tested may be a single patient or multiple patients.
[0022] In the above product, n multiple myeloma samples are 551 samples or Alternatively, the ratio greater than 1:1 mentioned above is to randomly split the training set and validation set according to a 2:1 ratio.
[0023] A fourth object of the present invention is to provide a method for constructing a model for classifying multiple myeloma patients.
[0024] The methods provided by the present invention include: 1) obtaining expression data for 97 classifier genes in n MM samples; The 97 gene expression data of MM samples is obtained from an existing database or is 97 gene expression data of multiple myeloma samples constructed using more than 100 samples; n is 100 or greater, The expression levels of the 97 genes are the expression levels of the 97 genes in multiple myeloma cells; 2) assigning MM samples to MCL1-M high or MCL1-M low subtypes by consensus clustering; 3) constructing a Bayesian classifier using naive Bayes method based on the two subtypes of step 2), the 97 gene expression data of the n multiple myeloma samples of step 1), and the prognosis survival data of the n multiple myeloma samples, to obtain a desired model.
[0025] The expression of 97 genes in multiple myeloma patients was derived from the expression of 97 genes in tumor cells of multiple myeloma patients.
[0026] The application of the above-mentioned method for obtaining or detecting the expression of 97 genes in multiple myeloma patients to be tested and / or a device for performing a multiple myeloma Bayesian classifier or a model obtained by the above-mentioned method to predict the prognosis survival rate of multiple myeloma patients to be tested is also within the scope of protection of the present invention.
[0027] The expression of 97 genes in multiple myeloma patients was derived from the expression of 97 genes in tumor cells of multiple myeloma patients.
[0028] The application of the above-mentioned device for obtaining or detecting substances expressed by the 97 genes in a tested multiple myeloma patient and / or for operating a Bayesian classifier for multiple myeloma, or the model obtained by the above-mentioned method in the preparation of a product for predicting the prognostic survival rate of a tested multiple myeloma patient, are all within the scope of protection of the present invention.
[0029] The application of the above-mentioned device for obtaining or detecting substances expressed by the 97 genes in a tested multiple myeloma patient and / or for operating a Bayesian classifier for multiple myeloma, or the model obtained by the above-mentioned method in the preparation of a product for predicting the prognosis survival of a tested multiple myeloma patient, are all within the scope of protection of the present invention.
[0030] The application of the above-mentioned device for obtaining or detecting substances expressed by the 97 genes in a tested multiple myeloma patient and / or for operating a Bayesian classifier for multiple myeloma, or the model obtained by the above-mentioned method in the preparation of a product for predicting the degree of survival risk of a tested multiple myeloma patient, are all within the scope of protection of the present invention.
[0031] The present invention also provides a method for classifying multiple myeloma patients, comprising: The method includes the steps of obtaining or detecting the expression of 97 genes in a test multiple myeloma patient to obtain expression data of the 97 genes in the test multiple myeloma patient, and then classifying the expression data of the 97 genes of the test multiple myeloma patient using a Bayesian multiple myeloma classifier to determine whether the test multiple myeloma patient belongs to the MCL1-M high subtype or the MCL1-M low subtype.
[0032] The present invention further provides a method for predicting the prognosis of a test multiple myeloma patient, comprising the steps of obtaining or detecting expression of 97 genes in the test multiple myeloma patient to obtain expression data of the 97 genes in the test multiple myeloma patient; and then classifying the expression data of the 97 genes in the test multiple myeloma patient using a Bayesian classifier for multiple myeloma, wherein the prognosis prediction for multiple myeloma patients belonging to the high MCL1-M subtype is significantly poorer or worse than the prognosis prediction for multiple myeloma patients belonging to the low MCL1-M subtype.
[0033] Prognosis is reflected in the prognosis of survival, survival time, or degree of survival risk.
[0034] The prognosis of multiple myeloma patients belonging to the high MCL1-M subtype is significantly worse than that of multiple myeloma patients belonging to the low MCL1-M subtype, and this is reselected as at least one of the following 1) to 3): 1) the predicted prognostic survival rate of the tested multiple myeloma patients belonging to the high MCL1-M subtype is significantly lower than the predicted prognostic survival rate of the tested multiple myeloma patients belonging to the low MCL1-M subtype; 2) the predicted prognostic survival of the tested multiple myeloma patients belonging to the high MCL1-M subtype is significantly lower than the predicted prognostic survival of the tested multiple myeloma patients belonging to the low MCL1-M subtype; 3) The predicted survival risk of the tested multiple myeloma patients belonging to the high MCL1-M subtype is significantly lower than the predicted survival risk of the tested multiple myeloma patients belonging to the low MCL1-M subtype.
[0035] The present invention provides a method for detecting the effectiveness of bortezomib or a bortezomib-containing drug in a test multiple myeloma patient, which includes the steps of obtaining or detecting the expression of 97 genes in the test multiple myeloma patient to obtain expression data of the 97 genes in the test multiple myeloma patient, and then classifying the expression data of the 97 genes in the test multiple myeloma patient using a Bayesian classifier for multiple myeloma, wherein the prognosis prediction for multiple myeloma patients belonging to the high MCL1-M subtype is better than the prognosis prediction for multiple myeloma patients belonging to the low MCL1-M subtype.
[0036] The present invention provides guidance for administering bortezomib or a bortezomib-containing drug to a test multiple myeloma patient, which includes the steps of obtaining or detecting the expression of 97 genes in the test multiple myeloma patient to obtain expression data of the 97 genes in the test multiple myeloma patient, and then classifying the expression data of the 97 genes in the test multiple myeloma patient using a Bayesian classifier for multiple myeloma, wherein if the test multiple myeloma patient belongs to a high MCL1-M subtype, bortezomib or a bortezomib-containing drug is used for treatment, and if the test multiple myeloma patient belongs to a low MCL1-M subtype, bortezomib or a bortezomib-containing drug is not used for treatment.
[0037] The expression of the 97 classifier genes can be obtained from the MM database or detected directly from MM samples.
[0038] The expression levels of the above-mentioned genes are gene expression levels in multiple myeloma tumor cells. [Brief explanation of the drawings]
[0039] [Figure 1] A plot of the ROC curve for Bayesian classification on the GSE2658 dataset. [Figure 2]This is a plot of the ROC curve for Bayesian classification on the MMRF dataset. [Figure 3] A plot of the ROC curve for Bayesian classification on the GSE19784 dataset. [Figure 4] Overall survival rates of MM patients with high or low MCL1-M in GSE2658. [Figure 5] Overall survival rates of MM patients with high or low MCL1-M in GSE2658. [Figure 6a] 1 shows the overall survival rate of MM patients with high or low MCL1-M in GSE19784. [Figure 6b] 1 shows progression-free survival rates for MM patients with high or low MCL1-M in GSE19784. [Figure 7a] Figures 7a-7d show the different responses of MM patients with high or low MCL1-M to bortezomib-containing treatment in GSE19784. [Figure 7b] See above. [Figure 7c] See above. [Figure 7d] See above. DETAILED DESCRIPTION OF THE INVENTION
[0040] All experimental methods used in the following examples are conventional methods unless otherwise indicated.
[0041] All materials, reagents, etc. used in the following examples are commercially available unless otherwise indicated.
[0042] Example 1. Screening for molecular diagnostic markers and molecular typing of multiple myeloma From the MM gene expression dataset GSE2658 published by NCBI, a gene module containing 87 genes co-expressed with MCL1 (MCL1-M) was identified using Pearson correlation coefficient analysis. Based on the above, 46 genes up-regulated in MM samples with low MCL1-M expression were also identified. To obtain stable classification results, 36 genes with low classification ability were excluded from the 133 genes mentioned above, and 97 genes with relatively high levels of differential expression were selected.
[0043] These 97 genes are: ACBD3, ADAR, ADSS, ALDH2, ANP32E, ANXA2, ATF3, ATP8B2, CACYBP, CAPN2, CCND1, CCT3, CDC42SE1, CERS2, CHSY3, CLIC1, CLMN, COPA, CSNK1G3, DAP3, DENND1B, ENSA, EPRS, EPSTI1, EVL, FAM13A, FAM49A, FLAD1, FRZB, GLRX2, HAX1, HDGF, HLA-A, HLA-B, HLA-C, HLA-F, HLA-G, IL6R, ISG20L2, JTB, KLF2, LAMTOR2, LDHA, MCL1, MOXD1, MRPL24, MRPL9, MVP, MYL6, NDUFS2, NOP58, NOTCH2NL, NTAN1, PAK1, PI4KB, PIEZO1, PIK3AP1, PIM2, PIP5K1B, PMVK, POGZ, PPIA, PRCC, PRKCA, PRRC2C, PSMB4, PSMD4, RAB29, RCBTB2, SCAMP3, SCAPER, SDHC, SEL1L3, SELP LG, SHC1, SIDT1, SSR2, STAP1, TAP1, TIMM17A, TLR10, TMCO1, TOR1AIP2, TOR3A, TP53INP1, TPM 3, TRANK1, TROVE2, UAP1, UBE2Q1, UBQLN4, UHMK1, VPS45, YY1AP1, ZC3H11A, ZFP36, and ZNF593.
[0044] These 97 genes were selected as classifier genes for classification. Based on the expression data of these 97 genes, 551 MM samples in GSE2658 were clustered into high MCL1-M and low MCL1-M subtypes using consensus clustering. However, clustering-based methods cannot classify individual samples. To enable classification of individual MM samples, the 551 samples were randomly divided into a training set (369 samples) and a validation set (182 samples) in a 2:1 ratio. A stratified sampling method was performed based on the results of consensus clustering to ensure that the proportions of high MCL1-M and MCL1-M samples in the training and validation sets were the same as those in the original dataset.
[0045] Based on the expression data of these 97 classifier genes in 369 samples from the training set and the results of the high or low MCL1-M subtype classification for these samples in the consensus clustering, a Bayesian classifier for assigning individual MM samples to the high or low MCL1-M subtype was trained using the naive Bayes classification algorithm in R in the klaR package.
[0046] The code for the MM Bayes classifier is as follows: options(warn=-1) #Install the machine learning package klaR install.packages(“klaR”) #Load expression data of 97 classifier genes of GSE2658 from file and preprocess library(klaR) i=0 while(TRUE) { GSE2658.data<-read.delim(“gse2658.batch_removed.txt“,row.names=1,stringsAsFactors=T) GSE2658<-apply(GSE2658.data[,-1],1,scale) GSE2658<-t(GSE2658) GSE2658<-data.frame(GSE2658.data[,1],GSE2658) colnames(GSE2658)[1]<-'subtype' colnames(GSE2658)<-colnames(GSE2658.data) rownames(GSE2658)<-rownames(GSE2658.data) # Split the samples into training and validation sets in a 2:1 ratio while(TRUE) { split_train_test<-function(data, ratio){ train_indices<-sample(length(data[,1]),as.integer(length(data[,1])*ratio)) return(train_indices) } train_sets<-GSE2658[split_train_test(GSE2658,0.67),] test_sets<-GSE2658[-split_train_test(GSE2658,0.67), #Build a Naive Bayes classification model using the training set GSE2658.NB<-NaiveBayes(subtype ~ .,data=train_sets,fL=1) if(as.vector(GSE2658.NB$apriori)[1]<0.453&as.vector(GSE2658.NB$apriori[1])>0.451){ break } } #Verifying the performance of a naive Bayes classification model on a validation set results<-predict(GSE2658.NB,test_sets[,-1],threshold=0.1,type='raw') predicted_class<-as.data.frame(results) predicted_class[,2:3]<-apply(predicted_class[,2:3],2,round,3) #Identifying a naive Bayes classification model with accuracy over 97% using cross-labeling method compare_table<-data.frame(predicted_class$class,test_sets$subtype) colnames(compare_table)<-c(”predicted_class”,”original_class”) table<-prop.table(table(compare_table),2) accuracy=c(table[1,1],table[2,2]) if(accuracy[1]>0.95&accuracy[2]>0.95){ break } else { i=i+1 } } print(paste('Both sensitivity and Specifity gets greater than 0.97 at the',i,'th','trial',sep='')) #Subtype prediction based on Bayesian classification on MMRF dataset mmrf.data<-read.delim(”mmrf.batch_removed.txt”,row.names=1,stringsAsFactors=T) mmrf<-apply(mmrf.data[,-1],1,scale) mmrf<-t(mmrf) rownames(mmrf)<-rownames(mmrf.data) colnames(mmrf)<-colnames(mmrf.data)[-1] mmrf<-data.frame(mmrf.data$subtype,mmrf) colnames(mmrf)[1]<-'subtype' results.mmrf<-predict(GSE2658.NB,mmrf[,-1],threshold=0.01,type='raw') predicted_class.mmrf<-as.data.frame(results.mmrf) predicted_class.mmrf[,2:3]<-apply(predicted_class.mmrf[,2:3],2,round,3) compare_table.mmrf<-data.frame(predicted_class.mmrf$class,mmrf.data$subtype) colnames(compare_table.mmrf)<-c(”predicted_class”,”original_class”) prop.table(table(compare_table.mmrf),2) Subtype prediction based on Bayesian classification on the #GSE19784 dataset gse19784.data<-read.delim(”gse19784.batch_removed.txt”,row.names=1,stringsAsFactors=T) gse19784<-apply(gse19784.data[,-1],1,scale) gse19784<-t(gse19784) colnames(gse19784)<-colnames(gse19784.data)[-1] gse19784<-data.frame(gse19784.data$subtype,gse19784) colnames(gse19784)[1]<-'subtype' results_19784<-predict(GSE2658.NB,gse19784[,-1],threshold=0.01,type='raw') predicted_class_19784<-as.data.frame(results_19784) predicted_class_19784[,2:3]<-apply(predicted_class_19784[,2:3],2,round,3) compare_table_19784<-data.frame(predicted_class_19784$class,gse19784$subtype) colnames(compare_table_19784)<-c(”predicted_class”,”original_class”) prop.table(table(compare_table_19784),2)
[0047] Additionally, 182 samples from the validation set were used to assess the accuracy of the classification.
[0048] The Bayesian classification model was optimized using the accuracy results from each run until the accuracy was above 95%. The accuracy results for Bayesian classification on GSE2658 are shown in Table 1, and the ROC curve data are shown in Figure 1. Table 1. Accuracy of the Naive Bayes prediction model for GSE2658 [Table 1]
[0049] To test whether the naive Bayes model developed using data from GSE2658 can be used generally, we used the naive Bayes model on the MM dataset MMRF published by the NCI and the GEO MM dataset GSE19784.
[0050] The MMRF dataset differed from GSE2658 because the expression data were obtained from mRNA-seq. The Bayesian classification model for the MMRF dataset is shown in Table 2, and the ROC curve plot is shown in Figure 2. Table 2. Accuracy of the Naive Bayes prediction model established using the GSE2658 dataset on the MMRF dataset. [Table 2]
[0051] The results show that the classifier can maintain high accuracy even in cross-platform cases, which indicates its high value for promotion and application.
[0052] Similar to GSE2658, the expression data in dataset GSE19784 was also generated using the Affymetrix U133 2.0 and 2.0 platforms. Because GSE19784 was generated by a different research institute at a different time, and the experimental conditions were unlikely to be the same as those for GSE2658, the two datasets may have different dynamics and noise in their gene expression profiles. Naive Bayes prediction results for GSE19784 are shown in Table 3, and ROC curve plots are shown in Figure 3. Accurate classification results were also generated for dataset GSE19784. Table 3. Accuracy of classifiers built using the GSE2658 dataset on the GSE19784 dataset. [Table 3]
[0053] The results show that the classifier can better overcome the above problems and still maintain high accuracy.
[0054] Example 2. Application of naive Bayes prediction model in predicting survival of MM patients.
[0055] I. Dataset GSE2658 Based on the expression data of 97 classifier genes in 551 preprocessed MM samples from the GSE2658 database, the naive Bayes prediction model developed in Example 1 was used to classify the 551 samples, resulting in 249 MCL1-M high MMs and 302 MCL1-M low MMs.
[0056] The follow-up period for 551 MM patients was 72 months. The results of survival analysis (Kaplan-Meier analysis and Cox regression analysis) are shown in Figure 4. Differential survival rates were observed between the high and low MCL1-M subtypes, with the overall survival rate of MM patients with high MCL1-M being significantly lower than that of patients with low MCL1-M (log-rank test, p=0.0201; hazard ratio 1.588, p=0.0212).
[0057] Therefore, based on the expression of 97 classifier genes in the MCL1 gene cluster, a naive Bayes prediction model enabled prognosis prediction for MM patients.
[0058] II. Database MMFR Based on the expression data of 97 classifier genes in 534 pre-treated MM samples (pre-treatment test), molecular classification was performed using the naive Bayes prediction model developed in Example 1. In classifying the 534 samples, 231 high MCL1-M MMs and 303 low MCL1-M MMs were obtained.
[0059] The follow-up period for 534 MM patients was 48 months. The results of survival analysis (Kaplan-Meier analysis and Cox regression analysis) are shown in Figure 5. Differential survival rates were observed between the high and low MCL1-M subtypes, with the overall survival rate of MM patients with high MCL1-M being significantly lower than that of patients with low MCL1-M (log-rank test, p=0.006663; hazard ratio 1.838, p=0.00706).
[0060] Therefore, regardless of the technological platform for detection of expression data, the expression of 97 classifier genes and the naive Bayes prediction model enabled prognosis prediction for MM patients.
[0061] III. Database GSE19784 Molecular classification was performed using the naive Bayes prediction model developed in Example 1 based on the expression data of 97 classifier genes in 304 pre-processed MM samples from the database GSE19784, resulting in 107 MCL1-M high MMs and 196 MCL1-M low MMs.
[0062] The follow-up period for 304 MM patients was 96 months. The results of survival analysis (Kaplan-Meier analysis and Cox regression analysis) are shown in Figure 6 (Panel A: Overall Survival Rate, Panel B: Progression-Free Survival Rate). Survival rates differed between high and low MCL1-M subtypes, with the overall survival rate of MM patients with high MCL1-M being significantly lower than that of patients with low MCL1-M subtypes (log-rank test, p<0.0001; hazard ratio 1.91, p=0.0002). GSE19784 also contains progression-free survival data. Similarly, the progression-free survival rate of MM patients with high MCL1-M was significantly lower than that of patients with low MCL1-M subtypes (log-rank test, p=0.0282; likelihood ratio test, hazard ratio 1.36, p=0.031). These results confirm that the expression of 97 classifier genes and the naive Bayes prediction model enabled prognosis prediction for MM patients.
[0063] Example 3. Molecular diagnostic markers and classification of multiple myeloma predict whether a test patient is treatable with bortezomib.
[0064] Gene expression data in GSE19784 were generated from MM patients enrolled in a randomized phase III clinical trial (HOVON-65 / GMMG-HD4 trial), with treatment details documented for all patients. Patients were randomly assigned to receive either a drug combination known as VAD (vincristine, doxorubicin, and dexamethasone; 155 patients) or a drug combination known as PAD (bortezomib, doxorubicin, and dexamethasone; 148 patients). The difference between the two is that the PAD combination included bortezomib. All expression data were obtained from pretreatment samples.
[0065] Naive Bayes prediction models were used to classify MM samples as MCL1-M high and MCL1-M low subtypes (described in Example 1). Survival analyses (Kaplan-Meier and Cox regression analyses) were performed separately for MCL1-M high (51 PAD-treated, 56 VAD-treated) or MCL1-M low samples (104 PAD-treated, 92 VAD-treated) according to treatment option.
[0066] The results are shown in Figure 7, where Panel A shows the overall survival rate for the high MCL1-M subtype, Panel B shows the overall survival rate for the low MCL1-M subtype, Panel C shows the progression-free survival rate for the high MCL1-M subtype, and Panel D shows the progression-free survival rate for the low MCL1-M subtype. PAD treatment with bortezomib only extended the survival time, especially the progression-free survival time, of MM patients with high MCL1-M (in Figure 7, the left panel shows the high MCL-M subtype, and the right panel shows the low MCL-M subtype; top: overall survival curve, bottom: progression-free survival curve). This indicates that PAD treatment with bortezomib can delay the progression of MM with high MCL1-M, but has no effect on MM patients with low MCL-M. In summary, the present invention enables stratification of MM patients for treatment decisions, thereby avoiding the need to treat MM with low MCL1-M with bortezomib. This reduces the financial burden associated with treatment and prevents side effects from treatment.
[0067] Example 4. Application of naive Bayes prediction models in stratifying MM patients into different risk groups Bone marrow samples from 30 newly diagnosed MM patients were collected at Beijing Chaoyang Hospital. CD138+ cells were purified using anti-CD138 antibody-coated beads and used to generate total RNA. RNA preparations were hybridized to Affymetrix Prime View arrays to detect the expression of 97 classifier genes.
[0068] Consensus clustering was performed to identify high or low MCL1-M samples within the group, and the naive Bayes prediction model developed in Example 1 was performed to individually identify high or low MCL1-M samples.
[0069] As shown in Table 4, the classification results of the consensus clustering and naive Bayes prediction models were highly consistent. Only one case of high MCL1-M MM was predicted as low MCL-M MM, suggesting that the naive Bayes prediction model can be used to predict MM subtypes in individual samples.
[0070] Due to the limited size of the dataset and the short follow-up period, survival analysis was not performed. However, based on conventional risk parameters (existing medical evidence index), 19 cases of MM with high MCL1-M included 14 cases of high-risk MM as defined by conventional risk parameters, and 11 cases of MM with low MCL1-M included only 3 cases of high-risk MM as defined by conventional risk parameters. This demonstrates that the established classification can still predict patient prognosis in this example. Table 4. Accuracy of the classifiers built using the GSE2658 dataset on the collected samples [Table 4] [Industrial Applicability]
[0071] Current molecular classification schemes do not correlate with the cellular origin of MM, nor can they predict treatment efficacy. To understand the pathogenesis and molecular classification of MM, we investigated gene co-expression networks surrounding key signaling pathways in germinal center development. Because gene networks involved in B cell to plasma cell development may play an important role in MM pathogenesis, we screened for dysregulation of these networks. After a series of analyses, we identified gene co-expression modules surrounding MCL1 (MCL1-M) and developed a classification scheme that assigns MM to MCL-high or MCL-low subtypes. These two subtypes differ in their prognosis and patterns of genomic alterations. More importantly, this classification scheme predicts response to bortezomib treatment and correlates with plasma cell development. This invention constitutes a new platform for developing personalized precision therapies for MM and also deepens our understanding of MM pathogenesis.
Claims
1. A method for predicting the prognosis of a multiple myeloma (MM) patient, comprising obtaining or detecting the expression of 97 classifier genes in an MM patient; The 97 classifier genes are ACBD3, ADAR, ADSS, ALDH2, ANP32E, ANXA2, ATF3, ATP8B2, CACYBP, CAPN2, CCND1, CCT3, CDC42SE1, CERS2, CHSY3, CLIC1, CLMN, COPA, CSNK1G3, DAP3, DENND1B, ENSA, EPRS, E PSTI1, EVL, FAM13A, FAM49A, FLAD1, FRZB, GLRX2, HAX1, HDGF, HLA-A, HLA-B, HLA-C, HLA-F, HL A-G, IL6R, ISG20L2, JTB, KLF2, LAMTOR2, LDHA, MCL1, MOXD1, MRPL24, MRPL9, MVP, MYL6, NDUFS 2, NOP58, NOTCH2NL, NTAN1, PAK1, PI4KB, PIEZO1, PIK3AP1, PIM2, PIP5K1B, PMVK, POGZ, PPIA, PRCC, PRKCA, PRRC2C, PSMB4, PSMD4, RAB29, RCBTB2, SCAMP3, SCAPER, SDHC, SEL1L3, SELPLG, S HC1, SIDT1, SSR2, STAP1, TAP1, TIMM17A, TLR10, TMCO1, TOR1AIP2, TOR3A, TP53INP1, TPM3, TRANK1, TROVE2, UAP1, UBE2Q1, UBQLN4, UHMK1, VPS45, YY1AP1, ZC3H11A, ZFP36, and ZNF593; dividing the n MM samples into a training set and a validation set, and performing stratified sampling based on the results of consensus clustering to assign the n MM samples to an MCL1-M high subtype or an MCL1-M low subtype; using a naive Bayes method to construct a naive Bayes prediction model based on the 97 classifier gene expression data, the two subtypes, and the prognostic survival data of the n MM samples; and using the constructed naive Bayes prediction model to predict the prognosis of the MM patient and the therapeutic effect of bortezomib or a combination of drugs containing bortezomib; wherein the constructed naive Bayes prediction model determines whether the MM patient belongs to the MCL1-M high subtype or the MCL1-M low subtype, and the prognosis prediction for the MM patient belonging to the MCL1-M high subtype is significantly poorer than the prognosis prediction for the MM patient belonging to the MCL1-M low subtype; If the MM patient belongs to the high MCL1-M subtype, the method supports guidance on the use of bortezomib or a combination of drugs containing bortezomib for treatment, and if the MM patient belongs to the low MCL1-M subtype, the method supports guidance on not using bortezomib or a drug containing bortezomib for treatment; The predicted efficacy of bortezomib or a bortezomib-containing drug combination for MM patients belonging to the MCL1-M high subtype based on prognostic survival rate or survival duration is superior to the predicted efficacy of bortezomib or a bortezomib-containing drug combination for MM patients belonging to the MCL1-M low subtype.
2. 1. An apparatus for implementing a naive Bayes predictive model of multiple myeloma (MM) for obtaining or detecting expression of 97 classifier genes in multiple myeloma (MM) patients, comprising: The 97 classifier genes are ACBD3, ADAR, ADSS, ALDH2, ANP32E, ANXA2, ATF3, ATP8B2, CACYBP, CAPN2, CCND1, CCT3, CDC42SE1, CERS2, CHSY3, CLIC1, CLMN, COPA, CSNK1G3, DAP3, DENND1B, ENSA, EPRS, E PSTI1, EVL, FAM13A, FAM49A, FLAD1, FRZB, GLRX2, HAX1, HDGF, HLA-A, HLA-B, HLA-C, HLA-F, HL A-G, IL6R, ISG20L2, JTB, KLF2, LAMTOR2, LDHA, MCL1, MOXD1, MRPL24, MRPL9, MVP, MYL6, NDUFS 2, NOP58, NOTCH2NL, NTAN1, PAK1, PI4KB, PIEZO1, PIK3AP1, PIM2, PIP5K1B, PMVK, POGZ, PPIA, PRCC, PRKCA, PRRC2C, PSMB4, PSMD4, RAB29, RCBTB2, SCAMP3, SCAPER, SDHC, SEL1L3, SELPLG, S HC1, SIDT1, SSR2, STAP1, TAP1, TIMM17A, TLR10, TMCO1, TOR1AIP2, TOR3A, TP53INP1, TPM3, TRANK1, TROVE2, UAP1, UBE2Q1, UBQLN4, UHMK1, VPS45, YY1AP1, ZC3H11A, ZFP36, and ZNF593; The device includes a carrier having recorded thereon a program for executing a naive Bayes prediction model and a method for predicting the prognosis of the MM patient and the therapeutic effect of using bortezomib or a combination of drugs containing bortezomib; The naive Bayes prediction model is dividing the n MM samples into a training set and a validation set, and performing stratified sampling based on the results of consensus clustering to assign the n MM samples to an MCL1-M high subtype or an MCL1-M low subtype; using a naive Bayes method to construct a naive Bayes prediction model based on the 97 classifier gene expression data, the two subtypes, and the prognostic survival data of the n MM samples; Constructed by where: The prediction method includes the steps of: acquiring or detecting expression of 97 classifier genes in a test MM patient to obtain expression data of the 97 classifier genes in the test MM patient; and then classifying the expression data of the 97 classifier genes in the test MM patient using the constructed naive Bayes prediction model, wherein prognosis prediction for MM patients belonging to the high MCL1-M subtype is significantly poorer than prognosis prediction for MM patients belonging to the low MCL1-M subtype; If the MM patient to be tested belongs to a high MCL1-M subtype, the method assists in providing guidance on the use of bortezomib or a combination of drugs containing bortezomib for treatment, and if the MM patient to be tested belongs to a low MCL1-M subtype, the method assists in providing guidance on not using bortezomib or a combination of drugs containing bortezomib for treatment; The predicted efficacy of bortezomib or a bortezomib-containing drug combination for MM patients belonging to the MCL1-M high subtype based on prognostic survival rate or survival duration is superior to the predicted efficacy of bortezomib or a bortezomib-containing drug combination for MM patients belonging to the MCL1-M low subtype.
3. 3. The apparatus of claim 2, wherein the test MM patient is a single patient or multiple patients.
4. 1. A method for constructing a naive Bayes predictive model for classifying multiple myeloma (MM) patients, comprising: 1) obtaining expression data for 97 classifier genes in n MM samples; 2) assigning MM samples to MCL1-M high or MCL1-M low subtypes by consensus clustering; 3) using a naive Bayes method to construct a naive Bayes prediction model based on the two subtypes of step 2), the 97 classifier gene expression data of the n MM samples of step 1), and the prognosis survival data of the n MM samples; The 97 classifier genes are ACBD3, ADAR, ADSS, ALDH2, ANP32E, ANXA2, ATF3, ATP8B2, CACYBP, CAPN2, CCND1, CCT3, CDC42SE1, CERS2, CHSY3, CLIC1, CLMN, COPA, CSNK1G3, DAP3, DENND1B, ENSA, EPRS, E PSTI1, EVL, FAM13A, FAM49A, FLAD1, FRZB, GLRX2, HAX1, HDGF, HLA-A, HLA-B, HLA-C, HLA-F, HL A-G, IL6R, ISG20L2, JTB, KLF2, LAMTOR2, LDHA, MCL1, MOXD1, MRPL24, MRPL9, MVP, MYL6, NDUFS 2, NOP58, NOTCH2NL, NTAN1, PAK1, PI4KB, PIEZO1, PIK3AP1, PIM2, PIP5K1B, PMVK, POGZ, PPIA, PRCC, PRKCA, PRRC2C, PSMB4, PSMD4, RAB29, RCBTB2, SCAMP3, SCAPER, SDHC, SEL1L3, SELPLG, S HC1, SIDT1, SSR2, STAP1, TAP1, TIMM17A, TLR10, TMCO1, TOR1AIP2, TOR3A, TP53INP1, TPM3, TRANK1, TROVE2, UAP1, UBE2Q1, UBQLN4, UHMK1, VPS45, YY1AP1, ZC3H11A, ZFP36, and ZNF593; wherein the constructed naive Bayes prediction model determines whether the MM patient belongs to the MCL1-M high subtype or the MCL1-M low subtype, and the prognosis prediction for the MM patient belonging to the MCL1-M high subtype is significantly poorer than the prognosis prediction for the MM patient belonging to the MCL1-M low subtype; If the MM patient belongs to the high MCL1-M subtype, the method supports guidance on the use of bortezomib or a combination of drugs containing bortezomib for treatment, and if the MM patient belongs to the low MCL1-M subtype, the method supports guidance on not using bortezomib or a drug containing bortezomib for treatment; A method in which the therapeutic effect of using bortezomib or a combination of drugs containing bortezomib extends the prognostic survival rate or survival period of MM patients belonging to the MCL1-M high subtype.
Citation Information
Patent Citations
Novel classification index for the molecular classification of multiple myeloma
JP2014520540A
Transcriptional classification and prediction of drug response (t-cap dr)
US20170159130A1