Computer device for predicting disease progression of multiple myeloma patient

By constructing a single-sample prediction algorithm based on the gene expression level data of CD138+ plasma cell samples in multiple myeloma patients, the problem of ineffective evaluation of the treatment response of multiple myeloma patients in the prior art is solved, and individualized evaluation and comparison of disease progression and treatment effects are achieved.

CN120565079AActive Publication Date: 2025-08-29BEIJING NORMAL UNIVERSITY
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510729832.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-08-29
Estimated Expiration
2045-06-03

Smart Images

  • Figure CN120565079A_ABST
    Figure CN120565079A_ABST
Patent Text Reader

Abstract

The invention discloses a computer device for predicting disease progression of a patient with multiple myeloma. The invention provides a data processing device, which comprises a memory, a processor and a computer program stored in the memory, and the computer program implements the following steps: S1, receiving mth sample data and nth sample data of a subject of the same multiple myeloma patient; s2, respectively inputting the mth sample data and the nth sample data of the same subject into a single sample prediction algorithm, and outputting an mth detection myeloma classification score (MCS) value and an nth detection MCS value; s3, subtracting the mth detection MCS value from the nth detection MCS value to obtain delta MCS, and grouping or comparing the disease progression of the subject according to the delta MCS. The single sample prediction algorithm and the delta MCS can be used for grouping or comparing disease progression states of multiple myeloma patient subjects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer information and relates to a computer device for predicting disease progression of multiple myeloma patients. Background Art

[0002] Multiple myeloma (MM) is the second most common hematologic malignancy, accounting for 10% of all hematologic malignancies. MM patients exhibit significant heterogeneity in prognosis and response to standardized treatment regimens. Consequently, numerous studies have attempted to develop molecular classification strategies to guide clinical treatment.

[0003] In recent years, numerous clinical trials have demonstrated at the population level the significant success of existing combination therapy regimens in prolonging survival in MM patients. However, individual patient responses to combination therapy are highly heterogeneous, and existing diagnostic and molecular pathology testing methods are unable to effectively predict whether an individual patient will respond to a combination therapy and the extent of their response. Therefore, biomarker-based selection of appropriate drug combinations is crucial for the survival of MM patients.

[0004] During the treatment of MM patients, timely and effective efficacy evaluation is particularly important for evaluating the response of MM patients to the current treatment plan and adjusting the subsequent treatment plan.

[0005] To date, MM remains incurable, and almost all MM patients will eventually relapse. Therefore, the development of highly sensitive efficacy assessment markers is of great practical significance for the treatment of MM patients. Current efficacy assessment is mainly based on the patient's tumor burden and the degree of organ damage. However, the residual plasma cell clone components in MM patients are closely related to the patient's relapse trend. Therefore, establishing an efficacy assessment standard based on gene expression profiles can more sensitively capture the overall expression signals at the bulk (tissue or cell population) level of MM patients, thereby providing MM patients with a new disease monitoring solution from the perspective of plasma cell development.

[0006] Therefore, there is an urgent need in the field for personalized testing tools that can predict disease progression and treatment response in MM patients. Summary of the Invention

[0007] The technical problem solved by the present invention is to provide a computer device for predicting disease progression in patients with multiple myeloma.

[0008] In order to solve the above technical problems, the present invention provides, in a first aspect, a data processing device for grouping disease progression in multiple myeloma patients, comprising a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the following steps:

[0009] S1. Receive data: Receive the mth sample data and the nth sample data of the same multiple myeloma patient, where n is greater than m;

[0010] The m sample data are ACBD3, ADAR, ADSS, ALDH2, ANP32E, ANXA2, ATF3, ATP8B2, CACYBP, CAPN2, CCND1, CCT3, CDC42SE1, CERS2, CHSY3, CLIC1, CLMN, COPA, CSNK1G3, DAP3, DENND1B, ENSA, EPRS, EPSTI1, EVL, FAM13A, FAM49A, FLAD1, FRZB, GLRX2, HAX1, HDGF, HLA-A, HLA-B, HLA-C, HLA-F, HLA-G, IL6R, ISG20L2, JTB, KLF2, LAMTOR2, LDHA, MCL1, MOXD1, MRPL24, MRPL9, MVP, Transcriptome expression data of 97 genes: MYL6, NDUFS2, NOP58, NOTCH2NL, NTAN1, PAK1, PI4KB, PIEZO1, PIK3AP1, PIM2, PIP5K1B, PMVK, POGZ, PPIA, PRCC, PRKCA, PRRC2C, PSMB4, PSMD4, RAB29, RCBTB2, SCAMP3, SCAPER, SDHC, SEL1L3, SELPLG, SHC1, SIDT1, SSR2, STAP1, TAP1, TIMM17A, TLR10, TMCO1, TOR1AIP2, TOR3A, TP53INP1, TPM3, TRANK1, TROVE2, UAP1, UBE2Q1, UBQLN4, UHMK1, VPS45, YY1AP1, ZC3H11A, ZFP36, and ZNF593;

[0011] The nth sample data is the transcriptome expression data of the 97 genes in the ex vivo bone marrow CD138+ plasma cell sample of the multiple myeloma patient during the nth test;

[0012] S2. Data processing:

[0013] Inputting the m-th sample data and the n-th sample data of the same multiple myeloma patient subject into the single sample prediction algorithm respectively, outputting the myeloma classification score of the m-th detection and the myeloma classification score of the n-th detection of the multiple myeloma patient subject, which are recorded as the MCS (Myeloma classification score) value of the m-th detection and the MCS value of the n-th detection respectively;

[0014] The single sample prediction algorithm is constructed according to a method comprising the following steps:

[0015] 1) Multiple ex vivo bone marrow CD138+ plasma cell samples from multiple myeloma patients were used as training sets, and the transcriptome expression data of the 97 genes in each sample in the training set was used as training set data. The training sets were divided into two groups using an unsupervised clustering method;

[0016] 2) using the grouping labels of each sample in each group obtained in 1) and the transcriptome expression data of the 97 genes of the corresponding sample as input data, training a support vector machine (SVM) model to generate a myeloma classification score for each sample in the training set, recorded as the MCS value;

[0017] 3) Using the transcriptome expression data of 97 genes in all samples of the training set as the explanatory variable and the MCS values ​​of all samples in the training set as the response variable, a single sample prediction (SSP) algorithm based on the multiple linear regression (MLR) model was constructed;

[0018] S3. Output results:

[0019] The value obtained by subtracting the MCS value of the mth test from the MCS value of the same multiple myeloma patient subject at the nth test is recorded as ΔMCS, and the disease progression of the multiple myeloma patient subject is grouped according to ΔMCS.

[0020] In the data processing device of the first aspect above, the disease progression of the multiple myeloma patient subjects may be grouped as follows: subjects with ΔMCS < -2 are defined as the MCS progression group, and subjects with ΔMCS ≥ -2 are defined as the MCS stable group;

[0021] MM patients in the MCS progression group were in a progressive state (the genomic instability of the nth test sample increased and / or the prognosis became worse compared with the mth test sample);

[0022] The MM patients in the MCS stable group were in a stable condition (the genomic status and / or prognosis of the nth test sample did not change significantly compared with the mth test sample).

[0023] The disease progression in the MCS stable group was slower than that in the MCS progressive group.

[0024] Specifically reflected in:

[0025] 1) The prognosis of the MCS stable group was better than that of the MCS progressive group;

[0026] 2) The genomic stability of the MCS stable group was better than that of the MCS progressive group.

[0027] The subject may be more than one subject.

[0028] In a second aspect, the present invention provides a data processing device for comparing disease progression in multiple myeloma patients, comprising a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the following steps:

[0029] S1. Receive data: Receive the mth sample data and the nth sample data of the same multiple myeloma patient, where n is greater than m;

[0030] The m sample data are ACBD3, ADAR, ADSS, ALDH2, ANP32E, ANXA2, ATF3, ATP8B2, CACYBP, CAPN2, CCND1, CCT3, CDC42SE1, CERS2, CHSY3, CLIC1, CLMN, COPA, CSNK1G3, DAP3, DENND1B, ENSA, EPRS, EPSTI1, EVL, FAM13A, FAM49A, FLAD1, FRZB, GLRX2, HAX1, HDGF, HLA-A, HLA-B, HLA-C, HLA-F, HLA-G, IL6R, ISG20L2, JTB, KLF2, LAMTOR2, LDHA, MCL1, MOXD1, MRPL24, MRPL9, MVP, Transcriptome expression data of 97 genes: MYL6, NDUFS2, NOP58, NOTCH2NL, NTAN1, PAK1, PI4KB, PIEZO1, PIK3AP1, PIM2, PIP5K1B, PMVK, POGZ, PPIA, PRCC, PRKCA, PRRC2C, PSMB4, PSMD4, RAB29, RCBTB2, SCAMP3, SCAPER, SDHC, SEL1L3, SELPLG, SHC1, SIDT1, SSR2, STAP1, TAP1, TIMM17A, TLR10, TMCO1, TOR1AIP2, TOR3A, TP53INP1, TPM3, TRANK1, TROVE2, UAP1, UBE2Q1, UBQLN4, UHMK1, VPS45, YY1AP1, ZC3H11A, ZFP36, and ZNF593;

[0031] The nth sample data is the transcriptome expression data of the 97 genes in the ex vivo bone marrow CD138+ plasma cell sample of the multiple myeloma patient during the nth test;

[0032] S2. Data processing:

[0033] Inputting the m-th sample data and the n-th sample data of the same multiple myeloma patient subject into the single sample prediction algorithm respectively, and outputting the myeloma classification score of the m-th detection and the myeloma classification score of the multiple myeloma patient subject, which are recorded as the MCS value of the m-th detection and the MCS value of the n-th detection respectively;

[0034] The single sample prediction algorithm is constructed according to a method comprising the following steps:

[0035] 1) Multiple ex vivo bone marrow CD138+ plasma cell samples from multiple myeloma patients were used as training sets, and the transcriptome expression data of the 97 genes in each sample in the training set was used as training set data. The training sets were divided into two groups using an unsupervised clustering method;

[0036] 2) using the grouping labels of each sample in each group obtained in 1) and the transcriptome expression data of the 97 genes of the corresponding sample as input data, training a support vector machine model to generate a myeloma classification score for each sample in the training set, recorded as an MCS value;

[0037] 3) Using the transcriptome expression data of 97 genes of all samples in the training set as the explanatory variable and the MCS values ​​of all samples in the training set as the response variable, a single-sample prediction algorithm based on the multivariate linear regression model was constructed;

[0038] S3. Output results:

[0039] The value obtained by subtracting the MCS value of the mth test from the MCS value of the same multiple myeloma patient subject at the nth test is recorded as ΔMCS, and the disease progression of the multiple myeloma patient subjects is compared based on ΔMCS.

[0040] In the data processing device of the second aspect above, the comparison of the disease progression of the multiple myeloma patient subjects may be as follows: the disease progression of subjects with ΔMCS ≥ -2 is slower than that of subjects with ΔMCS < -2;

[0041] Specifically reflected in:

[0042] 1) The prognosis of the subjects with ΔMCS ≥ -2 is better than that of the subjects with ΔMCS < -2;

[0043] 2) The genomic stability of the subjects with ΔMCS ≥ -2 was better than that of the subjects with ΔMCS < -2.

[0044] The subjects may be two or more subjects.

[0045] In a third aspect, the present invention provides a data processing device for comparing responses of multiple myeloma patients to combination therapy regimens, comprising a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the following steps:

[0046] S1. Receive data: Receive the mth sample data and the nth sample data of the same multiple myeloma patient, where n is greater than m;

[0047] The m sample data are ACBD3, ADAR, ADSS, ALDH2, ANP32E, ANXA2, ATF3, ATP8B2, CACYBP, CAPN2, CCND1, CCT3, CDC42SE1, CERS2, CHSY3, CLIC1, CLMN, COPA, CSNK1G3, DAP3, DENND1B, ENSA, EPRS, EPSTI1, EVL, FAM13A, FAM49A, FLAD1, FRZB, GLRX2, HAX1, HDGF, HLA-A, HLA-B, HLA-C, HLA-F, HLA-G, IL6R, ISG20L2, JTB, KLF2, LAMTOR2, LDHA, MCL1, MOXD1, MRPL24, MRPL9, MVP, Transcriptome expression data of 97 genes: MYL6, NDUFS2, NOP58, NOTCH2NL, NTAN1, PAK1, PI4KB, PIEZO1, PIK3AP1, PIM2, PIP5K1B, PMVK, POGZ, PPIA, PRCC, PRKCA, PRRC2C, PSMB4, PSMD4, RAB29, RCBTB2, SCAMP3, SCAPER, SDHC, SEL1L3, SELPLG, SHC1, SIDT1, SSR2, STAP1, TAP1, TIMM17A, TLR10, TMCO1, TOR1AIP2, TOR3A, TP53INP1, TPM3, TRANK1, TROVE2, UAP1, UBE2Q1, UBQLN4, UHMK1, VPS45, YY1AP1, ZC3H11A, ZFP36, and ZNF593;

[0048] The nth sample data is the transcriptome expression data of the 97 genes in the ex vivo bone marrow CD138+ plasma cell sample of the multiple myeloma patient during the nth test;

[0049] S2. Data processing:

[0050] Inputting the m-th sample data and the n-th sample data of the same multiple myeloma patient subject into the single sample prediction algorithm respectively, and outputting the myeloma classification score of the m-th detection and the myeloma classification score of the multiple myeloma patient subject, which are recorded as the MCS value of the m-th detection and the MCS value of the n-th detection respectively;

[0051] The single sample prediction algorithm is constructed according to a method comprising the following steps:

[0052] 1) Multiple ex vivo bone marrow CD138+ plasma cell samples from multiple myeloma patients were used as training sets, and the transcriptome expression data of the 97 genes in each sample in the training set was used as training set data. The training sets were divided into two groups using an unsupervised clustering method;

[0053] 2) using the grouping labels of each sample in each group obtained in 1) and the transcriptome expression data of the 97 genes of the corresponding sample as input data, training a support vector machine model to generate a myeloma classification score for each sample in the training set, recorded as an MCS value;

[0054] 3) Using the transcriptome expression data of 97 genes of all samples in the training set as the explanatory variable and the MCS values ​​of all samples in the training set as the response variable, a single-sample prediction algorithm based on the multivariate linear regression model was constructed;

[0055] S3. Output results:

[0056] The value obtained by subtracting the MCS value of the mth test from the MCS value of the same multiple myeloma patient subject at the nth test is recorded as ΔMCS, and the response of the multiple myeloma patient subject to the combination treatment regimen is compared based on ΔMCS.

[0057] In the data processing device of the third aspect above, the response of the multiple myeloma patient subjects to the combination treatment regimen is compared, and the smaller the ΔMCS value, the worse the response of the multiple myeloma patient subjects to the combination treatment regimen. Furthermore, the multiple myeloma patient subjects with a large ΔMCS value respond better to the combination treatment regimen than the multiple myeloma patient subjects with a small ΔMCS value.

[0058] The subjects may be two or more subjects.

[0059] In a fourth aspect, the present invention provides a computer program product, comprising a computer program, characterized in that when the computer program is executed by a processor, the steps in the data processing device described in any one of the first to third aspects are implemented.

[0060] In a fifth aspect, the present invention provides a computer-readable storage medium storing a computer program, wherein the computer program enables a computer to execute the steps in the data processing device in any one of the first to third aspects.

[0061] In a sixth aspect, the present invention provides any one of the following devices A1-A3:

[0062] A1. A device for grouping disease progression in multiple myeloma patients, comprising:

[0063] S1, data receiving module: used to receive the mth sample data and the nth sample data of the same multiple myeloma patient subject, where n is greater than m;

[0064] The m sample data are ACBD3, ADAR, ADSS, ALDH2, ANP32E, ANXA2, ATF3, ATP8B2, CACYBP, CAPN2, CCND1, CCT3, CDC42SE1, CERS2, CHSY3, CLIC1, CLMN, COPA, CSNK1G3, DAP3, DENND1B, ENSA, EPRS, EPSTI1, EVL, FAM13A, FAM49A, FLAD1, FRZB, GLRX2, HAX1, HDGF, HLA-A, HLA-B, HLA-C, HLA-F, HLA-G, IL6R, ISG20L2, JTB, KLF2, LAMTOR2, LDHA, MCL1, MOXD1, MRPL24, MRPL9, MVP, Transcriptome expression data of 97 genes: MYL6, NDUFS2, NOP58, NOTCH2NL, NTAN1, PAK1, PI4KB, PIEZO1, PIK3AP1, PIM2, PIP5K1B, PMVK, POGZ, PPIA, PRCC, PRKCA, PRRC2C, PSMB4, PSMD4, RAB29, RCBTB2, SCAMP3, SCAPER, SDHC, SEL1L3, SELPLG, SHC1, SIDT1, SSR2, STAP1, TAP1, TIMM17A, TLR10, TMCO1, TOR1AIP2, TOR3A, TP53INP1, TPM3, TRANK1, TROVE2, UAP1, UBE2Q1, UBQLN4, UHMK1, VPS45, YY1AP1, ZC3H11A, ZFP36, and ZNF593;

[0065] The nth sample data is the transcriptome expression data of the 97 genes in the ex vivo bone marrow CD138+ plasma cell sample of the multiple myeloma patient during the nth test;

[0066] S2, a data processing module: used to input the m-th sample data and the n-th sample data of the same multiple myeloma patient subject into the single sample prediction algorithm, and output the myeloma classification score of the m-th detection and the myeloma classification score of the multiple myeloma patient subject, which are recorded as the MCS value of the m-th detection and the MCS value of the n-th detection respectively;

[0067] The single sample prediction algorithm is constructed according to a method comprising the following steps:

[0068] 1) Multiple ex vivo bone marrow CD138+ plasma cell samples from multiple myeloma patients were used as training sets, and the transcriptome expression data of the 97 genes in each sample in the training set was used as training set data. The training sets were divided into two groups using an unsupervised clustering method;

[0069] 2) using the grouping labels of each sample in each group obtained in 1) and the transcriptome expression data of the 97 genes of the corresponding sample as input data, training a support vector machine model to generate a myeloma classification score for each sample in the training set, recorded as an MCS value;

[0070] 3) Using the transcriptome expression data of 97 genes of all samples in the training set as the explanatory variable and the MCS values ​​of all samples in the training set as the response variable, a single-sample prediction algorithm based on the multivariate linear regression model was constructed;

[0071] S3. Output results:

[0072] The value obtained by subtracting the MCS value of the mth test from the MCS value of the same multiple myeloma patient subject at the nth test is recorded as ΔMCS, and the disease progression of the multiple myeloma patient subject is grouped according to ΔMCS.

[0073] A2. A device for comparing disease progression in multiple myeloma patient subjects, the device comprising:

[0074] S1, data receiving module: used to receive the mth sample data and the nth sample data of the same multiple myeloma patient subject, where n is greater than m;

[0075] The m sample data are ACBD3, ADAR, ADSS, ALDH2, ANP32E, ANXA2, ATF3, ATP8B2, CACYBP, CAPN2, CCND1, CCT3, CDC42SE1, CERS2, CHSY3, CLIC1, CLMN, COPA, CSNK1G3, DAP3, DENND1B, ENSA, EPRS, EPSTI1, EVL, FAM13A, FAM49A, FLAD1, FRZB, GLRX2, HAX1, HDGF, HLA-A, HLA-B, HLA-C, HLA-F, HLA-G, IL6R, ISG20L2, JTB, KLF2, LAMTOR2, LDHA, MCL1, MOXD1, MRPL24, MRPL9, MVP, Transcriptome expression data of 97 genes: MYL6, NDUFS2, NOP58, NOTCH2NL, NTAN1, PAK1, PI4KB, PIEZO1, PIK3AP1, PIM2, PIP5K1B, PMVK, POGZ, PPIA, PRCC, PRKCA, PRRC2C, PSMB4, PSMD4, RAB29, RCBTB2, SCAMP3, SCAPER, SDHC, SEL1L3, SELPLG, SHC1, SIDT1, SSR2, STAP1, TAP1, TIMM17A, TLR10, TMCO1, TOR1AIP2, TOR3A, TP53INP1, TPM3, TRANK1, TROVE2, UAP1, UBE2Q1, UBQLN4, UHMK1, VPS45, YY1AP1, ZC3H11A, ZFP36, and ZNF593;

[0076] The nth sample data is the transcriptome expression data of the 97 genes in the ex vivo bone marrow CD138+ plasma cell sample of the multiple myeloma patient during the nth test;

[0077] S2, a data processing module: used to input the m-th sample data and the n-th sample data of the same multiple myeloma patient subject into the single sample prediction algorithm, and output the myeloma classification score of the m-th detection and the myeloma classification score of the multiple myeloma patient subject, which are recorded as the MCS value of the m-th detection and the MCS value of the n-th detection respectively;

[0078] The single sample prediction algorithm is constructed according to a method comprising the following steps:

[0079] 1) Multiple ex vivo bone marrow CD138+ plasma cell samples from multiple myeloma patients were used as training sets, and the transcriptome expression data of the 97 genes in each sample in the training set was used as training set data. The training sets were divided into two groups using an unsupervised clustering method;

[0080] 2) using the grouping labels of each sample in each group obtained in 1) and the transcriptome expression data of the 97 genes of the corresponding sample as input data, training a support vector machine model to generate a myeloma classification score for each sample in the training set, recorded as an MCS value;

[0081] 3) Using the transcriptome expression data of 97 genes of all samples in the training set as the explanatory variable and the MCS values ​​of all samples in the training set as the response variable, a single-sample prediction algorithm based on the multivariate linear regression model was constructed;

[0082] S3. Output results:

[0083] The value obtained by subtracting the MCS value of the mth test from the MCS value of the same multiple myeloma patient subject at the nth test is recorded as ΔMCS, and the disease progression of the multiple myeloma patient subjects is compared based on ΔMCS.

[0084] A3. A device for comparing responses of multiple myeloma patient subjects to a combination therapy regimen, the device comprising:

[0085] S1, data receiving module: used to receive the mth sample data and the nth sample data of the same multiple myeloma patient subject, where n is greater than m;

[0086] The m sample data are ACBD3, ADAR, ADSS, ALDH2, ANP32E, ANXA2, ATF3, ATP8B2, CACYBP, CAPN2, CCND1, CCT3, CDC42SE1, CERS2, CHSY3, CLIC1, CLMN, COPA, CSNK1G3, DAP3, DENND1B, ENSA, EPRS, EPSTI1, EVL, FAM13A, FAM49A, FLAD1, FRZB, GLRX2, HAX1, HDGF, HLA-A, HLA-B, HLA-C, HLA-F, HLA-G, IL6R, ISG20L2, JTB, KLF2, LAMTOR2, LDHA, MCL1, MOXD1, MRPL24, MRPL9, MVP, Transcriptome expression data of 97 genes: MYL6, NDUFS2, NOP58, NOTCH2NL, NTAN1, PAK1, PI4KB, PIEZO1, PIK3AP1, PIM2, PIP5K1B, PMVK, POGZ, PPIA, PRCC, PRKCA, PRRC2C, PSMB4, PSMD4, RAB29, RCBTB2, SCAMP3, SCAPER, SDHC, SEL1L3, SELPLG, SHC1, SIDT1, SSR2, STAP1, TAP1, TIMM17A, TLR10, TMCO1, TOR1AIP2, TOR3A, TP53INP1, TPM3, TRANK1, TROVE2, UAP1, UBE2Q1, UBQLN4, UHMK1, VPS45, YY1AP1, ZC3H11A, ZFP36, and ZNF593;

[0087] The nth sample data is the transcriptome expression data of the 97 genes in the ex vivo bone marrow CD138+ plasma cell sample of the multiple myeloma patient during the nth test;

[0088] S2, a data processing module: used to input the m-th sample data and the n-th sample data of the same multiple myeloma patient subject into the single sample prediction algorithm, and output the myeloma classification score of the m-th detection and the myeloma classification score of the multiple myeloma patient subject, which are recorded as the MCS value of the m-th detection and the MCS value of the n-th detection respectively;

[0089] The single sample prediction algorithm is constructed according to a method comprising the following steps:

[0090] 1) Multiple ex vivo bone marrow CD138+ plasma cell samples from multiple myeloma patients were used as training sets, and the transcriptome expression data of the 97 genes in each sample in the training set was used as training set data. The training sets were divided into two groups using an unsupervised clustering method;

[0091] 2) using the grouping labels of each sample in each group obtained in 1) and the transcriptome expression data of the 97 genes of the corresponding sample as input data, training a support vector machine model to generate a myeloma classification score for each sample in the training set, recorded as an MCS value;

[0092] 3) Using the transcriptome expression data of 97 genes of all samples in the training set as the explanatory variable and the MCS values ​​of all samples in the training set as the response variable, a single-sample prediction algorithm based on the multivariate linear regression model was constructed;

[0093] S3. Output results:

[0094] The value obtained by subtracting the MCS value of the mth test from the MCS value of the same multiple myeloma patient subject at the nth test is recorded as ΔMCS, and the response of the multiple myeloma patient subject to the combination treatment regimen is compared based on ΔMCS.

[0095] In a seventh aspect, the present invention provides any one of the following methods B1-B3:

[0096] B1. A method for grouping disease progression in multiple myeloma patients, comprising the following steps:

[0097] S1. Obtaining sample data, for receiving the mth sample data and the nth sample data of the same multiple myeloma patient subject, where n is greater than m;

[0098] The m sample data are ACBD3, ADAR, ADSS, ALDH2, ANP32E, ANXA2, ATF3, ATP8B2, CACYBP, CAPN2, CCND1, CCT3, CDC42SE1, CERS2, CHSY3, CLIC1, CLMN, COPA, CSNK1G3, DAP3, DENND1B, ENSA, EPRS, EPSTI1, EVL, FAM13A, FAM49A, FLAD1, FRZB, GLRX2, HAX1, HDGF, HLA-A, HLA-B, HLA-C, HLA-F, HLA-G, IL6R, ISG20L2, JTB, KLF2, LAMTOR2, LDHA, MCL1, MOXD1, MRPL24, MRPL9, MVP, Transcriptome expression data of 97 genes: MYL6, NDUFS2, NOP58, NOTCH2NL, NTAN1, PAK1, PI4KB, PIEZO1, PIK3AP1, PIM2, PIP5K1B, PMVK, POGZ, PPIA, PRCC, PRKCA, PRRC2C, PSMB4, PSMD4, RAB29, RCBTB2, SCAMP3, SCAPER, SDHC, SEL1L3, SELPLG, SHC1, SIDT1, SSR2, STAP1, TAP1, TIMM17A, TLR10, TMCO1, TOR1AIP2, TOR3A, TP53INP1, TPM3, TRANK1, TROVE2, UAP1, UBE2Q1, UBQLN4, UHMK1, VPS45, YY1AP1, ZC3H11A, ZFP36, and ZNF593;

[0099] The nth sample data is the transcriptome expression data of the 97 genes in the ex vivo bone marrow CD138+ plasma cell sample of the multiple myeloma patient during the nth test;

[0100] S2. Inputting the m-th sample data and the n-th sample data of the same multiple myeloma patient subject into the single sample prediction algorithm, respectively, and outputting the m-th detection myeloma classification score and the n-th detection myeloma classification score of the multiple myeloma patient subject, respectively recorded as the MCS value of the m-th detection and the MCS value of the n-th detection;

[0101] The single sample prediction algorithm is constructed according to a method comprising the following steps:

[0102] 1) Multiple ex vivo bone marrow CD138+ plasma cell samples from multiple myeloma patients were used as training sets, and the transcriptome expression data of the 97 genes in each sample in the training set was used as training set data. The training sets were divided into two groups using an unsupervised clustering method;

[0103] 2) using the grouping labels of each sample in each group obtained in 1) and the transcriptome expression data of the 97 genes of the corresponding sample as input data, training a support vector machine model to generate a myeloma classification score for each sample in the training set, recorded as an MCS value;

[0104] 3) Using the transcriptome expression data of 97 genes of all samples in the training set as the explanatory variable and the MCS values ​​of all samples in the training set as the response variable, a single-sample prediction algorithm based on the multivariate linear regression model was constructed;

[0105] S3. Subtract the MCS value of the m-th test from the MCS value of the same multiple myeloma patient subject at the n-th test as ΔMCS, and group the disease progression of the multiple myeloma patient subject according to ΔMCS.

[0106] B2. A method for comparing disease progression in multiple myeloma patients, comprising the steps of:

[0107] S1. Obtaining sample data, for receiving the mth sample data and the nth sample data of the same multiple myeloma patient subject, where n is greater than m;

[0108] The m sample data are ACBD3, ADAR, ADSS, ALDH2, ANP32E, ANXA2, ATF3, ATP8B2, CACYBP, CAPN2, CCND1, CCT3, CDC42SE1, CERS2, CHSY3, CLIC1, CLMN, COPA, CSNK1G3, DAP3, DENND1B, ENSA, EPRS, EPSTI1, EVL, FAM13A, FAM49A, FLAD1, FRZB, GLRX2, HAX1, HDGF, HLA-A, HLA-B, HLA-C, HLA-F, HLA-G, IL6R, ISG20L2, JTB, KLF2, LAMTOR2, LDHA, MCL1, MOXD1, MRPL24, MRPL9, MVP, Transcriptome expression data of 97 genes: MYL6, NDUFS2, NOP58, NOTCH2NL, NTAN1, PAK1, PI4KB, PIEZO1, PIK3AP1, PIM2, PIP5K1B, PMVK, POGZ, PPIA, PRCC, PRKCA, PRRC2C, PSMB4, PSMD4, RAB29, RCBTB2, SCAMP3, SCAPER, SDHC, SEL1L3, SELPLG, SHC1, SIDT1, SSR2, STAP1, TAP1, TIMM17A, TLR10, TMCO1, TOR1AIP2, TOR3A, TP53INP1, TPM3, TRANK1, TROVE2, UAP1, UBE2Q1, UBQLN4, UHMK1, VPS45, YY1AP1, ZC3H11A, ZFP36, and ZNF593;

[0109] The nth sample data is the transcriptome expression data of the 97 genes in the ex vivo bone marrow CD138+ plasma cell sample of the multiple myeloma patient during the nth test;

[0110] S2. Inputting the m-th sample data and the n-th sample data of the same multiple myeloma patient subject into the single sample prediction algorithm, respectively, and outputting the m-th detection myeloma classification score and the n-th detection myeloma classification score of the multiple myeloma patient subject, respectively recorded as the MCS value of the m-th detection and the MCS value of the n-th detection;

[0111] The single sample prediction algorithm is constructed according to a method comprising the following steps:

[0112] 1) Multiple ex vivo bone marrow CD138+ plasma cell samples from multiple myeloma patients were used as training sets, and the transcriptome expression data of the 97 genes in each sample in the training set was used as training set data. The training sets were divided into two groups using an unsupervised clustering method;

[0113] 2) using the grouping labels of each sample in each group obtained in 1) and the transcriptome expression data of the 97 genes of the corresponding sample as input data, training a support vector machine model to generate a myeloma classification score for each sample in the training set, recorded as an MCS value;

[0114] 3) Using the transcriptome expression data of 97 genes of all samples in the training set as the explanatory variable and the MCS values ​​of all samples in the training set as the response variable, a single-sample prediction algorithm based on the multivariate linear regression model was constructed;

[0115] S3. Subtract the MCS value of the m-th test from the MCS value of the same multiple myeloma patient subject at the n-th test as ΔMCS, and compare the disease progression of the multiple myeloma patient subject based on ΔMCS.

[0116] B3. A method for comparing the responses of multiple myeloma patients to a combination therapy regimen, the method comprising the steps of:

[0117] S1. Obtaining sample data, for receiving the mth sample data and the nth sample data of the same multiple myeloma patient subject, where n is greater than m;

[0118] The m sample data are ACBD3, ADAR, ADSS, ALDH2, ANP32E, ANXA2, ATF3, ATP8B2, CACYBP, CAPN2, CCND1, CCT3, CDC42SE1, CERS2, CHSY3, CLIC1, CLMN, COPA, CSNK1G3, DAP3, DENND1B, ENSA, EPRS, EPSTI1, EVL, FAM13A, FAM49A, FLAD1, FRZB, GLRX2, HAX1, HDGF, HLA-A, HLA-B, HLA-C, HLA-F, HLA-G, IL6R, ISG20L2, JTB, KLF2, LAMTOR2, LDHA, MCL1, MOXD1, MRPL24, MRPL9, MVP, Transcriptome expression data of 97 genes: MYL6, NDUFS2, NOP58, NOTCH2NL, NTAN1, PAK1, PI4KB, PIEZO1, PIK3AP1, PIM2, PIP5K1B, PMVK, POGZ, PPIA, PRCC, PRKCA, PRRC2C, PSMB4, PSMD4, RAB29, RCBTB2, SCAMP3, SCAPER, SDHC, SEL1L3, SELPLG, SHC1, SIDT1, SSR2, STAP1, TAP1, TIMM17A, TLR10, TMCO1, TOR1AIP2, TOR3A, TP53INP1, TPM3, TRANK1, TROVE2, UAP1, UBE2Q1, UBQLN4, UHMK1, VPS45, YY1AP1, ZC3H11A, ZFP36, and ZNF593;

[0119] The nth sample data is the transcriptome expression data of the 97 genes in the ex vivo bone marrow CD138+ plasma cell sample of the multiple myeloma patient during the nth test;

[0120] S2. Inputting the m-th sample data and the n-th sample data of the same multiple myeloma patient subject into the single sample prediction algorithm, respectively, and outputting the m-th detection myeloma classification score and the n-th detection myeloma classification score of the multiple myeloma patient subject, respectively recorded as the MCS value of the m-th detection and the MCS value of the n-th detection;

[0121] The single sample prediction algorithm is constructed according to a method comprising the following steps:

[0122] 1) Multiple ex vivo bone marrow CD138+ plasma cell samples from multiple myeloma patients were used as training sets, and the transcriptome expression data of the 97 genes in each sample in the training set was used as training set data. The training sets were divided into two groups using an unsupervised clustering method;

[0123] 2) using the grouping labels of each sample in each group obtained in 1) and the transcriptome expression data of the 97 genes of the corresponding sample as input data, training a support vector machine model to generate a myeloma classification score for each sample in the training set, recorded as an MCS value;

[0124] 3) Using the transcriptome expression data of 97 genes of all samples in the training set as the explanatory variable and the MCS values ​​of all samples in the training set as the response variable, a single-sample prediction algorithm based on the multivariate linear regression model was constructed;

[0125] S3. Subtract the MCS value of the m-th test from the MCS value of the same multiple myeloma patient subject at the n-th test as ΔMCS, and compare the responses of the multiple myeloma patient subjects to the combination treatment regimen based on ΔMCS.

[0126] In an eighth aspect, the present invention provides a method for constructing a single sample prediction algorithm, the method comprising the following steps:

[0127] 1) Multiple ex vivo bone marrow CD138+ plasma cell samples from multiple myeloma patients were used as training sets, and the transcriptome expression data of the 97 genes in each sample in the training set was used as training set data. The training sets were divided into two groups using an unsupervised clustering method;

[0128] 2) using the grouping labels of each sample in each group obtained in 1) and the transcriptome expression data of the 97 genes of the corresponding sample as input data, training a support vector machine model to generate a myeloma classification score for each sample in the training set, recorded as an MCS value;

[0129] 3) Using the transcriptome expression data of 97 genes of all samples in the training set as explanatory variables and the MCS values ​​of all samples in the training set as response variables, a single-sample prediction algorithm based on a multiple linear regression model was constructed.

[0130] In the above text, the training set consists of more than 500 bone marrow CD138+ plasma cell samples from multiple myeloma patients, specifically 840 samples in the MMRF database.

[0131] If the expression data in the above text is in FPKM format, it is directly used as expression data (recorded as GEPs data); if it is an SRA format file generated by the RNA-seq platform, it is processed as follows: first, use fastq-dump in the SRA toolkit to convert the downloaded SRA file into a FASTQ file; then use Fastqc software for quality control; use STAR software to align the FASTQ file with the GRCh38 reference genome through the quality-controlled sample file, and further generate a BAM format file; finally, use Stringtie software to generate FPKM format expression spectrum data as expression data (recorded as GEPs data).

[0132] In the above, the combination treatment regimen is the VRD (bortezomib + lenalidomide + dexamethasone) regimen.

[0133] In the above text, the prognosis is reflected by overall survival (OS); specifically, the prognosis of patients in the MCS stable group is better or potentially better than that in the MCS progressive group;

[0134] In the above, genomic stability is reflected by whole-genome copy number variation and / or average expression of CDC20-M gene; specifically, the genomic stability of patients in the MCS stable group is better or potentially better than that in the MCS progressive group;

[0135] The above methods are not intended for disease diagnosis or treatment. The methods may not be diagnostic methods, which refer to processes for identifying, studying, and determining the cause of disease or the state of a lesion in a living human or animal body. The methods may not have the direct purpose of obtaining a disease diagnosis or health status. The methods may not be therapeutic methods, which refer to processes for blocking, alleviating, or eliminating the cause of disease or lesion in a living human or animal body in order to restore or achieve health or alleviate suffering.

[0136] Each of the above methods may not include the step of obtaining a biological sample from an animal. Each of the above methods may not be directed to a living human or animal body, but rather to data. Each of the above methods may be an information processing method in which all steps are performed by a data processing device such as a computer.

[0137] The data processing device is a computer processing device.

[0138] Experiments in this study demonstrate that the single-sample prediction SSP algorithm constructed by this study measures the MCS values ​​of the same sample at different times to determine the change in MCS values ​​between different time points, recorded as ΔMCS. ΔMCS is closely correlated with the prognosis, genomic alterations, and treatment response outcomes of multiple myeloma patients, providing new possibilities for dynamically monitoring disease progression in MM patients. BRIEF DESCRIPTION OF THE DRAWINGS

[0139] Figure 1 Flowchart of the in silico model for disease progression prediction in MM patients.

[0140] Figure 2 is the distribution of ΔMCS in the MCS progression group and the MCS stable group.

[0141] Figure 3 This is the prognostic analysis of patients in the MCS stable group and the MCS progressive group.

[0142] Figure 4 These are the genomic alteration characteristics of the MCS progression group and the MCS stable group.

[0143] Figure 5 The average expression level of CDC20-M gene was significantly increased in the follow-up samples of MCS progression group.

[0144] Figure 6 is the correlation between ΔMCS and induction treatment outcomes. DETAILED DESCRIPTION

[0145] The present invention will be further described in detail below in conjunction with specific embodiments. The examples provided are only for illustrating the present invention and are not intended to limit the scope of the present invention. The examples provided below can serve as a guide for further improvements by those skilled in the art and are not intended to limit the present invention in any way.

[0146] Unless otherwise specified, the experimental methods in the following examples are conventional methods and were performed according to the techniques or conditions described in the literature in the field or according to the product instructions. The materials and reagents used in the following examples, unless otherwise specified, were all commercially available.

[0147] Unless otherwise specified, the quantitative tests in the following examples were performed three times, and the results were averaged.

[0148] The database information in the following examples is shown in Table 1.

[0149] Table 1 is the mRNA transcriptome database

[0150]

[0151] The above database records the transcriptome expression levels of each gene in bone marrow CD138+ plasma cells, and these expression level data are recorded as Gene Expression Profile (GEP).

[0152] Table 2 shows the genbank numbers of the 97 genes

[0153]

[0154] Example 1: Establishment of a computer device for predicting disease progression in MM patients

[0155] Figure 1 Flowchart of the in silico model for disease progression prediction in MM patients.

[0156] Identification of 97 genes for predicting disease progression in MM patients

[0157] Using the multiple myeloma expression dataset GSE2658 provided by the NCBI GEO public database, 97 stable differentially expressed and relatively high abundance classification genes were retained through Pearson correlation analysis.

[0158] The names of the 97 genes are as follows (see Table 2 for details):

[0159] ACBD3, ADAR, ADSS, ALDH2, ANP32E, ANXA2, ATF3, ATP8B2, CACYBP, CAPN2, CCND1, CCT3, CDC42SE1, CERS2, CHSY3, CLIC1, CLMN, COPA, CSNK1G3, DAP3, DENND1B, ENSA, EPRS, EPSTI1, EVL , FAM13A, FAM49A, FLAD1, FRZB, GLRX2, HAX1, HDGF, HLA-A, HLA-B, HLA-C, HLA-F, HLA-G, IL6R, ISG20L2, JTB, KLF2, LAMTOR2, LDHA, MCL1, MOXD1, MRPL24, MRPL9, MVP, MYL6, NDUFS2, NO P58, NOTCH2NL, NTAN1, PAK1, PI4KB, PIEZO1, PIK3AP1, PIM2, PIP5K1B, PMVK, POGZ, PPIA, PRCC, PRKCA, PRRC2C, PSMB4, PSMD4, RAB29, RCBTB2, SCAMP3, SCAPER, SDHC, SEL1L3, SELPLG, SHC1, SIDT1, SSR2, STAP1, TAP1, TIMM17A, TLR10, TMCO1, TOR1AIP2, TOR3A, TP53INP1, TPM3, TRANK1, TROVE2, UAP1, UBE2Q1, UBQLN4, UHMK1, VPS45, YY1AP1, ZC3H11A, ZFP36, and ZNF593.

[0160] These 97 genes will then serve as classifier genes for disease progression prediction or clustering of MM patients.

[0161] 2. Establishment of Single Sample Prediction (SSP) Algorithm

[0162] Using the transcriptome data of 840 samples in the MMRF database as the training set, a single sample prediction (SSP) algorithm for the RNAseq platform was constructed:

[0163] 1. Based on the unsupervised clustering method, the dataset is divided into two groups and the MCL1-M of all samples is generated. high or TLR10-M high Grouping Labels

[0164] Firstly, the gene expression profile (GEP) data of 97 classifier genes of 840 samples in the MMRF database were log2 transformed and z-score normalized. Then, the normalized data were clustered using an unsupervised method to obtain the MCL1-M at the dataset level. high or TLR10-M high Group labels.

[0165] The above unsupervised clustering method is as follows: the ConsensusClusterPlus R package was used to perform cluster analysis using the pam clustering algorithm. When K = 2, the most stable clustering was achieved. The best grouping effect was determined based on the consistency matrix (K cluster blocks with clear boundaries). The clustering results divided the samples in the data set into two groups, named MCL1-M high group or TLR10-M high Group, that is, add MCL1-M high or TLR10-M high Group labels.

[0166] 2. Train the SVM model to generate a dataset-level myeloma classification score (MCS) for all samples

[0167] The grouping labels of each sample in the two groups obtained in 1 above and the GEPs data of the 97 classifier genes corresponding to each sample were used as input data to train the Support Vector Machine (SVM) model and generate the MCS of all samples in the training set.

[0168] During training, a grid search method was used to determine the optimal parameter C = 1. The model's robustness was tested within the training set using 10-fold cross-validation. The validation results demonstrated that recall, accuracy, and precision reached 100%. Under these conditions, the geometric margins of the training set samples (the distance between each sample and the separating hyperplane) were output.

[0169] The geometric interval is a continuous variable with a positive or negative sign and is defined as MCS (myeloma classification score).

[0170] MCL1-M high The MCS of the subtype is negative, TLR10-M high The MCS of the subtype is positive.

[0171] 3. Develop an SSP algorithm based on Multiple Linear Regression (MLR) to calculate the MCS of a single sample

[0172] The GEPs data of 97 classifier genes of all samples in the training set were used as explanatory variables, and the MCS of all samples in the training set were used as response variables to construct a single sample prediction (SSP) algorithm based on the MLR model.

[0173] The expression spectrum data in FPKM format is directly obtained from the MMRF data as GEPs data.

[0174] The GEPs data of 97 classifier genes in the in vitro bone marrow CD138+ plasma cell sample of a single MM patient to be predicted were used as the input data of the above-mentioned single sample prediction (SSP) algorithm, and the MCS of the single sample was output.

[0175] If the GEPs data of the 97 classifier genes to be predicted for the ex vivo bone marrow CD138+ plasma cell sample of a single MM patient is in the SRA format generated by the RNA-seq platform, it needs to be processed as follows: first, use fastq-dump in the SRA toolkit to convert the downloaded SRA file into a FASTQ file; then use Fastqc software for quality control; use STAR software to align the FASTQ file with the GRCh38 reference genome through the quality-controlled sample file to further generate a BAM format file; finally, use Stringtie software to generate expression profile data in FPKM format as GEPs data.

[0176] ΔMCS predicts disease progression in MM patients

[0177] 1. Definition of ΔMCS

[0178] MCS is closely related to the differential expression of 97 classifier genes. Genomic changes within cells and the evolution of dominant cell clones constantly influence the transcriptome state at the patient bulk (tissue or cell population) level. Therefore, we sought to investigate whether changes in MCS values ​​across samples from the same MM patient at different time points correlate with disease progression.

[0179] ΔMCS is used to measure the MCS difference between samples tested at different times, that is, ΔMCS = MCS value of the later (nth) test sample - MCS value of the previous (mth) test sample; n>m, m is a natural number greater than or equal to 1, and n is a natural number.

[0180] Since ΔMCS is an individualized indicator, it can be used to measure the change in MCS between any two samples from the same patient.

[0181] 2. Correlation between ΔMCS and patient prognosis and genomic alterations

[0182] The correlation between ΔMCS and disease progression was analyzed using longitudinal samples (ex vivo bone marrow CD138+ plasma cell samples) collected at different time points from 45 patients in the MMRF dataset. To ensure comparability of the results, the ΔMCS for each patient in this section was calculated by subtracting the MCS value of the baseline follow-up (first visit) sample from the MCS value of the final follow-up sample.

[0183] 1) Use the single sample prediction (SSP) algorithm to obtain the two MCS of the sample

[0184] The GEPs data of 97 classifier genes were collected from 45 patients in the MMRF dataset at two different time points: the baseline period (the mth time) and the last follow-up period (the nth time). The 97 classifier gene GEPs data of each sample were used as input data for the single sample prediction (SSP) algorithm constructed above, and the baseline follow-up MCS values ​​and the last follow-up MCS values ​​of the individual samples of the 45 patients were output.

[0185] 2) ΔMCS

[0186] ΔMCS of each patient = MCS value at the last follow-up - MCS value at the baseline follow-up

[0187] The results are shown in Tables 3 and 4.

[0188] Table 3 shows the characteristics of patients with ΔMCS ≥ -2 in the follow-up samples in the MMRF dataset.

[0189]

[0190]

[0191] Note: ΔMCS analysis was performed on paired longitudinal samples. The table shows the characteristics of patients with a ΔMCS ≥ -2. ASCT: autologous stem cell transplantation; NA: not available; Status: 0 = "alive," 1 = "dead"; Baseline sample represents the time of first visit, and Follow-up Samples 1, 2, 3, and 4 represent the time periods from the beginning to the end of follow-up. The MCS value of each patient's last follow-up sample was recorded as the final follow-up MCS value.

[0192] Table 4 shows the characteristics of patients with ΔMCS < -2 in the follow-up samples in the MMRF dataset

[0193]

[0194]

[0195] Note: The table shows the characteristics of patients with a ΔMCS < -2. ASCT: autologous stem cell transplantation; NA: not provided; status: 0 = "alive," 1 = "deceased." The baseline sample represents the initial visit, while follow-up samples 1, 2, 3, and 4 represent the follow-up period from the earliest to the latest. The MCS value of each patient's last follow-up sample was recorded as the final follow-up MCS value.

[0196] 3) According to the ΔMCS values, the 45 patients were divided into MCS progression group and MCS stable group

[0197] Patients with ΔMCS < -2 were defined as the MCS progression group, and patients with ΔMCS ≥ -2 were defined as the MCS stable group.

[0198] The disease progression in the MCS stable group was slower than that in the MCS progressive group.

[0199] The above 45 patients were divided into MCS progression group and MCS stable group according to ΔMCS value. Figure 2 As shown in the figure, it can be seen that the 45 samples were divided into 22 patients in the MCS stable group and 23 patients in the MCS progressive group with a cutoff value of -2.

[0200] MM patients in the MCS progression group were in a progressive state (the genomic instability of the nth test sample increased and / or the prognosis became worse compared with the mth test sample);

[0201] The MM patients in the MCS stable group were in a stable condition (the genomic status and / or prognosis of the nth test sample did not change significantly compared with the mth test sample).

[0202] 3. Verification

[0203] Based on the defined 23 MCS progression groups and 22 MCS stable groups, a comprehensive multifaceted analysis was conducted to analyze the characteristics of the MCS stable and MCS progression groups in terms of prognosis and genomic changes.

[0204] 1) Prognosis

[0205] According to the follow-up prognostic data, the prognostic analysis results are as follows Figure 3 As shown, it can be seen that compared with patients in the MCS progression group, patients in the MCS stable group had a longer OS, among which the MCS stable group had not reached the median OS, and the median OS of the MCS progression group was 2.42 years (HR: 0.36, 95% CI: 0.15-0.87, p=0.029), indicating that the prognosis of patients in the MCS progression group was worse than that in the MCS stable group.

[0206] This prognostic trend was not affected by the age of the two groups of patients, whether they received ASCT treatment, and the risk-related chromosomal translocation status.

[0207] 2) Copy number analysis based on whole genome sequencing data

[0208] (1) Detection of genomic copy number variation

[0209] Whole-genome sequencing data for the test samples are available in the MMRF database. Paired longitudinal sample bam files were downloaded from dbGaP (phs000748.v7.p4) (https: / / portal.gdc.cancer.gov). The bam files were processed using the QDNAseq R package to generate seg files. Copy number variation was calculated using a threshold of ±0.4. A log2 ratio greater than 0.4 was used as the threshold for chromosome amplification (one additional copy) or gain (two or more additional copies), and a log2 ratio less than −0.4 was used as the threshold for chromosome loss. The copy number variation ratio was calculated as the total length of the copy number variation divided by the total length of the tested genome.

[0210] For example, a patient in the MCS progression group and a patient in the MCS stable group were sequenced using whole genome sequencing to generate seg files for plotting. Figure 4Figures A and 4B show that A is a typical MCS stable group MM patient with a ΔMCS of 1.8 but carrying a high-risk t(4;14) chromosomal translocation. The figure shows the genomic copy number variation of paired baseline and last follow-up samples. In the follow-up samples, the patient did not show any significant new copy number change events. The figure above indicates the patient's age, whether or not he received ASCT treatment, common chromosomal translocation events, ISS stage, and OS (status). Figure B is a typical MCS progressive group MM patient but carrying a high-risk t(11;14) chromosomal translocation. The patient's ΔMCS was -4.2. The figure shows the genomic copy number variation of paired baseline and last follow-up samples. In the follow-up samples, the patient showed significant new copy number change events and even chromosome breakage events. The figure above indicates the patient's age, whether or not he received ASCT treatment, common chromosomal translocation events, ISS stage, and OS (status). The red short line in the last follow-up sample marks the subclonal gain and loss of the sample.

[0211] The copy number variation results of all patient samples with available data in the MCS stable group are as follows Figure 4 As shown in C, there was no significant increase in copy number variation in the follow-up samples of the MCS stable group, and the paired sample t test was not significant. The figure shows data from 17 pairs of longitudinal samples.

[0212] The copy number variation results of all patient samples with available data in the MCS progression group are as follows Figure 4 D shows a significant increase in copy number variation in the follow-up samples of the MCS progression group, *: p < 0.05, paired sample t-test. Data from 15 paired longitudinal samples are shown.

[0213] In summary, the genomic stability of patients in the MCS progression group was worse than that in the MCS stable group.

[0214] Next, the correlation between the MCS stable group and the MCS progressive group and MM patients' genomic changes was analyzed. In the MCS stable group, the genomic status of the patients' follow-up samples was basically similar to that of the paired baseline samples. This genomic stability is not affected by high-risk genetic events. Many studies have shown that 1q+ and t(4;14) are classic high-risk prognostic markers. Among the 17 MCS stable group cases analyzed, 9 patients had high-fold amplification of 1q. Among them, 5 patients were detected with t(4;14) translocation events ( Figure 4A and 4C). In contrast, in the MCS progression group, the copy number variation ratio in the follow-up samples increased significantly compared with the baseline samples. Most of the newly added genomic variation events were subsequent subclonal events. In the longitudinal analysis results of 15 patients, 5 MM patients had no 1q+ detected in both baseline and follow-up samples. In the other 4 MM patients, although 1q+ was present in the baseline samples, no further increase in 1q copy number was found in the follow-up samples. The t(11;14) translocation event is a recognized standard-risk prognostic marker in the field, and the occurrence of this chromosomal translocation event was detected in 6 patients in the MCS progression group ( Figure 4 B and 4D).

[0215] In the MCS stable group, 1q+ and t(4;14) were found to be frequently present, while in the MCS progressive group, 1q+ deletion, stability, and t(11;14) events were frequently present. These results, which seem to contradict clinical common sense, confirm that a single prognostic marker is difficult to assess the stage-by-stage disease progression of a single MM patient. Although classic prognostic markers such as 1q+, t(4;14), and t(11;14) are significantly correlated with patient prognosis at the population level, the detection rate of these classic prognostic markers in the MM population is generally not high, so they cannot be routinely used to monitor the disease progression status of MM patients.

[0216] The generation of ΔMCS relies solely on the expression of 97 genes closely related to plasma cell development and MM pathogenesis, and the expression of these genes can always be detected in patients at any stage. Therefore, the dynamic predictive function based on ΔMCS is not provided by traditional prognostic markers.

[0217] (2) Detection of CDC20-M gene expression

[0218] CDC20-M is a group of genes closely related to cell malignant proliferation and genomic instability. The 139 genes that make up CDC20-M are as follows: ASF1B (55723), ASPM (259266), AURKA (6790), AURKB (9212), BIRC5 (332), BRCA1 (672), BUB1 (699), BUB1B (701), BUD31 (8896), C11orf82 (220042), CASC5 (57082), CASP2 (835), CCNA2 (890), CCNB1 (891), CCNB2 (9133), CDC20 (991), and CDC2 5A(993), CDC45(8318), CDC6(990), CDCA2(157313), CDCA3(83461), CDCA7(83879), CDCA7L(55536), CDCA8(55143), CDK1(983), CDK2(1017), CDKN2C (1031), CDKN3(1033), CENPA(1058), CENPE(1062), CENPF(1063), CENPK(64105), CENPN(55839), CENPW(387103), CEP55(55165), CHEK1(1111), CKAP2 L(150468), CKS2(1164), DBF4(10926), DDX39A(10212), DEPDC1(55635), DEPDC1B(55789), DLGAP5(9787), DNMT1(1786), DSN1(79980), DTL(51514), DTYMK(1841), E2F7(144455), ECT2(1894), EME1(146956), ESPL1(9700), FAM64A(54478), FANCD2(2177), FANCI(55215), FBXO5(26271), FOXM1(2305 ), GAS2L3(283431), GINS1(9837), GINS2(51659), GJC1(10052), GPX7(2882), GTF2IRD2(84163), GTSE1(51512), HJURP(55355), HMGB2(3148), HMMR( 3161), IGF2BP3(10643), KIAA0101(9768), KIF11(3832), KIF14(9928), KIF15(56992), KIF20A(10112), KIF23(9493), KIF2C(11004), KIF4A(24137),KIFC1(3833), KNTC1(9735), KPNA2(3838), LMNB1(4001), LMNB2(84823), LRR1(122769), MAD2L1(4085), MCM2(4171), MCM3(4172), MCM6(4175), MCM8(84515), MELK(9833), MKI67(4288), MLF1IP(79682), MND1(84057), MYBL2(4605), NCAP G(64151), NCAPG2(54892), NCAPH(23397), NDC80(10403), NEK2(4751), NRM(11270), NUF2(83540), NUSAP1(51203), PA RPBP(55010), PBK(55872), PCNA(5111), PDIA4(9601), POC1A(25886), POLE2(5427), PRC1(9055), PTBP1(5725), PTTG1 (9232), RACGAP1(29127), RAE1(8480), RBBP8(5932), RFC2(5982), RFC3(5983), RFC4(5984), RNASEH2A(10535), RRM2 (6241), SGOL2(151246), SHCBP1(79801), SMC4(10051), SNRPB(6628), SPAG5(10615), SPC24(147841), STIL(6491), TA CC3 (10460), TCF3 (6929), TIMELESS (8914), TK1 (7083), TMEM48 (55706), TOP2A (7153), TPX2 (22974), TRIP13 (9319), TTK (7272), TYMS (7298), UBE2C (11065), UBE2S (27338), USP1 (7398), ZNF765 (91661), ZNF850 (342892), ZWINT (11130). The numbers in parentheses after the above genes are the gene IDs of the corresponding genes on the NCBI website.

[0219] The correlation between ΔMCS and genomic instability was analyzed using the CDC20-M gene profile.

[0220] The expression of CDC20-M in 45 samples was analyzed and the results were as follows: Figure 5 As shown, Figure 5Comparison of the mean CDC20-M expression levels between paired baseline and final follow-up samples in the MCS stable group (A) or MCS progressive group (B) is shown. *: P < 0.05, **: P < 0.01, paired sample t-test; CDC20-M expression levels in the final follow-up sample in the MCS progressive group were significantly elevated compared to baseline.

[0221] Therefore, the genomic stability of patients in the MCS progression group was worse than that in the MCS stable group, and patients in the MCS progression group had stronger genomic instability.

[0222] In summary, ΔMCS is significantly correlated with the prognosis and genomic alterations of MM patients. Lower ΔMCS or being in the MCS progression group indicates a poorer prognosis and increased genomic instability in patients.

[0223] ΔMCS predicts treatment efficacy in MM patients

[0224] PETHEMA / GEM2012MENOS65 (September 2013-November 2016): An open-label phase III clinical trial, data stored in the GSE147165 database. Patients first received 6 cycles of VRD induction therapy, followed by ASCT with busulfan-melphalan or melphalan pretreatment, and finally received 2 cycles of VRD consolidation therapy. The patients then entered

[0225] The PETHEMA / GEM2014MAIN clinical trial, which included 40 patients receiving two years of maintenance therapy with either Rd or Irradiated Receptor Drug (IRD), provided paired transcriptome data before and after induction therapy with VRD (bortezomib, lenalidomide, and dexamethasone). Therefore, this study investigated whether ΔMCS could be used to predict patient response to this combination therapy.

[0226] Bone marrow CD138+ plasma cell samples were collected from patients at their first diagnosis (NDMM) before induction therapy, and bone marrow CD138+ plasma cell samples were collected from patients at the time of measurable minimal residual disease (MRD) after induction therapy.

[0227] The GEPs data of 97 classifier genes at each sample at each time were obtained as the input data of the single sample prediction (SSP) algorithm constructed above, and the MCS of a single sample at each time was output: the MCS value at the time of initial diagnosis and the MCS value when MRD could be measured, recorded as MCS (MRD) and MCS (NDMM).

[0228] ΔMCS=MCS(MRD)-MCS(NDMM).

[0229] According to the ΔMCS value, a scatter plot was made using GraphPad Prism software.

[0230] The results are as follows Figure 6 As shown, CR: clinically defined complete response; VGPR: clinically defined very good partial response; PR: clinically defined partial response; SD: clinically defined stable disease; PD: clinically defined progressive disease. T-tests were used to analyze differences in ΔMCS between groups. Cases with 1q+ chromosomal abnormalities are indicated by solid circles. *: P < 0.05, t-test;

[0231] As can be seen, the ΔMCS of CR patients was significantly higher than that of PR or PD patients. Therefore, the smaller the ΔMCS value, the worse the patient's response to the combination therapy.

[0232] In summary, changes in MCS values ​​across samples tested at different times are closely associated with patient prognosis, genomic alterations, and treatment response. Lower ΔMCS values ​​indicate greater genomic instability and a poorer response to combination therapy.

[0233] The present invention has been described in detail above. It will be apparent to those skilled in the art that the present invention may be practiced over a wide range of parameters, concentrations, and conditions without departing from the spirit and scope of the present invention and without unnecessary experimentation. Although specific embodiments have been given herein, it should be understood that further modifications may be made to the present invention. In summary, this application is intended to encompass any variations, uses, or improvements to the present invention, including those made by conventional techniques known in the art that depart from the scope of the present invention. Applications of the essential features may be made within the scope of the following claims.

Claims

1. A data processing device for clustering disease progression in multiple myeloma patients, comprising a memory, a processor, and a computer program stored in the memory, wherein: The processor executes the computer program to implement the following steps: S1. Receive data: Receive the mth sample data and the nth sample data of the same multiple myeloma patient, where n is greater than m; The m sample data are ACBD3, ADAR, ADSS, ALDH2, ANP32E, ANXA2, ATF3, ATP8B2, CACYBP, CAPN2, CCND1, CCT3, CDC42SE1, CERS2, CHSY3, CLIC1, CLMN, COPA, CSNK1G3, DAP3, DENND1B, ENSA, EPRS, EPSTI1, EVL, FAM13A, FAM49A, FLAD1, FRZB, GLRX2, HAX1, HDGF, HLA-A, HLA-B, HLA-C, HLA-F, HLA-G, IL6R, ISG20L2, JTB, KLF2, LAMTOR2, LDHA, MCL1, MOXD1, MRPL24, MRPL9, MVP, Transcriptome expression data of 97 genes: MYL6, NDUFS2, NOP58, NOTCH2NL, NTAN1, PAK1, PI4KB, PIEZO1, PIK3AP1, PIM2, PIP5K1B, PMVK, POGZ, PPIA, PRCC, PRKCA, PRRC2C, PSMB4, PSMD4, RAB29, RCBTB2, SCAMP3, SCAPER, SDHC, SEL1L3, SELPLG, SHC1, SIDT1, SSR2, STAP1, TAP1, TIMM17A, TLR10, TMCO1, TOR1AIP2, TOR3A, TP53INP1, TPM3, TRANK1, TROVE2, UAP1, UBE2Q1, UBQLN4, UHMK1, VPS45, YY1AP1, ZC3H11A, ZFP36, and ZNF593; The nth sample data is the transcriptome expression data of the 97 genes in the ex vivo bone marrow CD138+ plasma cell sample of the multiple myeloma patient during the nth test; S2. Data processing: Inputting the m-th sample data and the n-th sample data of the same multiple myeloma patient subject into the single sample prediction algorithm respectively, and outputting the myeloma classification score of the m-th detection and the myeloma classification score of the multiple myeloma patient subject, which are recorded as the MCS value of the m-th detection and the MCS value of the n-th detection respectively; The single sample prediction algorithm is constructed according to a method comprising the following steps: 1) Multiple ex vivo bone marrow CD138+ plasma cell samples from multiple myeloma patients were used as training sets, and the transcriptome expression data of the 97 genes in each sample in the training set was used as training set data. The training sets were divided into two groups using an unsupervised clustering method; 2) using the grouping labels of each sample in each group obtained in 1) and the transcriptome expression data of the 97 genes of the corresponding sample as input data, training a support vector machine model to generate a myeloma classification score for each sample in the training set, recorded as an MCS value; 3) Using the transcriptome expression data of 97 genes of all samples in the training set as the explanatory variable and the MCS values ​​of all samples in the training set as the response variable, a single-sample prediction algorithm based on the multivariate linear regression model was constructed; S3. Output results: The value obtained by subtracting the MCS value of the mth test from the MCS value of the same multiple myeloma patient subject at the nth test is recorded as ΔMCS, and the disease progression of the multiple myeloma patient subject is grouped according to ΔMCS.

2. A data processing apparatus for comparing disease progression in multiple myeloma patients, comprising a memory, a processor, and a computer program stored in the memory, wherein: The processor executes the computer program to implement the following steps: S1. Receive data: Receive the mth sample data and the nth sample data of the same multiple myeloma patient, where n is greater than m; The m sample data are ACBD3, ADAR, ADSS, ALDH2, ANP32E, ANXA2, ATF3, ATP8B2, CACYBP, CAPN2, CCND1, CCT3, CDC42SE1, CERS2, CHSY3, CLIC1, CLMN, COPA, CSNK1G3, DAP3, DENND1B, ENSA, EPRS, EPSTI1, EVL, FAM13A, FAM49A, FLAD1, FRZB, GLRX2, HAX1, HDGF, HLA-A, HLA-B, HLA-C, HLA-F, HLA-G, IL6R, ISG20L2, JTB, KLF2, LAMTOR2, LDHA, MCL1, MOXD1, MRPL24, MRPL9, MVP, Transcriptome expression data of 97 genes: MYL6, NDUFS2, NOP58, NOTCH2NL, NTAN1, PAK1, PI4KB, PIEZO1, PIK3AP1, PIM2, PIP5K1B, PMVK, POGZ, PPIA, PRCC, PRKCA, PRRC2C, PSMB4, PSMD4, RAB29, RCBTB2, SCAMP3, SCAPER, SDHC, SEL1L3, SELPLG, SHC1, SIDT1, SSR2, STAP1, TAP1, TIMM17A, TLR10, TMCO1, TOR1AIP2, TOR3A, TP53INP1, TPM3, TRANK1, TROVE2, UAP1, UBE2Q1, UBQLN4, UHMK1, VPS45, YY1AP1, ZC3H11A, ZFP36, and ZNF593; The nth sample data is the transcriptome expression data of the 97 genes in the ex vivo bone marrow CD138+ plasma cell sample of the multiple myeloma patient during the nth test; S2. Data processing: Inputting the m-th sample data and the n-th sample data of the same multiple myeloma patient subject into the single sample prediction algorithm respectively, and outputting the myeloma classification score of the m-th detection and the myeloma classification score of the multiple myeloma patient subject, which are recorded as the MCS value of the m-th detection and the MCS value of the n-th detection respectively; The single sample prediction algorithm is constructed according to a method comprising the following steps: 1) Multiple ex vivo bone marrow CD138+ plasma cell samples from multiple myeloma patients were used as training sets, and the transcriptome expression data of the 97 genes in each sample in the training set was used as training set data. The training sets were divided into two groups using an unsupervised clustering method; 2) using the grouping labels of each sample in each group obtained in 1) and the transcriptome expression data of the 97 genes of the corresponding sample as input data, training a support vector machine model to generate a myeloma classification score for each sample in the training set, recorded as an MCS value; 3) Using the transcriptome expression data of 97 genes of all samples in the training set as the explanatory variable and the MCS values ​​of all samples in the training set as the response variable, a single-sample prediction algorithm based on the multivariate linear regression model was constructed; S3. Output results: The value obtained by subtracting the MCS value of the mth test from the MCS value of the same multiple myeloma patient subject at the nth test is recorded as ΔMCS, and the disease progression of the multiple myeloma patient subjects is compared based on ΔMCS.

3. A data processing apparatus for comparing responses of multiple myeloma patients to combination therapy regimens, comprising a memory, a processor, and a computer program stored in the memory, wherein: The processor executes the computer program to implement the following steps: S1. Receive data: Receive the mth sample data and the nth sample data of the same multiple myeloma patient, where n is greater than m; The m sample data are ACBD3, ADAR, ADSS, ALDH2, ANP32E, ANXA2, ATF3, ATP8B2, CACYBP, CAPN2, CCND1, CCT3, CDC42SE1, CERS2, CHSY3, CLIC1, CLMN, COPA, CSNK1G3, DAP3, DENND1B, ENSA, EPRS, EPSTI1, EVL, FAM13A, FAM49A, FLAD1, FRZB, GLRX2, HAX1, HDGF, HLA-A, HLA-B, HLA-C, HLA-F, HLA-G, IL6R, ISG20L2, JTB, KLF2, LAMTOR2, LDHA, MCL1, MOXD1, MRPL24, MRPL9, MVP, Transcriptome expression data of 97 genes: MYL6, NDUFS2, NOP58, NOTCH2NL, NTAN1, PAK1, PI4KB, PIEZO1, PIK3AP1, PIM2, PIP5K1B, PMVK, POGZ, PPIA, PRCC, PRKCA, PRRC2C, PSMB4, PSMD4, RAB29, RCBTB2, SCAMP3, SCAPER, SDHC, SEL1L3, SELPLG, SHC1, SIDT1, SSR2, STAP1, TAP1, TIMM17A, TLR10, TMCO1, TOR1AIP2, TOR3A, TP53INP1, TPM3, TRANK1, TROVE2, UAP1, UBE2Q1, UBQLN4, UHMK1, VPS45, YY1AP1, ZC3H11A, ZFP36, and ZNF593; The nth sample data is the transcriptome expression data of the 97 genes in the ex vivo bone marrow CD138+ plasma cell sample of the multiple myeloma patient during the nth test; S2. Data processing: Inputting the m-th sample data and the n-th sample data of the same multiple myeloma patient subject into the single sample prediction algorithm respectively, and outputting the myeloma classification score of the m-th detection and the myeloma classification score of the multiple myeloma patient subject, which are recorded as the MCS value of the m-th detection and the MCS value of the n-th detection respectively; The single sample prediction algorithm is constructed according to a method comprising the following steps: 1) Multiple ex vivo bone marrow CD138+ plasma cell samples from multiple myeloma patients were used as training sets, and the transcriptome expression data of the 97 genes in each sample in the training set was used as training set data. The training sets were divided into two groups using an unsupervised clustering method; 2) using the grouping labels of each sample in each group obtained in 1) and the transcriptome expression data of the 97 genes of the corresponding sample as input data, training a support vector machine model to generate a myeloma classification score for each sample in the training set, recorded as an MCS value; 3) Using the transcriptome expression data of 97 genes of all samples in the training set as the explanatory variable and the MCS values ​​of all samples in the training set as the response variable, a single-sample prediction algorithm based on the multivariate linear regression model was constructed; S3. Output results: The value obtained by subtracting the MCS value of the mth test from the MCS value of the same multiple myeloma patient subject at the nth test is recorded as ΔMCS, and the response of the multiple myeloma patient subject to the combination treatment regimen is compared based on ΔMCS.

4. Computer program product, including a computer program, characterized in that: When the computer program is executed by a processor, the steps in the data processing device according to any one of claims 1 to 3 are implemented.

5. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program causes a computer to execute the steps in the data processing device according to any one of claims 1 to 3.

6. Any of the following devices: A1. A device for grouping disease progression in multiple myeloma patients, comprising: S1, data receiving module: used to receive the mth sample data and the nth sample data of the same multiple myeloma patient subject, where n is greater than m; The m sample data are ACBD3, ADAR, ADSS, ALDH2, ANP32E, ANXA2, ATF3, ATP8B2, CACYBP, CAPN2, CCND1, CCT3, CDC42SE1, CERS2, CHSY3, CLIC1, CLMN, COPA, CSNK1G3, DAP3, DENND1B, ENSA, EPRS, EPSTI1, EVL, FAM13A, FAM49A, FLAD1, FRZB, GLRX2, HAX1, HDGF, HLA-A, HLA-B, HLA-C, HLA-F, HLA-G, IL6R, ISG20L2, JTB, KLF2, LAMTOR2, LDHA, MCL1, MOXD1, MRPL24, MRPL9, MVP, Transcriptome expression data of 97 genes: MYL6, NDUFS2, NOP58, NOTCH2NL, NTAN1, PAK1, PI4KB, PIEZO1, PIK3AP1, PIM2, PIP5K1B, PMVK, POGZ, PPIA, PRCC, PRKCA, PRRC2C, PSMB4, PSMD4, RAB29, RCBTB2, SCAMP3, SCAPER, SDHC, SEL1L3, SELPLG, SHC1, SIDT1, SSR2, STAP1, TAP1, TIMM17A, TLR10, TMCO1, TOR1AIP2, TOR3A, TP53INP1, TPM3, TRANK1, TROVE2, UAP1, UBE2Q1, UBQLN4, UHMK1, VPS45, YY1AP1, ZC3H11A, ZFP36, and ZNF593; The nth sample data is the transcriptome expression data of the 97 genes in the ex vivo bone marrow CD138+ plasma cell sample of the multiple myeloma patient during the nth test; S2, a data processing module: used to input the m-th sample data and the n-th sample data of the same multiple myeloma patient subject into the single sample prediction algorithm, and output the myeloma classification score of the m-th detection and the myeloma classification score of the multiple myeloma patient subject, which are recorded as the MCS value of the m-th detection and the MCS value of the n-th detection respectively; The single sample prediction algorithm is constructed according to a method comprising the following steps: 1) Multiple ex vivo bone marrow CD138+ plasma cell samples from multiple myeloma patients were used as training sets, and the transcriptome expression data of the 97 genes in each sample in the training set was used as training set data. The training sets were divided into two groups using an unsupervised clustering method; 2) using the grouping labels of each sample in each group obtained in 1) and the transcriptome expression data of the 97 genes of the corresponding sample as input data, training a support vector machine model to generate a myeloma classification score for each sample in the training set, recorded as an MCS value; 3) Using the transcriptome expression data of 97 genes of all samples in the training set as the explanatory variable and the MCS values ​​of all samples in the training set as the response variable, a single-sample prediction algorithm based on the multivariate linear regression model was constructed; S3. Output results: The value obtained by subtracting the MCS value of the m-th test from the MCS value of the same multiple myeloma patient subject at the n-th test is recorded as ΔMCS, and the disease progression of the multiple myeloma patient subject is grouped according to ΔMCS; A2. A device for comparing disease progression in multiple myeloma patient subjects, the device comprising: S1, data receiving module: used to receive the mth sample data and the nth sample data of the same multiple myeloma patient subject, where n is greater than m; The m sample data are ACBD3, ADAR, ADSS, ALDH2, ANP32E, ANXA2, ATF3, ATP8B2, CACYBP, CAPN2, CCND1, CCT3, CDC42SE1, CERS2, CHSY3, CLIC1, CLMN, COPA, CSNK1G3, DAP3, DENND1B, ENSA, EPRS, EPSTI1, EVL, FAM13A, FAM49A, FLAD1, FRZB, GLRX2, HAX1, HDGF, HLA-A, HLA-B, HLA-C, HLA-F, HLA-G, IL6R, ISG20L2, JTB, KLF2, LAMTOR2, LDHA, MCL1, MOXD1, MRPL24, MRPL9, MVP, Transcriptome expression data of 97 genes: MYL6, NDUFS2, NOP58, NOTCH2NL, NTAN1, PAK1, PI4KB, PIEZO1, PIK3AP1, PIM2, PIP5K1B, PMVK, POGZ, PPIA, PRCC, PRKCA, PRRC2C, PSMB4, PSMD4, RAB29, RCBTB2, SCAMP3, SCAPER, SDHC, SEL1L3, SELPLG, SHC1, SIDT1, SSR2, STAP1, TAP1, TIMM17A, TLR10, TMCO1, TOR1AIP2, TOR3A, TP53INP1, TPM3, TRANK1, TROVE2, UAP1, UBE2Q1, UBQLN4, UHMK1, VPS45, YY1AP1, ZC3H11A, ZFP36, and ZNF593; The nth sample data is the transcriptome expression data of the 97 genes in the ex vivo bone marrow CD138+ plasma cell sample of the multiple myeloma patient during the nth test; S2, a data processing module: used to input the m-th sample data and the n-th sample data of the same multiple myeloma patient subject into the single sample prediction algorithm, and output the myeloma classification score of the m-th detection and the myeloma classification score of the multiple myeloma patient subject, which are recorded as the MCS value of the m-th detection and the MCS value of the n-th detection respectively; The single sample prediction algorithm is constructed according to a method comprising the following steps: 1) Multiple ex vivo bone marrow CD138+ plasma cell samples from multiple myeloma patients were used as training sets, and the transcriptome expression data of the 97 genes in each sample in the training set was used as training set data. The training sets were divided into two groups using an unsupervised clustering method; 2) using the grouping labels of each sample in each group obtained in 1) and the transcriptome expression data of the 97 genes of the corresponding sample as input data, training a support vector machine model to generate a myeloma classification score for each sample in the training set, recorded as an MCS value; 3) Using the transcriptome expression data of 97 genes of all samples in the training set as the explanatory variable and the MCS values ​​of all samples in the training set as the response variable, a single-sample prediction algorithm based on the multivariate linear regression model was constructed; S3. Output results: The value obtained by subtracting the MCS value of the m-th test from the MCS value of the same multiple myeloma patient subject at the n-th test is recorded as ΔMCS, and the disease progression of the multiple myeloma patient subject is compared based on ΔMCS; A3. A device for comparing responses of multiple myeloma patient subjects to a combination therapy regimen, the device comprising: S1, data receiving module: used to receive the mth sample data and the nth sample data of the same multiple myeloma patient subject, where n is greater than m; The m sample data are ACBD3, ADAR, ADSS, ALDH2, ANP32E, ANXA2, ATF3, ATP8B2, CACYBP, CAPN2, CCND1, CCT3, CDC42SE1, CERS2, CHSY3, CLIC1, CLMN, COPA, CSNK1G3, DAP3, DENND1B, ENSA, EPRS, EPSTI1, EVL, FAM13A, FAM49A, FLAD1, FRZB, GLRX2, HAX1, HDGF, HLA-A, HLA-B, HLA-C, HLA-F, HLA-G, IL6R, ISG20L2, JTB, KLF2, LAMTOR2, LDHA, MCL1, MOXD1, MRPL24, MRPL9, MVP, Transcriptome expression data of 97 genes: MYL6, NDUFS2, NOP58, NOTCH2NL, NTAN1, PAK1, PI4KB, PIEZO1, PIK3AP1, PIM2, PIP5K1B, PMVK, POGZ, PPIA, PRCC, PRKCA, PRRC2C, PSMB4, PSMD4, RAB29, RCBTB2, SCAMP3, SCAPER, SDHC, SEL1L3, SELPLG, SHC1, SIDT1, SSR2, STAP1, TAP1, TIMM17A, TLR10, TMCO1, TOR1AIP2, TOR3A, TP53INP1, TPM3, TRANK1, TROVE2, UAP1, UBE2Q1, UBQLN4, UHMK1, VPS45, YY1AP1, ZC3H11A, ZFP36, and ZNF593; The nth sample data is the transcriptome expression data of the 97 genes in the ex vivo bone marrow CD138+ plasma cell sample of the multiple myeloma patient during the nth test; S2, a data processing module: used to input the m-th sample data and the n-th sample data of the same multiple myeloma patient subject into the single sample prediction algorithm, and output the myeloma classification score of the m-th detection and the myeloma classification score of the multiple myeloma patient subject, which are recorded as the MCS value of the m-th detection and the MCS value of the n-th detection respectively; The single sample prediction algorithm is constructed according to a method comprising the following steps: 1) Multiple ex vivo bone marrow CD138+ plasma cell samples from multiple myeloma patients were used as training sets, and the transcriptome expression data of the 97 genes in each sample in the training set was used as training set data. The training sets were divided into two groups using an unsupervised clustering method; 2) using the grouping labels of each sample in each group obtained in 1) and the transcriptome expression data of the 97 genes of the corresponding sample as input data, training a support vector machine model to generate a myeloma classification score for each sample in the training set, recorded as an MCS value; 3) Using the transcriptome expression data of 97 genes of all samples in the training set as the explanatory variable and the MCS values ​​of all samples in the training set as the response variable, a single-sample prediction algorithm based on the multivariate linear regression model was constructed; S3. Output results: The value obtained by subtracting the MCS value of the mth test from the MCS value of the same multiple myeloma patient subject at the nth test is recorded as ΔMCS, and the response of the multiple myeloma patient subject to the combination treatment regimen is compared based on ΔMCS.

7. Use any of the following methods: B1. A method for clustering disease progression in multiple myeloma patients, comprising: The method comprises the following steps: S1. Obtaining sample data, for receiving the mth sample data and the nth sample data of the same multiple myeloma patient subject, where n is greater than m; The m sample data are ACBD3, ADAR, ADSS, ALDH2, ANP32E, ANXA2, ATF3, ATP8B2, CACYBP, CAPN2, CCND1, CCT3, CDC42SE1, CERS2, CHSY3, CLIC1, CLMN, COPA, CSNK1G3, DAP3, DENND1B, ENSA, EPRS, EPSTI1, EVL, FAM13A, FAM49A, FLAD1, FRZB, GLRX2, HAX1, HDGF, HLA-A, HLA-B, HLA-C, HLA-F, HLA-G, IL6R, ISG20L2, JTB, KLF2, LAMTOR2, LDHA, MCL1, MOXD1, MRPL24, MRPL9, MVP, Transcriptome expression data of 97 genes: MYL6, NDUFS2, NOP58, NOTCH2NL, NTAN1, PAK1, PI4KB, PIEZO1, PIK3AP1, PIM2, PIP5K1B, PMVK, POGZ, PPIA, PRCC, PRKCA, PRRC2C, PSMB4, PSMD4, RAB29, RCBTB2, SCAMP3, SCAPER, SDHC, SEL1L3, SELPLG, SHC1, SIDT1, SSR2, STAP1, TAP1, TIMM17A, TLR10, TMCO1, TOR1AIP2, TOR3A, TP53INP1, TPM3, TRANK1, TROVE2, UAP1, UBE2Q1, UBQLN4, UHMK1, VPS45, YY1AP1, ZC3H11A, ZFP36, and ZNF593; The nth sample data is the transcriptome expression data of the 97 genes in the ex vivo bone marrow CD138+ plasma cell sample of the multiple myeloma patient during the nth test; S2. Inputting the m-th sample data and the n-th sample data of the same multiple myeloma patient subject into the single sample prediction algorithm, respectively, and outputting the m-th detection myeloma classification score and the n-th detection myeloma classification score of the multiple myeloma patient subject, respectively recorded as the MCS value of the m-th detection and the MCS value of the n-th detection; The single sample prediction algorithm is constructed according to a method comprising the following steps: 1) Multiple ex vivo bone marrow CD138+ plasma cell samples from multiple myeloma patients were used as training sets, and the transcriptome expression data of the 97 genes in each sample in the training set was used as training set data. The training sets were divided into two groups using an unsupervised clustering method; 2) using the grouping labels of each sample in each group obtained in 1) and the transcriptome expression data of the 97 genes of the corresponding sample as input data, training a support vector machine model to generate a myeloma classification score for each sample in the training set, recorded as an MCS value; 3) Using the transcriptome expression data of 97 genes of all samples in the training set as the explanatory variable and the MCS values ​​of all samples in the training set as the response variable, a single-sample prediction algorithm based on the multivariate linear regression model was constructed; S3. Subtracting the MCS value of the mth test from the MCS value of the same multiple myeloma patient subject at the nth test as ΔMCS, and grouping the disease progression of the multiple myeloma patient subject according to ΔMCS; B2. A method for comparing disease progression in multiple myeloma patients, comprising the steps of: S1. Obtaining sample data, for receiving the mth sample data and the nth sample data of the same multiple myeloma patient subject, where n is greater than m; The m sample data are ACBD3, ADAR, ADSS, ALDH2, ANP32E, ANXA2, ATF3, ATP8B2, CACYBP, CAPN2, CCND1, CCT3, CDC42SE1, CERS2, CHSY3, CLIC1, CLMN, COPA, CSNK1G3, DAP3, DENND1B, ENSA, EPRS, EPSTI1, EVL, FAM13A, FAM49A, FLAD1, FRZB, GLRX2, HAX1, HDGF, HLA-A, HLA-B, HLA-C, HLA-F, HLA-G, IL6R, ISG20L2, JTB, KLF2, LAMTOR2, LDHA, MCL1, MOXD1, MRPL24, MRPL9, MVP, Transcriptome expression data of 97 genes: MYL6, NDUFS2, NOP58, NOTCH2NL, NTAN1, PAK1, PI4KB, PIEZO1, PIK3AP1, PIM2, PIP5K1B, PMVK, POGZ, PPIA, PRCC, PRKCA, PRRC2C, PSMB4, PSMD4, RAB29, RCBTB2, SCAMP3, SCAPER, SDHC, SEL1L3, SELPLG, SHC1, SIDT1, SSR2, STAP1, TAP1, TIMM17A, TLR10, TMCO1, TOR1AIP2, TOR3A, TP53INP1, TPM3, TRANK1, TROVE2, UAP1, UBE2Q1, UBQLN4, UHMK1, VPS45, YY1AP1, ZC3H11A, ZFP36, and ZNF593; The nth sample data is the transcriptome expression data of the 97 genes in the ex vivo bone marrow CD138+ plasma cell sample of the multiple myeloma patient during the nth test; S2. Inputting the m-th sample data and the n-th sample data of the same multiple myeloma patient subject into the single sample prediction algorithm, respectively, and outputting the m-th detection myeloma classification score and the n-th detection myeloma classification score of the multiple myeloma patient subject, respectively recorded as the MCS value of the m-th detection and the MCS value of the n-th detection; The single sample prediction algorithm is constructed according to a method comprising the following steps: 1) Multiple ex vivo bone marrow CD138+ plasma cell samples from multiple myeloma patients were used as training sets, and the transcriptome expression data of the 97 genes in each sample in the training set was used as training set data. The training sets were divided into two groups using an unsupervised clustering method; 2) using the grouping labels of each sample in each group obtained in 1) and the transcriptome expression data of the 97 genes of the corresponding sample as input data, training a support vector machine model to generate a myeloma classification score for each sample in the training set, recorded as an MCS value; 3) Using the transcriptome expression data of 97 genes of all samples in the training set as the explanatory variable and the MCS values ​​of all samples in the training set as the response variable, a single-sample prediction algorithm based on the multivariate linear regression model was constructed; S3. Subtracting the MCS value of the mth test from the MCS value of the same multiple myeloma patient subject at the nth test as ΔMCS, and comparing the disease progression of the multiple myeloma patient subject based on ΔMCS; B3. A method for comparing responses of multiple myeloma patients to a combination therapy regimen, comprising the steps of: S1. Obtaining sample data, for receiving the mth sample data and the nth sample data of the same multiple myeloma patient subject, where n is greater than m; The m sample data are ACBD3, ADAR, ADSS, ALDH2, ANP32E, ANXA2, ATF3, ATP8B2, CACYBP, CAPN2, CCND1, CCT3, CDC42SE1, CERS2, CHSY3, CLIC1, CLMN, COPA, CSNK1G3, DAP3, DENND1B, ENSA, EPRS, EPSTI1, EVL, FAM13A, FAM49A, FLAD1, FRZB, GLRX2, HAX1, HDGF, HLA-A, HLA-B, HLA-C, HLA-F, HLA-G, IL6R, ISG20L2, JTB, KLF2, LAMTOR2, LDHA, MCL1, MOXD1, MRPL24, MRPL9, MVP, Transcriptome expression data of 97 genes: MYL6, NDUFS2, NOP58, NOTCH2NL, NTAN1, PAK1, PI4KB, PIEZO1, PIK3AP1, PIM2, PIP5K1B, PMVK, POGZ, PPIA, PRCC, PRKCA, PRRC2C, PSMB4, PSMD4, RAB29, RCBTB2, SCAMP3, SCAPER, SDHC, SEL1L3, SELPLG, SHC1, SIDT1, SSR2, STAP1, TAP1, TIMM17A, TLR10, TMCO1, TOR1AIP2, TOR3A, TP53INP1, TPM3, TRANK1, TROVE2, UAP1, UBE2Q1, UBQLN4, UHMK1, VPS45, YY1AP1, ZC3H11A, ZFP36, and ZNF593; The nth sample data is the transcriptome expression data of the 97 genes in the ex vivo bone marrow CD138+ plasma cell sample of the multiple myeloma patient during the nth test; S2. Inputting the m-th sample data and the n-th sample data of the same multiple myeloma patient subject into the single sample prediction algorithm, respectively, and outputting the m-th detection myeloma classification score and the n-th detection myeloma classification score of the multiple myeloma patient subject, respectively recorded as the MCS value of the m-th detection and the MCS value of the n-th detection; The single sample prediction algorithm is constructed according to a method comprising the following steps: 1) Multiple ex vivo bone marrow CD138+ plasma cell samples from multiple myeloma patients were used as training sets, and the transcriptome expression data of the 97 genes in each sample in the training set was used as training set data. The training sets were divided into two groups using an unsupervised clustering method; 2) using the grouping labels of each sample in each group obtained in 1) and the transcriptome expression data of the 97 genes of the corresponding sample as input data, training a support vector machine model to generate a myeloma classification score for each sample in the training set, recorded as an MCS value; 3) Using the transcriptome expression data of 97 genes of all samples in the training set as the explanatory variable and the MCS values ​​of all samples in the training set as the response variable, a single-sample prediction algorithm based on the multivariate linear regression model was constructed; S3. Subtract the MCS value of the m-th test from the MCS value of the same multiple myeloma patient subject at the n-th test as ΔMCS, and compare the responses of the multiple myeloma patient subjects to the combination treatment regimen based on ΔMCS.

8. A method for constructing a single sample prediction algorithm, characterized by: The method comprises the following steps: 1) Multiple ex vivo bone marrow CD138+ plasma cell samples from multiple myeloma patients were used as training sets, and the transcriptome expression data of the 97 genes in each sample in the training set was used as training set data. The training sets were divided into two groups using an unsupervised clustering method; 2) using the grouping labels of each sample in each group obtained in 1) and the transcriptome expression data of the 97 genes of the corresponding sample as input data, training a support vector machine model to generate a myeloma classification score for each sample in the training set, recorded as an MCS value; 3) Using the transcriptome expression data of 97 genes of all samples in the training set as explanatory variables and the MCS values ​​of all samples in the training set as response variables, a single-sample prediction algorithm based on a multiple linear regression model was constructed.

Citation Information

Patent Citations

  • Multiple myeloma molecular subtype and application thereof to medication guidance

    CN108559778A

  • Methods and compositions for classification of samples

    US20180122508A1

  • Unique cancer associated fibroblast subsets predict response to immunotherapy

    US20240254565A1