Computer apparatus for predicting disease progression in multiple myeloma patients

By analyzing transcriptomic data from multiple myeloma patients and using unsupervised clustering and support vector machine models, myeloma classification scores are generated. This addresses the problem that existing technologies cannot effectively predict the treatment response of multiple myeloma patients, enabling individualized assessment of disease progression and treatment response.

CN120565079BActive Publication Date: 2026-02-10BEIJING NORMAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510729832.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2026-02-10
Estimated Expiration
2045-06-03

AI Technical Summary

Technical Problem

Existing diagnostic and molecular pathological testing methods cannot effectively predict the response of multiple myeloma patients to combination therapy regimens, resulting in strong heterogeneity of treatment regimens for individual patients, making it impossible to assess efficacy in a timely and effective manner, and affecting treatment outcomes.

Method used

Using a data processing device, transcriptome expression data of CD138+ plasma cell samples from ex vivo bone marrow of multiple myeloma patients were received. A single-sample prediction algorithm was constructed using unsupervised clustering and support vector machine models to generate myeloma classification scores and calculate ΔMCS values ​​to group or compare patients’ disease progression and treatment response.

Benefits of technology

It enables individualized detection of disease progression and treatment response in multiple myeloma patients, accurately grouping and comparing patients' disease progression and treatment response, and providing a more sensitive disease monitoring solution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120565079B_ABST
    Figure CN120565079B_ABST
Patent Text Reader

Abstract

The application discloses a computer device for predicting disease progression of a multiple myeloma patient. The application provides a data processing device, which comprises a memory, a processor and a computer program stored in the memory, and the computer program implements the following steps: S1, receiving sample data of the mth time and sample data of the nth time of a same multiple myeloma patient subject; S2, inputting the sample data of the mth time and the sample data of the nth time of the same subject into a single sample prediction algorithm respectively, and outputting a myeloma classification score (MCS) value detected in the mth time and a MCS value detected in the nth time; and S3, subtracting the MCS value detected in the mth time from the MCS value detected in the nth time to obtain a delta MCS, and grouping or comparing the disease progression of the subject according to the delta MCS. The single sample prediction algorithm and the delta MCS can be used for grouping or comparing the disease progression state of the multiple myeloma patient subject.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of computer information, and relates to a computer device for predicting disease progression of a multiple myeloma patient. BACKGROUND

[0002] Multiple myeloma (MM) is the second most common hematological malignancy, accounting for 10% of all hematological malignancies. There is a strong heterogeneity in the prognosis and response to standard treatment regimens of MM patients. Therefore, many studies attempt to construct a molecular typing scheme for guiding clinical treatment.

[0003] In recent years, a number of clinical trials have demonstrated that existing combination treatment regimens have achieved remarkable results in prolonging the survival of MM patients at the population level. However, there is a strong heterogeneity in the response of individual patients to combination treatment regimens, and existing diagnostic and molecular pathology detection methods cannot effectively predict whether individual patients can respond to combination treatment regimens and the degree of response. Therefore, it is important to select appropriate drug combinations based on biomarkers for the survival of MM patients.

[0004] In the treatment process of MM patients, timely and effective efficacy evaluation is particularly important for evaluating the response of MM patients to the current treatment regimen and adjusting the subsequent treatment regimen.

[0005] So far, MM is still incurable, and almost all MM patients will eventually relapse, so the development of a highly sensitive efficacy evaluation marker is of great practical significance for the treatment of MM patients. Current efficacy evaluation is mainly carried out from the tumor burden and organ damage of patients. However, the residual plasma cell clone component in the body of MM patients is closely related to the relapse trend of patients, so the establishment of an efficacy evaluation standard based on gene expression profile can more sensitively capture the overall expression signal at the bulk (tissue or cell population) level of MM patients, and thus provide a new disease monitoring scheme for MM patients from the perspective of plasma cell development.

[0006] Therefore, there is an urgent need in the art for an individualized detection tool that can predict the disease progression and treatment response of MM patients. SUMMARY

[0007] The technical problem solved by the present application is to provide a computer device for predicting disease progression of a multiple myeloma patient.

[0008] In order to solve the above technical problem, the first aspect of the present application provides a data processing device, which is a device for grouping disease progression of a multiple myeloma patient subject, comprising a memory, a processor and a computer program stored in the memory, wherein the processor executes the computer program to realize the following steps:

[0009] S1, receiving data: receiving the mth sample data and the nth sample data of the same multiple myeloma patient subject, n is greater than m;

[0010] The mth sample data is the transcriptome expression data of the 97 genes ACBD3, ADAR, ADSS, ALDH2, ANP32E, ANXA2, ATF3, ATP8B2, CACYBP, CAPN2, CCND1, CCT3, CDC42SE1, CERS2, CHSY3, CLIC1, CLMN, COPA, CSNK1G3, DAP3, DENND1B, ENSA, EPRS, EPSTI1, EVL, FAM13A, FAM49A, FLAD1, FRZB, GLRX2, HAX1, HDGF, HLA-A, HLA-B, HLA-C, HLA-F, HLA-G, IL6R, ISG20L2, JTB, KLF2, LAMTOR2, LDHA, MCL1, MOXD1, MRPL24, MRPL9, MVP, MYL6, NDUFS2, NOP58, NOTCH2NL, NTAN1, PAK1, PI4KB, PIEZO1, PIK3AP1, PIM2, PIP5K1B, PMVK, POGZ, PPIA, PRCC, PRKCA, PRRC2C, PSMB4, PSMD4, RAB29, RCBTB2, SCAMP3, SCAPER, SDHC, SEL1L3, SELPLG, SHC1, SIDT1, SSR2, STAP1, TAP1, TIMM17A, TLR10, TMCO1, TOR1AIP2, TOR3A, TP53INP1, TPM3, TRANK1, TROVE2, UAP1, UBE2Q1, UBQLN4, UHMK1, VPS45, YY1AP1, ZC3H11A, ZFP36 and ZNF593 in the ex vivo bone marrow CD138+ plasma cell sample of the multiple myeloma patient subject at the mth detection;

[0011] The nth sample data is the transcriptome expression data of the 97 genes in the ex vivo bone marrow CD138+ plasma cell sample of the multiple myeloma patient subject at the nth detection;

[0012] S2, data processing:

[0013] The mth sample data and the nth sample data of the same multiple myeloma patient subject are respectively input into a single sample prediction algorithm, and the myeloma classification scores of the mth detection and the nth detection of the multiple myeloma patient subject are output, which are respectively recorded as the MCS (Myeloma Classification Score) value of the mth detection and the MCS value of the nth detection;

[0014] The single sample prediction algorithm is constructed according to a method comprising the following steps:

[0015] 1) The CD138+ plasma cell samples of multiple myeloma patients in vitro are used as a training set, and the transcriptome expression data of the 97 genes of each sample in the training set are used as training set data. An unsupervised clustering method is used to divide the training set into two groups;

[0016] 2) The grouping label of each sample in each group obtained in 1) and the transcriptome expression data of the 97 genes of the corresponding sample are used as input data, and a support vector machine (SVM) model is trained to generate the myeloma classification score of each sample in the training set, which is recorded as the MCS value;

[0017] 3) The transcriptome expression data of the 97 genes of all samples in the training set are used as explanatory variables, and the MCS values of all samples in the training set are used as response variables, and a single sample prediction (SSP) algorithm based on a multiple linear regression (MLR) model is constructed;

[0018] S3, output result:

[0019] The MCS value of the nth detection of the same multiple myeloma patient subject is subtracted from the MCS value of the mth detection, and the value obtained is recorded as ΔMCS. According to ΔMCS, the disease progression of the multiple myeloma patient subject is grouped.

[0020] In the data processing device of the first aspect above, the disease progression of the multiple myeloma patient subject is grouped as follows: the subjects with ΔMCS <-2 are defined as the MCS progression group, and the subjects with ΔMCS ≥-2 are defined as the MCS stable group;

[0021] The MM patients in the MCS progression group are in a progressive state (compared with the mth detection sample, the genomic instability of the nth detection sample increases and / or the prognosis worsens);

[0022] The MM patients in the MCS stable group are in a stable state (no significant change in genomic condition and / or prognosis condition of the sample at the n th detection compared with the sample at the m th detection).

[0023] The disease progression of the MCS stable group is slower than that of the MCS progression group.

[0024] Specifically embodied in:

[0025] 1) The prognosis condition of the MCS stable group is better than that of the MCS progression group;

[0026] 2) The genomic stability of the MCS stable group is better than that of the MCS progression group.

[0027] The subject can be more than one subject.

[0028] In a second aspect, the present application provides a data processing device, which is a device for comparing the disease progression of multiple myeloma patient subjects, comprising a memory, a processor and a computer program stored in the memory, wherein the processor executes the computer program to realize the following steps:

[0029] S1, receiving data: receiving the m th sample data and the n th sample data of the same multiple myeloma patient subject, wherein n is greater than m;

[0030] the mth sample data is the transcriptome expression data of the 97 genes in the CD138+ plasma cell sample of the multiple myeloma patient subject at the mth detection;

[0031] the nth sample data is the transcriptome expression data of the 97 genes in the CD138+ plasma cell sample of the multiple myeloma patient subject at the nth detection;

[0032] S2, data processing:

[0033] the mth sample data and the nth sample data of the same multiple myeloma patient subject are respectively input into a single sample prediction algorithm, and the multiple myeloma patient subject's myeloma classification score at the mth detection and the multiple myeloma patient subject's myeloma classification score at the nth detection are output, respectively recorded as the mth detection MCS value and the nth detection MCS value;

[0034] the single sample prediction algorithm is constructed according to a method comprising the following steps:

[0035] 1) using bone marrow CD138+ plasma cell samples of multiple myeloma patients as a training set, using the transcriptome expression data of the 97 genes of each sample in the training set as training set data, using unsupervised clustering method to divide the training set into 2 groups;

[0036] 2) using the grouping label of each sample in each group obtained in 1) and the transcriptome expression data of the 97 genes of the corresponding sample as input data, training a support vector machine model, generating the myeloma classification score of each sample in the training set, denoted as MCS value;

[0037] 3) using the transcriptome expression data of the 97 genes of all samples in the training set as explanatory variables, and the MCS value of all samples in the training set as response variables, constructing a single-sample prediction algorithm based on a multiple linear regression model;

[0038] S3, output result:

[0039] The MCS value of the n-th detection of the same multiple myeloma patient subject is subtracted from the MCS value of the m-th detection, and the value obtained is denoted as ΔMCS. According to ΔMCS, the disease progression of the multiple myeloma patient subject is compared.

[0040] In the above second aspect data processing device, the comparison of the disease progression of the multiple myeloma patient subject can be as follows: the disease progression of the subject with ΔMCS≥-2 is slower than that of the subject with ΔMCS<-2;

[0041] Specifically embodied in:

[0042] 1) the prognosis of the subject with ΔMCS≥-2 is better than that of the subject with ΔMCS<-2;

[0043] 2) the genomic stability of the subject with ΔMCS≥-2 is better than that of the subject with ΔMCS<-2.

[0044] The subject can be more than two subjects.

[0045] In a third aspect, the present application provides a data processing device for comparing the response of multiple myeloma patient subjects to a combination treatment regimen, which comprises a memory, a processor and a computer program stored in the memory, and the processor executes the computer program to realize the following steps:

[0046] S1, receiving data: receiving the m-th sample data and the n-th sample data of the same multiple myeloma patient subject, n is greater than m;

[0047] the mth sample data is the transcriptome expression data of the 97 genes in the CD138+ plasma cell sample of the multiple myeloma patient subject at the mth detection;

[0048] the nth sample data is the transcriptome expression data of the 97 genes in the CD138+ plasma cell sample of the multiple myeloma patient subject at the nth detection;

[0049] S2, data processing:

[0050] The mth sample data and the nth sample data of the same multiple myeloma patient subject are respectively input into a single sample prediction algorithm, and the multiple myeloma patient subject's myeloma classification score at the mth detection and the multiple myeloma patient subject's myeloma classification score at the nth detection are output, respectively recorded as the mth detection MCS value and the nth detection MCS value;

[0051] The single sample prediction algorithm is constructed according to a method comprising the following steps:

[0052] 1) using multiple myeloma patient bone marrow CD138+ plasma cell samples as a training set, the transcriptome expression data of the 97 genes of each sample in the training set as training set data, using unsupervised clustering method to divide the training set into 2 groups;

[0053] 2) using the grouping label of each sample in each group obtained in 1) and the transcriptome expression data of the 97 genes of the corresponding sample as input data, training a support vector machine model to generate the myeloma classification score of each sample in the training set, denoted as MCS value;

[0054] 3) using the transcriptome expression data of the 97 genes of all samples in the training set as explanatory variables, and the MCS values of all samples in the training set as response variables, constructing a single-sample prediction algorithm based on a multiple linear regression model;

[0055] S3, output result:

[0056] Subtracting the MCS value of the mth detection from the MCS value of the nth detection of the same multiple myeloma patient subject to obtain a value denoted as ΔMCS, and comparing the response of the multiple myeloma patient subject to the combination treatment scheme according to ΔMCS.

[0057] In the above third aspect data processing device, the smaller the ΔMCS value, the worse the response of the multiple myeloma patient subject to the combination treatment scheme, and further, the multiple myeloma patient subject with a larger ΔMCS value responds better to the combination treatment scheme than the multiple myeloma patient subject with a smaller ΔMCS value.

[0058] The subject can be two or more subjects.

[0059] In a fourth aspect, the present application provides a computer program product, comprising a computer program, characterized in that: the computer program is executed by a processor to realize the steps in the data processing device in any one of the first to third aspects.

[0060] In a fifth aspect, the present application provides a computer readable storage medium, which stores a computer program, and the computer program causes a computer to execute the steps in the data processing device in any one of the first to third aspects.

[0061] In a sixth aspect, the present application provides any one of A1-A3 devices:

[0062] A1, a device for grouping disease progression of multiple myeloma patient subjects, the device comprising:

[0063] S1, a data receiving module, configured to receive the mth sample data and the nth sample data of the same multiple myeloma patient subject, wherein n is greater than m;

[0064] The mth sample data is the transcriptome expression data of the 97 genes in the CD138+ plasma cell sample of the multiple myeloma patient subject in the mth detection;

[0065] The nth sample data is the transcriptome expression data of the 97 genes in the CD138+ plasma cell sample of the multiple myeloma patient subject in the nth detection;

[0066] S2, a data processing module, configured to input the mth sample data and the nth sample data of the same multiple myeloma patient subject into a single sample prediction algorithm respectively, and output the myeloma classification score of the mth detection and the myeloma classification score of the nth detection of the multiple myeloma patient subject, which are recorded as the MCS value of the mth detection and the MCS value of the nth detection respectively;

[0067] The single-sample prediction algorithm is constructed according to a method comprising the following steps:

[0068] 1) The bone marrow CD138+ plasma cell sample of a multiple myeloma patient is used as a training set, the transcriptome expression data of the 97 genes in each sample in the training set is used as training set data, and an unsupervised clustering method is used to divide the training set into 2 groups;

[0069] 2) The grouping label of each sample in each group obtained in 1) and the transcriptome expression data of the 97 genes of the corresponding sample are used as input data, a support vector machine model is trained, and a multiple myeloma classification score of each sample in the training set is generated, denoted as MCS value;

[0070] 3) The transcriptome expression data of the 97 genes of all samples in the training set is used as the explanatory variable, and the MCS value of all samples in the training set is used as the response variable, and a single-sample prediction algorithm based on a multiple linear regression model is constructed;

[0071] S3, output result:

[0072] The value obtained by subtracting the MCS value of the mth detection from the MCS value of the nth detection of the same multiple myeloma patient subject is denoted as ΔMCS, and the disease progression of the multiple myeloma patient subject is grouped according to ΔMCS.

[0073] A2, a device for comparing the disease progression of a multiple myeloma patient subject, the device comprising:

[0074] S1, a data receiving module: for receiving the mth sample data and the nth sample data of the same multiple myeloma patient subject, n is greater than m;

[0075] The m sample data refers to the data from ex vivo bone marrow CD138+ plasma cell samples from multiple myeloma patients during the m-th test, including ACBD3, ADAR, ADSS, ALDH2, ANP32E, ANXA2, ATF3, ATP8B2, CACYBP, CAPN2, CCND1, CCT3, CDC42SE1, CERS2, CHSY3, CLIC1, CLMN, COPA, CSNK1G3, DAP3, DENND1B, ENSA, EPRS, EPSTI1, EVL, FAM13A, FAM49A, FLAD1, FRZB, GLRX2, HAX1, HDGF, HLA-A, HLA-B, HLA-C, HLA-F, HLA-G, IL6R, ISG20L2, JTB, KLF2, LAMTOR2, LDHA, MCL1, MOXD1, MRPL24, MRPL9, and MVP. Transcriptome expression data for 97 genes: MYL6, NDUFS2, NOP58, NOTCH2NL, NTAN1, PAK1, PI4KB, PIEZO1, PIK3AP1, PIM2, PIP5K1B, PMVK, POGZ, PPIA, PRCC, PRKCA, PRRC2C, PSMB4, PSMD4, RAB29, RCBTB2, SCAMP3, SCAPER, SDHC, SEL1L3, SELPLG, SHC1, SIDT1, SSR2, STAP1, TAP1, TIMM17A, TLR10, TMCO1, TOR1AIP2, TOR3A, TP53INP1, TPM3, TRANK1, TROVE2, UAP1, UBE2Q1, UBQLN4, UHMK1, VPS45, YY1AP1, ZC3H11A, ZFP36, and ZNF593.

[0076] The nth sample data refers to the transcriptome expression data of the 97 genes in the ex vivo bone marrow CD138+ plasma cell sample of the multiple myeloma patient subject during the nth test.

[0077] S2, Data Processing Module: Used to input the m-th sample data and n-th sample data of the same multiple myeloma patient subject into the single-sample prediction algorithm, and output the myeloma classification score of the m-th test and the myeloma classification score of the n-th test of the multiple myeloma patient subject, which are recorded as the MCS value of the m-th test and the MCS value of the n-th test, respectively.

[0078] The single-sample prediction algorithm is constructed according to the following steps:

[0079] 1) Multiple bone marrow CD138+ plasma cell samples from multiple myeloma patients were used as the training set. The transcriptome expression data of the 97 genes in each sample in the training set were used as the training set data. The training set was divided into two groups using an unsupervised clustering method.

[0080] 2) Take the grouping labels of each sample in each group obtained in 1) and the transcriptome expression data of the 97 genes in the corresponding sample as input data, train the support vector machine model, generate the myeloma classification score of each sample in the training set, and record it as the MCS value.

[0081] 3) Using the transcriptome expression data of 97 genes in all samples of the training set as explanatory variables and the MCS value of all samples in the training set as response variables, a single-sample prediction algorithm based on a multiple linear regression model is constructed.

[0082] S3, Output Results:

[0083] The value obtained by subtracting the MCS value of the mth test from the MCS value of the nth test of the same multiple myeloma patient subject is denoted as ΔMCS. The disease progression of the multiple myeloma patient subject is compared based on ΔMCS.

[0084] A3. An apparatus for comparing the response of multiple myeloma patients to combination therapy regimens, the apparatus comprising:

[0085] S1, Data Receiving Module: Used to receive the m-th and n-th sample data of the same multiple myeloma patient subject, where n is greater than m;

[0086] The m sample data refers to the data from ex vivo bone marrow CD138+ plasma cell samples from multiple myeloma patients during the m-th test, including ACBD3, ADAR, ADSS, ALDH2, ANP32E, ANXA2, ATF3, ATP8B2, CACYBP, CAPN2, CCND1, CCT3, CDC42SE1, CERS2, CHSY3, CLIC1, CLMN, COPA, CSNK1G3, DAP3, DENND1B, ENSA, EPRS, EPSTI1, EVL, FAM13A, FAM49A, FLAD1, FRZB, GLRX2, HAX1, HDGF, HLA-A, HLA-B, HLA-C, HLA-F, HLA-G, IL6R, ISG20L2, JTB, KLF2, LAMTOR2, LDHA, MCL1, MOXD1, MRPL24, MRPL9, and MVP. Transcriptome expression data for 97 genes: MYL6, NDUFS2, NOP58, NOTCH2NL, NTAN1, PAK1, PI4KB, PIEZO1, PIK3AP1, PIM2, PIP5K1B, PMVK, POGZ, PPIA, PRCC, PRKCA, PRRC2C, PSMB4, PSMD4, RAB29, RCBTB2, SCAMP3, SCAPER, SDHC, SEL1L3, SELPLG, SHC1, SIDT1, SSR2, STAP1, TAP1, TIMM17A, TLR10, TMCO1, TOR1AIP2, TOR3A, TP53INP1, TPM3, TRANK1, TROVE2, UAP1, UBE2Q1, UBQLN4, UHMK1, VPS45, YY1AP1, ZC3H11A, ZFP36, and ZNF593.

[0087] The nth sample data refers to the transcriptome expression data of the 97 genes in the ex vivo bone marrow CD138+ plasma cell sample of the multiple myeloma patient subject during the nth test.

[0088] S2, Data Processing Module: Used to input the m-th sample data and n-th sample data of the same multiple myeloma patient subject into the single-sample prediction algorithm, and output the myeloma classification score of the m-th test and the myeloma classification score of the n-th test of the multiple myeloma patient subject, which are recorded as the MCS value of the m-th test and the MCS value of the n-th test, respectively.

[0089] The single-sample prediction algorithm is constructed according to the following steps:

[0090] 1) Multiple bone marrow CD138+ plasma cell samples from multiple myeloma patients were used as the training set. The transcriptome expression data of the 97 genes in each sample in the training set were used as the training set data. The training set was divided into two groups using an unsupervised clustering method.

[0091] 2) Take the grouping labels of each sample in each group obtained in 1) and the transcriptome expression data of the 97 genes in the corresponding sample as input data, train the support vector machine model, generate the myeloma classification score of each sample in the training set, and record it as the MCS value.

[0092] 3) Using the transcriptome expression data of 97 genes in all samples of the training set as explanatory variables and the MCS value of all samples in the training set as response variables, a single-sample prediction algorithm based on a multiple linear regression model is constructed.

[0093] S3, Output Results:

[0094] The value obtained by subtracting the MCS value of the mth test from the MCS value of the nth test of the same multiple myeloma patient subject is denoted as ΔMCS. The response of the multiple myeloma patient subject to the combination treatment regimen is compared based on ΔMCS.

[0095] In a seventh aspect, the present invention provides any one of the following methods B1-B3:

[0096] B1. A method for grouping subjects with multiple myeloma based on disease progression, the method comprising the following steps:

[0097] S1. Obtain sample data, used to receive the m-th and n-th sample data of the same multiple myeloma patient subject, where n is greater than m;

[0098] The m sample data refers to the data from ex vivo bone marrow CD138+ plasma cell samples from multiple myeloma patients during the m-th test, including ACBD3, ADAR, ADSS, ALDH2, ANP32E, ANXA2, ATF3, ATP8B2, CACYBP, CAPN2, CCND1, CCT3, CDC42SE1, CERS2, CHSY3, CLIC1, CLMN, COPA, CSNK1G3, DAP3, DENND1B, ENSA, EPRS, EPSTI1, EVL, FAM13A, FAM49A, FLAD1, FRZB, GLRX2, HAX1, HDGF, HLA-A, HLA-B, HLA-C, HLA-F, HLA-G, IL6R, ISG20L2, JTB, KLF2, LAMTOR2, LDHA, MCL1, MOXD1, MRPL24, MRPL9, and MVP. Transcriptome expression data for 97 genes: MYL6, NDUFS2, NOP58, NOTCH2NL, NTAN1, PAK1, PI4KB, PIEZO1, PIK3AP1, PIM2, PIP5K1B, PMVK, POGZ, PPIA, PRCC, PRKCA, PRRC2C, PSMB4, PSMD4, RAB29, RCBTB2, SCAMP3, SCAPER, SDHC, SEL1L3, SELPLG, SHC1, SIDT1, SSR2, STAP1, TAP1, TIMM17A, TLR10, TMCO1, TOR1AIP2, TOR3A, TP53INP1, TPM3, TRANK1, TROVE2, UAP1, UBE2Q1, UBQLN4, UHMK1, VPS45, YY1AP1, ZC3H11A, ZFP36, and ZNF593.

[0099] The nth sample data refers to the transcriptome expression data of the 97 genes in the ex vivo bone marrow CD138+ plasma cell sample of the multiple myeloma patient subject during the nth test.

[0100] S2. Input the m-th sample data and n-th sample data of the same multiple myeloma patient subject into the single-sample prediction algorithm, and output the myeloma classification score of the m-th test and the myeloma classification score of the n-th test of the multiple myeloma patient subject, which are recorded as the MCS value of the m-th test and the MCS value of the n-th test, respectively.

[0101] The single-sample prediction algorithm is constructed according to the following steps:

[0102] 1) Multiple bone marrow CD138+ plasma cell samples from multiple myeloma patients were used as the training set. The transcriptome expression data of the 97 genes in each sample in the training set were used as the training set data. The training set was divided into two groups using an unsupervised clustering method.

[0103] 2) Take the grouping labels of each sample in each group obtained in 1) and the transcriptome expression data of the 97 genes in the corresponding sample as input data, train the support vector machine model, generate the myeloma classification score of each sample in the training set, and record it as the MCS value.

[0104] 3) Using the transcriptome expression data of 97 genes in all samples of the training set as explanatory variables and the MCS value of all samples in the training set as response variables, a single-sample prediction algorithm based on a multiple linear regression model is constructed.

[0105] S3. Subtract the MCS value of the mth test from the MCS value of the nth test of the same multiple myeloma patient subject and denote the value as ΔMCS. Based on ΔMCS, the disease progression of the multiple myeloma patient subject is grouped.

[0106] B2. A method for comparing disease progression in subjects with multiple myeloma, the method comprising the following steps:

[0107] S1. Obtain sample data, used to receive the m-th and n-th sample data of the same multiple myeloma patient subject, where n is greater than m;

[0108] The m sample data refers to the data from ex vivo bone marrow CD138+ plasma cell samples from multiple myeloma patients during the m-th test, including ACBD3, ADAR, ADSS, ALDH2, ANP32E, ANXA2, ATF3, ATP8B2, CACYBP, CAPN2, CCND1, CCT3, CDC42SE1, CERS2, CHSY3, CLIC1, CLMN, COPA, CSNK1G3, DAP3, DENND1B, ENSA, EPRS, EPSTI1, EVL, FAM13A, FAM49A, FLAD1, FRZB, GLRX2, HAX1, HDGF, HLA-A, HLA-B, HLA-C, HLA-F, HLA-G, IL6R, ISG20L2, JTB, KLF2, LAMTOR2, LDHA, MCL1, MOXD1, MRPL24, MRPL9, and MVP. Transcriptome expression data for 97 genes: MYL6, NDUFS2, NOP58, NOTCH2NL, NTAN1, PAK1, PI4KB, PIEZO1, PIK3AP1, PIM2, PIP5K1B, PMVK, POGZ, PPIA, PRCC, PRKCA, PRRC2C, PSMB4, PSMD4, RAB29, RCBTB2, SCAMP3, SCAPER, SDHC, SEL1L3, SELPLG, SHC1, SIDT1, SSR2, STAP1, TAP1, TIMM17A, TLR10, TMCO1, TOR1AIP2, TOR3A, TP53INP1, TPM3, TRANK1, TROVE2, UAP1, UBE2Q1, UBQLN4, UHMK1, VPS45, YY1AP1, ZC3H11A, ZFP36, and ZNF593.

[0109] The nth sample data refers to the transcriptome expression data of the 97 genes in the ex vivo bone marrow CD138+ plasma cell sample of the multiple myeloma patient subject during the nth test.

[0110] S2. Input the m-th sample data and n-th sample data of the same multiple myeloma patient subject into the single-sample prediction algorithm, and output the myeloma classification score of the m-th test and the myeloma classification score of the n-th test of the multiple myeloma patient subject, which are recorded as the MCS value of the m-th test and the MCS value of the n-th test, respectively.

[0111] The single-sample prediction algorithm is constructed according to the following steps:

[0112] 1) Multiple bone marrow CD138+ plasma cell samples from multiple myeloma patients were used as the training set. The transcriptome expression data of the 97 genes in each sample in the training set were used as the training set data. The training set was divided into two groups using an unsupervised clustering method.

[0113] 2) Take the grouping labels of each sample in each group obtained in 1) and the transcriptome expression data of the 97 genes in the corresponding sample as input data, train the support vector machine model, generate the myeloma classification score of each sample in the training set, and record it as the MCS value.

[0114] 3) Using the transcriptome expression data of 97 genes in all samples of the training set as explanatory variables and the MCS value of all samples in the training set as response variables, a single-sample prediction algorithm based on a multiple linear regression model is constructed.

[0115] S3. The value obtained by subtracting the MCS value of the mth test from the MCS value of the nth test of the same multiple myeloma patient subject is denoted as ΔMCS. The disease progression of the multiple myeloma patient subject is compared based on ΔMCS.

[0116] B3. A method for comparing the response of multiple myeloma patients to combination therapy regimens, the method comprising the following steps:

[0117] S1. Obtain sample data, used to receive the m-th and n-th sample data of the same multiple myeloma patient subject, where n is greater than m;

[0118] The m sample data refers to the data from ex vivo bone marrow CD138+ plasma cell samples from multiple myeloma patients during the m-th test, including ACBD3, ADAR, ADSS, ALDH2, ANP32E, ANXA2, ATF3, ATP8B2, CACYBP, CAPN2, CCND1, CCT3, CDC42SE1, CERS2, CHSY3, CLIC1, CLMN, COPA, CSNK1G3, DAP3, DENND1B, ENSA, EPRS, EPSTI1, EVL, FAM13A, FAM49A, FLAD1, FRZB, GLRX2, HAX1, HDGF, HLA-A, HLA-B, HLA-C, HLA-F, HLA-G, IL6R, ISG20L2, JTB, KLF2, LAMTOR2, LDHA, MCL1, MOXD1, MRPL24, MRPL9, and MVP. Transcriptome expression data for 97 genes: MYL6, NDUFS2, NOP58, NOTCH2NL, NTAN1, PAK1, PI4KB, PIEZO1, PIK3AP1, PIM2, PIP5K1B, PMVK, POGZ, PPIA, PRCC, PRKCA, PRRC2C, PSMB4, PSMD4, RAB29, RCBTB2, SCAMP3, SCAPER, SDHC, SEL1L3, SELPLG, SHC1, SIDT1, SSR2, STAP1, TAP1, TIMM17A, TLR10, TMCO1, TOR1AIP2, TOR3A, TP53INP1, TPM3, TRANK1, TROVE2, UAP1, UBE2Q1, UBQLN4, UHMK1, VPS45, YY1AP1, ZC3H11A, ZFP36, and ZNF593.

[0119] The nth sample data refers to the transcriptome expression data of the 97 genes in the ex vivo bone marrow CD138+ plasma cell sample of the multiple myeloma patient subject during the nth test.

[0120] S2. Input the m-th sample data and n-th sample data of the same multiple myeloma patient subject into the single-sample prediction algorithm, and output the myeloma classification score of the m-th test and the myeloma classification score of the n-th test of the multiple myeloma patient subject, which are recorded as the MCS value of the m-th test and the MCS value of the n-th test, respectively.

[0121] The single-sample prediction algorithm is constructed according to the following steps:

[0122] 1) Multiple bone marrow CD138+ plasma cell samples from multiple myeloma patients were used as the training set. The transcriptome expression data of the 97 genes in each sample in the training set were used as the training set data. The training set was divided into two groups using an unsupervised clustering method.

[0123] 2) Take the grouping labels of each sample in each group obtained in 1) and the transcriptome expression data of the 97 genes in the corresponding sample as input data, train the support vector machine model, generate the myeloma classification score of each sample in the training set, and record it as the MCS value.

[0124] 3) Using the transcriptome expression data of 97 genes in all samples of the training set as explanatory variables and the MCS value of all samples in the training set as response variables, a single-sample prediction algorithm based on a multiple linear regression model is constructed.

[0125] S3. The value obtained by subtracting the MCS value of the mth test from the MCS value of the nth test of the same multiple myeloma patient subject is denoted as ΔMCS. The response of the multiple myeloma patient subject to the combination treatment regimen is compared based on ΔMCS.

[0126] Eighthly, the present invention provides a method for constructing a single-sample prediction algorithm, the method comprising the following steps:

[0127] 1) Multiple bone marrow CD138+ plasma cell samples from multiple myeloma patients were used as the training set. The transcriptome expression data of the 97 genes in each sample in the training set were used as the training set data. The training set was divided into two groups using an unsupervised clustering method.

[0128] 2) Take the grouping labels of each sample in each group obtained in 1) and the transcriptome expression data of the 97 genes in the corresponding sample as input data, train the support vector machine model, generate the myeloma classification score of each sample in the training set, and record it as the MCS value.

[0129] 3) Using the transcriptome expression data of 97 genes in all samples of the training set as explanatory variables and the MCS value of all samples in the training set as response variables, a single-sample prediction algorithm based on a multiple linear regression model is constructed.

[0130] In the above text, the training set consists of CD138+ plasma cell samples from the bone marrow of more than 500 patients with multiple myeloma, specifically 840 samples from the MMRF database.

[0131] If the expression data mentioned above is in FPKM format, it will be directly used as expression data (denoted as GEPs data). If it is an SRA format file generated by the RNA-seq platform, it will be processed as follows: First, use fastq-dump in the SRA toolkit to convert the downloaded SRA file into a FASTQ file; then use Fastqc software for quality control; use STAR software to align the FASTQ file with the GRCh38 reference genome using the quality control sample file, and further generate a BAM format file; finally, use Stringtie software to generate FPKM format expression profile data, which will be used as expression data (denoted as GEPs data).

[0132] The combination therapy regimen mentioned above is the VRD (bortezomib + lenalidomide + dexamethasone) regimen.

[0133] In the above text, the prognosis is reflected by overall survival (OS); specifically, the prognosis of patients in the stable MCS group is better or better than that of patients in the progressive MCS group.

[0134] In the above text, genomic stability is reflected by genome-wide copy number variation and / or the average expression level of the CDC20-M gene; specifically, the genomic stability of patients in the MCS stable group is better or better than that of patients in the MCS progression group.

[0135] The methods described above are not for disease diagnosis or treatment. These methods are not necessarily diagnostic methods; diagnostic methods refer to the process of identifying, studying, and determining the cause or lesion state of a living human or animal. These methods are not necessarily aimed at obtaining a disease diagnosis or health status. These methods are not necessarily treatment methods; treatment methods refer to the process of blocking, alleviating, or eliminating the cause or lesion to restore or restore health or reduce suffering in a living human or animal.

[0136] The methods described above may not include the step of obtaining biological samples from animals. All methods may not target living human or animal bodies, but only data. All methods may be information processing methods in which all steps are performed by a data processing device such as a computer.

[0137] The data processing device is a computer processing device.

[0138] Experiments of this invention demonstrate that the single-sample prediction SSP algorithm constructed in this invention obtains the change in MCS value between different time points by detecting the MCS value of the same sample at different times, denoted as ΔMCS. ΔMCS is closely related to the prognosis, genomic alterations, and treatment response outcomes of multiple myeloma patients, providing a new possibility for dynamically monitoring the disease progression status of MM patients. Attached Figure Description

[0139] Figure 1 Flowchart of a computer device model for predicting disease progression in MM patients.

[0140] Figure 2 The distribution of ΔMCS in the MCS progression group and the MCS stability group.

[0141] Figure 3 Prognostic analysis for patients in the stable MCS group and the progressive MCS group.

[0142] Figure 4 Genomic alteration characteristics of the MCS progressive group and the MCS stable group.

[0143] Figure 5 The mean expression level of the CDC20-M gene was significantly increased in follow-up samples from the MCS progression group.

[0144] Figure 6 The correlation between ΔMCS and induction therapy outcomes. Detailed Implementation

[0145] The present invention will now be described in further detail with reference to specific embodiments. The given embodiments are merely illustrative of the invention and not intended to limit its scope. The embodiments provided below can serve as a guide for further improvements by those skilled in the art and do not constitute a limitation on the invention in any way.

[0146] Unless otherwise specified, the experimental methods used in the following examples are conventional methods, performed according to the techniques or conditions described in the literature in this field or according to the product instructions. Unless otherwise specified, the materials and reagents used in the following examples are commercially available.

[0147] Unless otherwise specified, the quantitative experiments in the following examples are all repeated three times, and the results are averaged.

[0148] The database information in the following examples is shown in Table 1.

[0149] Table 1 shows the mRNA transcriptome database.

[0150]

[0151] The aforementioned database records the transcriptome expression levels of various genes in bone marrow CD138+ plasma cells, and these expression levels are referred to as gene expression profiles (GEPs).

[0152] Table 2 lists the genbank numbers for the 97 genes.

[0153]

[0154] Example 1: Establishment of a computer device for predicting disease progression in MM patients

[0155] Figure 1 Flowchart of a computer device model for predicting disease progression in MM patients.

[0156] I. Identification of 97 genes used for predicting disease progression in MM patients

[0157] Using the GSE2658 multiple myeloma expression dataset provided by the NCBI GEO public database, 97 stable differentially expressed and relatively abundant taxonomic genes were retained through Pearson correlation analysis.

[0158] The names of the 97 genes are as follows (see Table 2 for details):

[0159] ACBD3, ADAR, ADSS, ALDH2, ANP32E, ANXA2, ATF3, ATP8B2, CACYBP, CAPN2, CCND1, CCT3, CDC42SE1, CERS2, CHSY3, CLIC1, CLMN, COPA, CSNK1G3, DAP3, DENND1B, ENSA, EPRS, EPSTI1, EVL , FAM13A, FAM49A, FLAD1, FRZB, GLRX2, HAX1, HDGF, HLA-A, HLA-B, HLA-C, HLA-F, HLA-G, IL6R, ISG20L2, JTB, KLF2, LAMTOR2, LDHA, MCL1, MOXD1, MRPL24, MRPL9, MVP, MYL6, NDUFS2, NO P58, NOTCH2NL, NTAN1, PAK1, PI4KB, PIEZO1, PIK3AP1, PIM2, PIP5K1B, PMVK, POGZ, PPIA, PRCC, PRKCA, PRRC2C, PSMB4, PSMD4, RAB29, RCBTB2, SCAMP3, SCAPER, SDHC, SEL1L3, SELPLG, SHC1, SIDT1, SSR2, STAP1, TAP1, TIMM17A, TLR10, TMCO1, TOR1AIP2, TOR3A, TP53INP1, TPM3, TRANK1, TROVE2, UAP1, UBE2Q1, UBQLN4, UHMK1, VPS45, YY1AP1, ZC3H11A, ZFP36, and ZNF593.

[0160] These 97 genes will then serve as classifier genes for predicting or grouping disease progression in MM patients.

[0161] II. Establishment of the Single Sample Prediction (SSP) Algorithm

[0162] Using transcriptome data from 840 samples in the MMRF database as the training set, a single-sample prediction (SSP) algorithm for the RNAseq platform was constructed:

[0163] 1. Based on unsupervised clustering methods, the dataset is divided into two groups at the dataset level, generating MCL1-M for all samples. high Or TLR10-M high Grouping tags

[0164] First, the gene expression profile (GEP) data of 97 classifier genes from 840 samples in the MMRF database were log2 transformed and z-score normalized. Then, unsupervised clustering was used to obtain the dataset-level MCL1-M from the normalized data. high Or TLR10-M high Grouping tags.

[0165] The unsupervised clustering method described above is as follows: The PAM clustering algorithm is used with the ConsensusClusterPlus R package for cluster analysis. The most stable clustering is achieved when K=2. The optimal grouping effect is determined based on the consistency matrix (K clusters have clear boundaries). The clustering results divide the samples in the dataset into two groups, named MCL1-M... high Group or TLR10-M high Group, i.e., add MCL1-M high Or TLR10-M high Grouping tags.

[0166] 2. Train the SVM model to generate dataset-level myeloma classification scores (MCS) for all samples.

[0167] Using the grouping labels of each sample in the two groups obtained above and the GEPs data of the 97 classifier genes corresponding to each sample as input data, a Support Vector Machine (SVM) model is trained to generate the MCS of all samples in the training set.

[0168] During training, a grid search method was used to determine the optimal parameter C=1. The model's robustness was tested within the training set using 10-fold cross-validation. The validation results demonstrated that recall, accuracy, and precision reached 100%. Under these conditions, the geometric margin of the training set samples (the distance from a single sample to the segmentation hyperplane) was output.

[0169] The geometric interval is a continuous variable with a positive or negative sign, defined as the MCS (myeloma classification score).

[0170] MCL1-M high The subtype has a negative MCS value, TLR10-M high The MCS of the subtype is positive.

[0171] 3. Develop an SSP algorithm based on Multiple Linear Regression (MLR) to calculate the MCS of a single sample.

[0172] Using GEPs data of 97 classifier genes from all samples in the training set as explanatory variables and MCS data of all samples in the training set as response variables, a single-sample prediction (SSP) algorithm based on the MLR model is constructed.

[0173] Expression profile data in FPKM format were directly obtained from MMRF data and used as GEPs data.

[0174] The GEPs data of 97 classifier genes in an ex vivo bone marrow CD138+ plasma cell sample from a single MM patient to be predicted are used as the input data for the Single Sample Prediction (SSP) algorithm described above, and the MCS of the single sample is output.

[0175] If the GEPs data of 97 classifier genes from ex vivo bone marrow CD138+ plasma cell samples from a single MM patient are generated in SRA format using an RNA-seq platform, the following processing steps are required: First, use fastq-dump from the SRA toolkit to convert the downloaded SRA file to a FASTQ file; then, use Fastqc software for quality control; using STAR software, align the FASTQ file with the GRCh38 reference genome using the quality-controlled sample file to further generate a BAM format file; finally, use Stringtie software to generate FPKM format expression profile data as GEPs data.

[0176] III. ΔMCS Predicts Disease Progression in MM Patients

[0177] 1. Definition of ΔMCS

[0178] MCS is closely related to the differential expression status of 97 classifier genes. Intracellular genomic alterations and the evolution of dominant cell clones constantly influence the transcriptomic status at the bulk (tissue or cell population) level in patients. Therefore, this study attempts to investigate whether changes in MCS values ​​at different time points in samples from the same MM patient are related to disease progression.

[0179] ΔMCS is used to measure the difference in MCS between two samples detected at different times. That is, ΔMCS = MCS value of the sample detected in the second (nth) test - MCS value of the sample detected in the first (mth) test; n>m, where m is a natural number greater than or equal to 1 and n is a natural number.

[0180] Since ΔMCS is an individualized indicator, it can be used to measure the change in MCS between any two samples from the same patient.

[0181] 2. Correlation between ΔMCS and patient prognosis and genomic alterations

[0182] We analyzed the correlation between ΔMCS and disease progression using longitudinal samples (ex vivo bone marrow CD138+ plasma cell samples) collected from 45 patients at different time points in the MMRF dataset. To make the analysis results more comparable, the ΔMCS for each patient in this section is the MCS value of the last follow-up sample minus the MCS value of the baseline follow-up (first visit test) sample.

[0183] 1) Obtain two MCS values ​​for the sample using the Single Sample Prediction (SSP) algorithm.

[0184] Data on 97 classifier gene GEPs were collected from 45 patients in the MMRF dataset at two different time points: the baseline period (m-th time) and the last follow-up period (n-th time). The data on 97 classifier gene GEPs of each sample were used as input data for the single sample prediction (SSP) algorithm constructed above, and the baseline follow-up MCS value and the last follow-up MCS value of each of the 45 patients were output.

[0185] 2)ΔMCS

[0186] ΔMCS for each patient = MCS value at last follow-up - MCS value at baseline follow-up

[0187] The results are shown in Tables 3 and 4.

[0188] Table 3 shows the characteristics of patients with ΔMCS≥-2 in the follow-up samples of the MMRF dataset.

[0189]

[0190]

[0191] Note: ΔMCS analysis was performed on paired longitudinal samples. The table shows the characteristics of patients with ΔMCS ≥ -2. ASCT: Autologous stem cell transplantation; NA: Not provided; Status: 0 = "Live", 1 = "Dead"; The baseline sample is from the time of the first visit for testing. Follow-up samples 1, 2, 3, and 4 are from the time of follow-up in chronological order. The MCS value of the last follow-up sample for each patient is recorded as the final follow-up MCS value.

[0192] Table 4 shows the characteristics of patients with ΔMCS<-2 in the follow-up samples of the MMRF dataset.

[0193]

[0194]

[0195] Note: The table shows the characteristics of patients with ΔMCS < -2. ASCT: Autologous stem cell transplantation; NA: Not provided; Status: 0 = "Live", 1 = "Dead". The baseline sample is from the initial visit testing period. Follow-up samples 1, 2, 3, and 4 are from the beginning to the end of the follow-up period. The MCS value of the last follow-up sample for each patient is recorded as the final follow-up MCS value.

[0196] 3) Based on the ΔMCS values ​​of the 45 patients, they were divided into an MCS progression group and an MCS stable group.

[0197] Patients with ΔMCS < -2 were defined as the MCS progression group, and patients with ΔMCS ≥ -2 were defined as the MCS stable group.

[0198] Disease progression was slower in the stable MCS group than in the progressive MCS group.

[0199] The 45 patients were divided into an MCS progression group and an MCS stable group based on their ΔMCS values. The results are as follows: Figure 2 As shown, the 45 samples were divided into 22 patients in the stable MCS group and 23 patients in the progressive MCS group, with -2 as the cutoff value.

[0200] MM patients in the MCS progression group are in a progressive state (compared to the m-th test sample, the n-th test sample has increased genomic instability and / or worse prognosis);

[0201] MM patients in the MCS stable group have a stable condition (compared to the m-th test sample, the genomic status and / or prognosis of the n-th test sample are not significantly different).

[0202] 3. Verification

[0203] Based on the defined 23 MCS progression group and 22 MCS stable group patients, a comprehensive analysis was conducted to analyze the prognostic and genomic alteration characteristics of the MCS stable and MCS progression groups.

[0204] 1) Prognosis

[0205] Based on the tracked prognostic data, the prognostic analysis results are as follows: Figure 3 As shown, patients in the MCS stable group had a longer overall survival (OS) compared to those in the MCS progression group. The median OS in the MCS stable group did not reach the target, while the median OS in the MCS progression group was 2.42 years (HR: 0.36, 95% CI: 0.15-0.87, p = 0.029). This indicates that the prognosis of patients in the MCS progression group was worse than that in the MCS stable group.

[0206] The prognostic trend was not affected by the age of the two groups of patients, whether they received ASCT treatment, or the presence of risk-related chromosomal translocations.

[0207] 2) Copy number analysis based on whole genome sequencing data

[0208] (1) Detection of genomic copy number variation

[0209] The MMRF database provides whole-genome sequencing data for the tested samples. Paired longitudinal sample BAM files can be downloaded from dbGaP (phs000748.v7.p4) (https: / / portal.gdc.cancer.gov). The BAM files were processed using the QDNAseq R package to generate SEG files. Copy number variation was calculated with a threshold of ±0.4. A log2 ratio greater than 0.4 was used as the threshold for chromosome amplification (one extra copy) or gain (two or more extra copies), and less than -0.4 was used as the threshold for chromosome loss. The copy number variation ratio was calculated by dividing the total length of copy number variation by the total length of the tested genome.

[0210] For example, a patient in the MCS progression group and a patient in the MCS stable group were used to generate seg files through whole genome sequencing for mapping. Figure 4As shown in Figures A and B, A represents a typical MCS-stable MM patient with a ΔMCS of 1.8 but carrying a high-risk t(4;14) chromosomal translocation. The figure shows the genomic copy number variation in paired baseline and last follow-up samples. No significant new copy number alteration events were observed in the follow-up samples. The patient's age, ASCT treatment status, common chromosomal translocation events, ISS stage, and OS (status) are indicated above the figure. Figure B represents a typical MCS-progressive MM patient carrying a high-risk t(11;14) chromosomal translocation with a ΔMCS of -4.2. The figure shows the genomic copy number variation in paired baseline and last follow-up samples. Significant new copy number alteration events were observed in the follow-up samples, including chromosomal fragmentation. The patient's age, ASCT treatment status, common chromosomal translocation events, ISS stage, and OS (status) are indicated above the figure. Short red lines in the last follow-up sample indicate subcloning gains and losses.

[0211] Copy number variation results for all available patient samples in the MCS stable group are as follows: Figure 4 As shown in Figure C, there was no significant increase in copy number variation in the follow-up samples of the MCS stable group, and the paired-samples t-test showed no significant difference. The figure shows data from 17 pairs of longitudinal samples.

[0212] Copy number variation results for all available patient samples in the MCS progression group, as follows: Figure 4 As shown in Figure D, the copy number variation was significantly increased in the follow-up samples of the MCS progression group (*: p < 0.05, paired samples t-test). The figure shows data from 15 pairs of longitudinal samples.

[0213] In conclusion, the genomic stability of patients in the MCS progression group was worse than that in the MCS stable group.

[0214] The correlation between genomic alterations in MCS stable and progressive groups and MM patients was then analyzed. In the MCS stable group, the genomic status of follow-up samples was largely similar to that of paired baseline samples. This genomic stability was not affected by high-risk genetic events. Multiple studies have shown that 1q+ and t(4;14) are classic high-risk prognostic markers. Among the 17 MCS stable group cases analyzed, 9 patients had high-fold amplification of 1q. Of these, 5 patients were found to have t(4;14) translocation events. Figure 4A and 4C). Conversely, in the MCS progression group, the proportion of copy number variations in follow-up samples was significantly increased compared to baseline samples. Most of the newly added genomic variation events were secondary subclonal events. In the longitudinal analysis of 15 patients, 5 MM patients did not have 1q+ detected in either baseline or follow-up samples. In another 4 MM patients, although 1q+ was present in baseline samples, no further increase in 1q copy number was found in follow-up samples. The t(11;14) translocation event is a recognized standard-risk prognostic marker in the field, and this chromosomal translocation event was detected in 6 patients in the MCS progression group. Figure 4 (B and 4D).

[0215] In the stable MCS group, 1q+ and t(4;14) events were frequently observed, while in the progressive MCS group, 1q+ deletion, stabilization, and t(11;14) events were frequently observed. These results, which seem to contradict clinical common sense, precisely confirm that a single prognostic biomarker is insufficient to assess the staged disease progression in an individual MM patient. Although classic prognostic biomarkers such as 1q+, t(4;14), and t(11;14) are significantly associated with patient prognosis at the population level, their detection rate in the MM population is generally not high, so they cannot be routinely used to monitor the disease progression status of MM patients.

[0216] The generation of ΔMCS depends solely on the expression of 97 genes closely related to plasma cell development and the pathogenesis of multiple myeloma (MM), and the expression of these genes can always be detected in patients at any stage. Therefore, the dynamic predictive function based on ΔMCS is something that traditional prognostic biomarkers cannot provide.

[0217] (2) Detection of CDC20-M gene expression level

[0218] CDC20-M is a group of genes closely related to malignant cell proliferation and genomic instability. The 139 genes comprising CDC20-M are as follows: ASF1B (55723), ASPM (259266), AURKA (6790), AURKB (9212), BIRC5 (332), BRCA1 (672), BUB1 (699), BUB1B (701), BUD31 (8896), C11orf82 (220042), CASC5 (57082), CASP2 (835), CCNA2 (890), CCNB1 (891), CCNB2 (9133), CDC20 (991), CDC2 5A(993), CDC45(8318), CDC6(990), CDCA2(157313), CDCA3(83461), CDCA7(83879), CDCA7L(55536), CDCA8(55143), CDK1(983), CDK2(1017), CDKN2C (1031), CDKN3(1033), CENPA(1058), CENPE(1062), CENPF(1063), CENPK(64105), CENPN(55839), CENPW(387103), CEP55(55165), CHEK1(1111), CKAP2 L(150468), CKS2(1164), DBF4(10926), DDX39A(10212), DEPDC1(55635), DEPDC1B(55789), DLGAP5(9787), DNMT1(1786), DSN1(79980), DTL(51514), DTYMK(1841), E2F7(144455), ECT2(1894), EME1(146956), ESPL1(9700), FAM64A(54478), FANCD2(2177), FANCI(55215), FBXO5(26271), FOXM1(2305 ), GAS2L3(283431), GINS1(9837), GINS2(51659), GJC1(10052), GPX7(2882), GTF2IRD2(84163), GTSE1(51512), HJURP(55355), HMGB2(3148), HMMR( 3161), IGF2BP3(10643), KIAA0101(9768), KIF11(3832), KIF14(9928), KIF15(56992), KIF20A(10112), KIF23(9493), KIF2C(11004), KIF4A(24137),KIFC1(3833), KNTC1(9735), KPNA2(3838), LMNB1(4001), LMNB2(84823), LRR1(122769), MAD2L1(4085), MCM2(4171), MCM3(4172), MCM6(4175), MCM8(84515), MELK(9833), MKI67(4288), MLF1IP(79682), MND1(84057), MYBL2(4605), NCAP G(64151), NCAPG2(54892), NCAPH(23397), NDC80(10403), NEK2(4751), NRM(11270), NUF2(83540), NUSAP1(51203), PA RPBP(55010), PBK(55872), PCNA(5111), PDIA4(9601), POC1A(25886), POLE2(5427), PRC1(9055), PTBP1(5725), PTTG1 (9232), RACGAP1(29127), RAE1(8480), RBBP8(5932), RFC2(5982), RFC3(5983), RFC4(5984), RNASEH2A(10535), RRM2 (6241), SGOL2(151246), SHCBP1(79801), SMC4(10051), SNRPB(6628), SPAG5(10615), SPC24(147841), STIL(6491), TA CC3 (10460), TCF3 (6929), TIMELESS (8914), TK1 (7083), TMEM48 (55706), TOP2A (7153), TPX2 (22974), TRIP13 (9319), TTK (7272), TYMS (7298), UBE2C (11065), UBE2S (27338), USP1 (7398), ZNF765 (91661), ZNF850 (342892), ZWINT (11130). The numbers in parentheses after each gene are their corresponding gene IDs on the NCBI website.

[0219] The correlation between ΔMCS and genomic instability was analyzed using the CDC20-M gene.

[0220] Analysis was performed on the CDC20-M expression levels of 45 samples obtained, and the results are as follows: Figure 5 As shown, Figure 5The figure shows a comparison of the mean CDC20-M expression levels between paired baseline samples and the last follow-up samples in the MCS stable group (A) or the MCS progressive group (B). *: P<0.05, **: P<0.01, paired samples t-test; it can be seen that compared with the baseline samples, the CDC20-M expression level in the last follow-up samples of the MCS progressive group was significantly increased.

[0221] Therefore, patients in the MCS progression group have poorer genomic stability than those in the MCS stable group, and patients in the MCS progression group have stronger genomic instability.

[0222] In summary, ΔMCS is significantly associated with prognosis and genomic alterations in MM patients. Lower ΔMCS or being in the MCS progression group predicts poorer prognosis and increased genomic instability.

[0223] IV. ΔMCS Predicts Treatment Outcomes in MM Patients

[0224] PETHEMA / GEM2012MENOS65 (September 2013 – November 2016): An open-label phase III clinical trial, data stored in the GSE147165 database. Patients first received 6 cycles of VRD induction therapy, followed by ASCT pretreated with busulfan-melphalan or melphalan, and finally 2 cycles of VRD consolidation therapy. Patients then entered...

[0225] The PETHEMA / GEM2014MAIN clinical trial involved patients receiving two years of maintenance therapy with either Rd or IRD. This clinical trial provided paired transcriptome data from 40 patients before and after VRD (bortezomib + lenalidomide + dexamethasone) induction therapy. Therefore, this study attempts to investigate whether ΔMCS can be used to predict patient response to combination therapy.

[0226] Bone marrow CD138+ plasma cell samples were collected from patients at initial diagnosis (NDMM) before induction therapy, and bone marrow CD138+ plasma cell samples were collected from patients at the time of measurable minimal residual disease (MRD) after induction therapy.

[0227] The GEPs data of 97 classifier genes for each sample at each time point are obtained as input data for the Single Sample Prediction (SSP) algorithm constructed above. The output is the MCS of a single sample at each time point: the MCS value at the initial diagnosis time and the MCS value at the measurable MRD time, denoted as MCS(MRD) and MCS(NDMM).

[0228] ΔMCS = MCS(MRD) - MCS(NDMM).

[0229] Based on the ΔMCS value, use GraphPad Prism software to create a scatter plot.

[0230] The results are as follows Figure 6 As shown, CR: clinically defined complete remission, VGPR: clinically defined very good partial remission, PR: clinically defined partial remission, SD: clinically defined stable disease, and PD: clinically defined progressive disease. The t-test was used to analyze the difference in ΔMCS among the groups. Cases with 1q+ chromosomal abnormalities are indicated by solid circles. *: P < 0.05, t-test;

[0231] It can be seen that the ΔMCS of CR patients is significantly higher than that of PR or PD patients. Therefore, the smaller the ΔMCS value, the worse the patient's response to the combination therapy regimen.

[0232] In summary, changes in MCS values ​​at different time points are closely related to patient prognosis, genomic alterations, and treatment response. Lower ΔMCS values ​​indicate greater genomic instability and poorer response to combination therapy regimens.

[0233] The present invention has been described in detail above. For those skilled in the art, the invention can be practiced in a wide range of ways with equivalent parameters, concentrations, and conditions without departing from its spirit and scope, and without requiring unnecessary experiments. Although specific embodiments have been given, it should be understood that further modifications can be made to the invention. In summary, according to the principles of the invention, this application is intended to include any changes, uses, or improvements to the invention, including changes made using conventional techniques known in the art that depart from the scope disclosed herein. Some of the essential features can be applied within the scope of the following appended claims.

Claims

1. A data processing apparatus for grouping multiple myeloma patient subjects based on disease progression, comprising a memory, a processor, and a computer program stored in the memory, characterized in that: The processor executes the computer program to perform the following steps: S1. Data reception: Receiving the m-th and n-th sample data of the same multiple myeloma patient subject, where n is greater than m; The m sample data refers to the CD138+ excised bone marrow samples from multiple myeloma patients during the m-th test. ACBD3, ADAR, ADSS, ALDH2, ANP32E, ANXA2, ATF3, ATP8B2, CACYBP, CAPN2, CCND1, CCT3, CDC42SE1, CERS2, CHSY3, CLIC1, CLMN, COPA, CSNK1G3, DAP3, DENND1B, ENSA, EPRS, EPSTI1, EV in plasma cell samples L, FAM13A, FAM49A, FLAD1, FRZB, GLRX2, HAX1, HDGF, HLA-A, HLA-B, HLA-C, HLA-F, HLA-G, IL6R , ISG20L2, JTB, KLF2, LAMTOR2, LDHA, MCL1, MOXD1, MRPL24, MRPL9, MVP, MYL6, NDUFS2, NOP58, Transcriptome expression data for 97 genes including NOTCH2NL, NTAN1, PAK1, PI4KB, PIEZO1, PIK3AP1, PIM2, PIP5K1B, PMVK, POGZ, PPIA, PRCC, PRKCA, PRRC2C, PSMB4, PSMD4, RAB29, RCBTB2, SCAMP3, SCAPER, SDHC, SEL1L3, SELPLG, SHC1, SIDT1, SSR2, STAP1, TAP1, TIMM17A, TLR10, TMCO1, TOR1AIP2, TOR3A, TP53INP1, TPM3, TRANK1, TROVE2, UAP1, UBE2Q1, UBQLN4, UHMK1, VPS45, YY1AP1, ZC3H11A, ZFP36, and ZNF593; The nth sample data refers to the transcriptome expression data of the 97 genes in the ex vivo bone marrow CD138+ plasma cell sample of the multiple myeloma patient subject during the nth test. S2, Data Processing: The m-th and n-th sample data of the same multiple myeloma patient are respectively input into the single-sample prediction algorithm, and the myeloma classification score of the m-th test and the myeloma classification score of the n-th test are output, which are recorded as the MCS value of the m-th test and the MCS value of the n-th test, respectively. The single-sample prediction algorithm is constructed according to the following steps: 1) Multiple bone marrow CD138+ plasma cell samples from multiple multiple myeloma patients were used as the training set. The transcriptome expression data of the 97 genes in each sample in the training set were used as the training data. Unsupervised clustering was used to divide the training set into two groups, labeled MCL1-M... high and TLR10-M high ; 2) Take the grouping labels of each sample in each group obtained in 1) and the transcriptome expression data of the 97 genes of the corresponding sample as input data, train the support vector machine model, generate the myeloma classification score of each sample in the training set, and record it as the MCS value. 3) Using the transcriptome expression data of 97 genes in all samples of the training set as explanatory variables and the MCS value of all samples in the training set as response variables, a single-sample prediction algorithm based on a multiple linear regression model is constructed. S3, Output Results: The value obtained by subtracting the MCS value of the mth test from the MCS value of the nth test of the same multiple myeloma patient subject is denoted as ΔMCS. The disease progression of the multiple myeloma patient subject is grouped according to ΔMCS.

2. A data processing apparatus for comparing disease progression in multiple myeloma patients, comprising a memory, a processor, and a computer program stored in the memory, characterized in that: The processor executes the computer program to perform the following steps: S1. Data reception: Receiving the m-th and n-th sample data of the same multiple myeloma patient subject, where n is greater than m; The m sample data refers to the CD138+ excised bone marrow samples from multiple myeloma patients during the m-th test. ACBD3, ADAR, ADSS, ALDH2, ANP32E, ANXA2, ATF3, ATP8B2, CACYBP, CAPN2, CCND1, CCT3, CDC42SE1, CERS2, CHSY3, CLIC1, CLMN, COPA, CSNK1G3, DAP3, DENND1B, ENSA, EPRS, EPSTI1, EV in plasma cell samples L, FAM13A, FAM49A, FLAD1, FRZB, GLRX2, HAX1, HDGF, HLA-A, HLA-B, HLA-C, HLA-F, HLA-G, IL6R , ISG20L2, JTB, KLF2, LAMTOR2, LDHA, MCL1, MOXD1, MRPL24, MRPL9, MVP, MYL6, NDUFS2, NOP58, Transcriptome expression data for 97 genes including NOTCH2NL, NTAN1, PAK1, PI4KB, PIEZO1, PIK3AP1, PIM2, PIP5K1B, PMVK, POGZ, PPIA, PRCC, PRKCA, PRRC2C, PSMB4, PSMD4, RAB29, RCBTB2, SCAMP3, SCAPER, SDHC, SEL1L3, SELPLG, SHC1, SIDT1, SSR2, STAP1, TAP1, TIMM17A, TLR10, TMCO1, TOR1AIP2, TOR3A, TP53INP1, TPM3, TRANK1, TROVE2, UAP1, UBE2Q1, UBQLN4, UHMK1, VPS45, YY1AP1, ZC3H11A, ZFP36, and ZNF593; The nth sample data refers to the transcriptome expression data of the 97 genes in the ex vivo bone marrow CD138+ plasma cell sample of the multiple myeloma patient subject during the nth test. S2, Data Processing: The m-th and n-th sample data of the same multiple myeloma patient are respectively input into the single-sample prediction algorithm, and the myeloma classification score of the m-th test and the myeloma classification score of the n-th test are output, which are recorded as the MCS value of the m-th test and the MCS value of the n-th test, respectively. The single-sample prediction algorithm is constructed according to the following steps: 1) Multiple bone marrow CD138+ plasma cell samples from multiple multiple myeloma patients were used as the training set. The transcriptome expression data of the 97 genes in each sample in the training set were used as the training data. Unsupervised clustering was used to divide the training set into two groups, labeled MCL1-M... high and TLR10-M high ; 2) Take the grouping labels of each sample in each group obtained in 1) and the transcriptome expression data of the 97 genes of the corresponding sample as input data, train the support vector machine model, generate the myeloma classification score of each sample in the training set, and record it as the MCS value. 3) Using the transcriptome expression data of 97 genes in all samples of the training set as explanatory variables and the MCS value of all samples in the training set as response variables, a single-sample prediction algorithm based on a multiple linear regression model is constructed. S3, Output Results: The value obtained by subtracting the MCS value of the mth test from the MCS value of the nth test of the same multiple myeloma patient subject is denoted as ΔMCS. The disease progression of the multiple myeloma patient subject is compared based on ΔMCS.

3. A data processing apparatus for comparing the responses of multiple myeloma patients to combination therapy regimens, comprising a memory, a processor, and a computer program stored in the memory, characterized in that: The processor executes the computer program to perform the following steps: S1. Data reception: Receiving the m-th and n-th sample data of the same multiple myeloma patient subject, where n is greater than m; The m sample data refers to the CD138+ excised bone marrow samples from multiple myeloma patients during the m-th test. ACBD3, ADAR, ADSS, ALDH2, ANP32E, ANXA2, ATF3, ATP8B2, CACYBP, CAPN2, CCND1, CCT3, CDC42SE1, CERS2, CHSY3, CLIC1, CLMN, COPA, CSNK1G3, DAP3, DENND1B, ENSA, EPRS, EPSTI1, EV in plasma cell samples L, FAM13A, FAM49A, FLAD1, FRZB, GLRX2, HAX1, HDGF, HLA-A, HLA-B, HLA-C, HLA-F, HLA-G, IL6R , ISG20L2, JTB, KLF2, LAMTOR2, LDHA, MCL1, MOXD1, MRPL24, MRPL9, MVP, MYL6, NDUFS2, NOP58, Transcriptome expression data for 97 genes including NOTCH2NL, NTAN1, PAK1, PI4KB, PIEZO1, PIK3AP1, PIM2, PIP5K1B, PMVK, POGZ, PPIA, PRCC, PRKCA, PRRC2C, PSMB4, PSMD4, RAB29, RCBTB2, SCAMP3, SCAPER, SDHC, SEL1L3, SELPLG, SHC1, SIDT1, SSR2, STAP1, TAP1, TIMM17A, TLR10, TMCO1, TOR1AIP2, TOR3A, TP53INP1, TPM3, TRANK1, TROVE2, UAP1, UBE2Q1, UBQLN4, UHMK1, VPS45, YY1AP1, ZC3H11A, ZFP36, and ZNF593; The nth sample data refers to the transcriptome expression data of the 97 genes in the ex vivo bone marrow CD138+ plasma cell sample of the multiple myeloma patient subject during the nth test. S2, Data Processing: The m-th and n-th sample data of the same multiple myeloma patient are respectively input into the single-sample prediction algorithm, and the myeloma classification score of the m-th test and the myeloma classification score of the n-th test are output, which are recorded as the MCS value of the m-th test and the MCS value of the n-th test, respectively. The single-sample prediction algorithm is constructed according to the following steps: 1) Multiple bone marrow CD138+ plasma cell samples from multiple multiple myeloma patients were used as the training set. The transcriptome expression data of the 97 genes in each sample in the training set were used as the training data. Unsupervised clustering was used to divide the training set into two groups, labeled MCL1-M... high and TLR10-M high ; 2) Take the grouping labels of each sample in each group obtained in 1) and the transcriptome expression data of the 97 genes of the corresponding sample as input data, train the support vector machine model, generate the myeloma classification score of each sample in the training set, and record it as the MCS value. 3) Using the transcriptome expression data of 97 genes in all samples of the training set as explanatory variables and the MCS value of all samples in the training set as response variables, a single-sample prediction algorithm based on a multiple linear regression model is constructed. S3, Output Results: The value obtained by subtracting the MCS value of the mth test from the MCS value of the nth test of the same multiple myeloma patient subject is denoted as ΔMCS. The response of the multiple myeloma patient subject to the combination treatment regimen is compared based on ΔMCS.

4. A computer program product, including a computer program, characterized in that: When executed by a processor, the computer program performs the steps described in the data processing apparatus according to any one of claims 1-3.

5. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program that causes a computer to perform the steps in the data processing apparatus of any one of claims 1-3.

6. Any of the following devices: A1. An apparatus for grouping subjects with multiple myeloma based on disease progression, the apparatus comprising: S1, Data Receiving Module: Used to receive the m-th and n-th sample data of the same multiple myeloma patient subject, where n is greater than m; The m sample data refers to the CD138+ excised bone marrow samples from multiple myeloma patients during the m-th test. ACBD3, ADAR, ADSS, ALDH2, ANP32E, ANXA2, ATF3, ATP8B2, CACYBP, CAPN2, CCND1, CCT3, CDC42SE1, CERS2, CHSY3, CLIC1, CLMN, COPA, CSNK1G3, DAP3, DENND1B, ENSA, EPRS, EPSTI1, EV in plasma cell samples L, FAM13A, FAM49A, FLAD1, FRZB, GLRX2, HAX1, HDGF, HLA-A, HLA-B, HLA-C, HLA-F, HLA-G, IL6R , ISG20L2, JTB, KLF2, LAMTOR2, LDHA, MCL1, MOXD1, MRPL24, MRPL9, MVP, MYL6, NDUFS2, NOP58, Transcriptome expression data for 97 genes including NOTCH2NL, NTAN1, PAK1, PI4KB, PIEZO1, PIK3AP1, PIM2, PIP5K1B, PMVK, POGZ, PPIA, PRCC, PRKCA, PRRC2C, PSMB4, PSMD4, RAB29, RCBTB2, SCAMP3, SCAPER, SDHC, SEL1L3, SELPLG, SHC1, SIDT1, SSR2, STAP1, TAP1, TIMM17A, TLR10, TMCO1, TOR1AIP2, TOR3A, TP53INP1, TPM3, TRANK1, TROVE2, UAP1, UBE2Q1, UBQLN4, UHMK1, VPS45, YY1AP1, ZC3H11A, ZFP36, and ZNF593; The nth sample data refers to the transcriptome expression data of the 97 genes in the ex vivo bone marrow CD138+ plasma cell sample of the multiple myeloma patient subject during the nth test. S2, Data Processing Module: Used to input the m-th sample data and n-th sample data of the same multiple myeloma patient subject into the single-sample prediction algorithm, and output the myeloma classification score of the m-th test and the myeloma classification score of the n-th test of the multiple myeloma patient subject, which are recorded as the MCS value of the m-th test and the MCS value of the n-th test, respectively. The single-sample prediction algorithm is constructed according to the following steps: 1) Multiple bone marrow CD138+ plasma cell samples from multiple multiple myeloma patients were used as the training set. The transcriptome expression data of the 97 genes in each sample in the training set were used as the training data. Unsupervised clustering was used to divide the training set into two groups, labeled MCL1-M... high and TLR10-M high ; 2) Take the grouping labels of each sample in each group obtained in 1) and the transcriptome expression data of the 97 genes of the corresponding sample as input data, train the support vector machine model, generate the myeloma classification score of each sample in the training set, and record it as the MCS value. 3) Using the transcriptome expression data of 97 genes in all samples of the training set as explanatory variables and the MCS value of all samples in the training set as response variables, a single-sample prediction algorithm based on a multiple linear regression model is constructed. S3, Output Results: The value obtained by subtracting the MCS value of the mth test from the MCS value of the nth test of the same multiple myeloma patient subject is denoted as ΔMCS. The disease progression of the multiple myeloma patient subject is grouped according to ΔMCS. A2. A device for comparing disease progression in subjects with multiple myeloma, the device comprising: S1, Data Receiving Module: Used to receive the m-th and n-th sample data of the same multiple myeloma patient subject, where n is greater than m; The m sample data refers to the CD138+ excised bone marrow samples from multiple myeloma patients during the m-th test. ACBD3, ADAR, ADSS, ALDH2, ANP32E, ANXA2, ATF3, ATP8B2, CACYBP, CAPN2, CCND1, CCT3, CDC42SE1, CERS2, CHSY3, CLIC1, CLMN, COPA, CSNK1G3, DAP3, DENND1B, ENSA, EPRS, EPSTI1, EV in plasma cell samples L, FAM13A, FAM49A, FLAD1, FRZB, GLRX2, HAX1, HDGF, HLA-A, HLA-B, HLA-C, HLA-F, HLA-G, IL6R , ISG20L2, JTB, KLF2, LAMTOR2, LDHA, MCL1, MOXD1, MRPL24, MRPL9, MVP, MYL6, NDUFS2, NOP58, Transcriptome expression data for 97 genes including NOTCH2NL, NTAN1, PAK1, PI4KB, PIEZO1, PIK3AP1, PIM2, PIP5K1B, PMVK, POGZ, PPIA, PRCC, PRKCA, PRRC2C, PSMB4, PSMD4, RAB29, RCBTB2, SCAMP3, SCAPER, SDHC, SEL1L3, SELPLG, SHC1, SIDT1, SSR2, STAP1, TAP1, TIMM17A, TLR10, TMCO1, TOR1AIP2, TOR3A, TP53INP1, TPM3, TRANK1, TROVE2, UAP1, UBE2Q1, UBQLN4, UHMK1, VPS45, YY1AP1, ZC3H11A, ZFP36, and ZNF593; The nth sample data refers to the transcriptome expression data of the 97 genes in the ex vivo bone marrow CD138+ plasma cell sample of the multiple myeloma patient subject during the nth test. S2, Data Processing Module: Used to input the m-th sample data and n-th sample data of the same multiple myeloma patient subject into the single-sample prediction algorithm, and output the myeloma classification score of the m-th test and the myeloma classification score of the n-th test of the multiple myeloma patient subject, which are recorded as the MCS value of the m-th test and the MCS value of the n-th test, respectively. The single-sample prediction algorithm is constructed according to the following steps: 1) Multiple bone marrow CD138+ plasma cell samples from multiple multiple myeloma patients were used as the training set. The transcriptome expression data of the 97 genes in each sample in the training set were used as the training data. Unsupervised clustering was used to divide the training set into two groups, labeled MCL1-M... high and TLR10-M high ; 2) Take the grouping labels of each sample in each group obtained in 1) and the transcriptome expression data of the 97 genes of the corresponding sample as input data, train the support vector machine model, generate the myeloma classification score of each sample in the training set, and record it as the MCS value. 3) Using the transcriptome expression data of 97 genes in all samples of the training set as explanatory variables and the MCS value of all samples in the training set as response variables, a single-sample prediction algorithm based on a multiple linear regression model is constructed. S3, Output Results: The value obtained by subtracting the MCS value of the mth test from the MCS value of the nth test of the same multiple myeloma patient subject is denoted as ΔMCS. The disease progression of the multiple myeloma patient subject is compared based on ΔMCS. A3. An apparatus for comparing the response of multiple myeloma patients to combination therapy regimens, the apparatus comprising: S1, Data Receiving Module: Used to receive the m-th and n-th sample data of the same multiple myeloma patient subject, where n is greater than m; The m sample data refers to the CD138+ excised bone marrow samples from multiple myeloma patients during the m-th test. ACBD3, ADAR, ADSS, ALDH2, ANP32E, ANXA2, ATF3, ATP8B2, CACYBP, CAPN2, CCND1, CCT3, CDC42SE1, CERS2, CHSY3, CLIC1, CLMN, COPA, CSNK1G3, DAP3, DENND1B, ENSA, EPRS, EPSTI1, EV in plasma cell samples L, FAM13A, FAM49A, FLAD1, FRZB, GLRX2, HAX1, HDGF, HLA-A, HLA-B, HLA-C, HLA-F, HLA-G, IL6R , ISG20L2, JTB, KLF2, LAMTOR2, LDHA, MCL1, MOXD1, MRPL24, MRPL9, MVP, MYL6, NDUFS2, NOP58, Transcriptome expression data for 97 genes including NOTCH2NL, NTAN1, PAK1, PI4KB, PIEZO1, PIK3AP1, PIM2, PIP5K1B, PMVK, POGZ, PPIA, PRCC, PRKCA, PRRC2C, PSMB4, PSMD4, RAB29, RCBTB2, SCAMP3, SCAPER, SDHC, SEL1L3, SELPLG, SHC1, SIDT1, SSR2, STAP1, TAP1, TIMM17A, TLR10, TMCO1, TOR1AIP2, TOR3A, TP53INP1, TPM3, TRANK1, TROVE2, UAP1, UBE2Q1, UBQLN4, UHMK1, VPS45, YY1AP1, ZC3H11A, ZFP36, and ZNF593; The nth sample data refers to the transcriptome expression data of the 97 genes in the ex vivo bone marrow CD138+ plasma cell sample of the multiple myeloma patient subject during the nth test. S2, Data Processing Module: Used to input the m-th sample data and n-th sample data of the same multiple myeloma patient subject into the single-sample prediction algorithm, and output the myeloma classification score of the m-th test and the myeloma classification score of the n-th test of the multiple myeloma patient subject, which are recorded as the MCS value of the m-th test and the MCS value of the n-th test, respectively. The single-sample prediction algorithm is constructed according to the following steps: 1) Multiple bone marrow CD138+ plasma cell samples from multiple multiple myeloma patients were used as the training set. The transcriptome expression data of the 97 genes in each sample in the training set were used as the training data. Unsupervised clustering was used to divide the training set into two groups, labeled MCL1-M... high and TLR10-M high ; 2) Take the grouping labels of each sample in each group obtained in 1) and the transcriptome expression data of the 97 genes of the corresponding sample as input data, train the support vector machine model, generate the myeloma classification score of each sample in the training set, and record it as the MCS value. 3) Using the transcriptome expression data of 97 genes in all samples of the training set as explanatory variables and the MCS value of all samples in the training set as response variables, a single-sample prediction algorithm based on a multiple linear regression model is constructed. S3, Output Results: The value obtained by subtracting the MCS value of the mth test from the MCS value of the nth test of the same multiple myeloma patient subject is denoted as ΔMCS. The response of the multiple myeloma patient subject to the combination treatment regimen is compared based on ΔMCS.

7. Any of the following methods: B1. A method for grouping subjects with multiple myeloma based on disease progression, characterized in that: The method includes the following steps: S1. Obtain sample data, used to receive the m-th and n-th sample data of the same multiple myeloma patient subject, where n is greater than m; The m sample data refers to the CD138+ excised bone marrow samples from multiple myeloma patients during the m-th test. ACBD3, ADAR, ADSS, ALDH2, ANP32E, ANXA2, ATF3, ATP8B2, CACYBP, CAPN2, CCND1, CCT3, CDC42SE1, CERS2, CHSY3, CLIC1, CLMN, COPA, CSNK1G3, DAP3, DENND1B, ENSA, EPRS, EPSTI1, EV in plasma cell samples L, FAM13A, FAM49A, FLAD1, FRZB, GLRX2, HAX1, HDGF, HLA-A, HLA-B, HLA-C, HLA-F, HLA-G, IL6R , ISG20L2, JTB, KLF2, LAMTOR2, LDHA, MCL1, MOXD1, MRPL24, MRPL9, MVP, MYL6, NDUFS2, NOP58, Transcriptome expression data for 97 genes including NOTCH2NL, NTAN1, PAK1, PI4KB, PIEZO1, PIK3AP1, PIM2, PIP5K1B, PMVK, POGZ, PPIA, PRCC, PRKCA, PRRC2C, PSMB4, PSMD4, RAB29, RCBTB2, SCAMP3, SCAPER, SDHC, SEL1L3, SELPLG, SHC1, SIDT1, SSR2, STAP1, TAP1, TIMM17A, TLR10, TMCO1, TOR1AIP2, TOR3A, TP53INP1, TPM3, TRANK1, TROVE2, UAP1, UBE2Q1, UBQLN4, UHMK1, VPS45, YY1AP1, ZC3H11A, ZFP36, and ZNF593; The nth sample data refers to the transcriptome expression data of the 97 genes in the ex vivo bone marrow CD138+ plasma cell sample of the multiple myeloma patient subject during the nth test. S2. Input the m-th sample data and n-th sample data of the same multiple myeloma patient subject into the single-sample prediction algorithm, and output the myeloma classification score of the m-th test and the myeloma classification score of the n-th test of the multiple myeloma patient subject, which are recorded as the MCS value of the m-th test and the MCS value of the n-th test, respectively. The single-sample prediction algorithm is constructed according to the following steps: 1) Multiple bone marrow CD138+ plasma cell samples from multiple multiple myeloma patients were used as the training set. The transcriptome expression data of the 97 genes in each sample in the training set were used as the training data. Unsupervised clustering was used to divide the training set into two groups, labeled MCL1-M... high and TLR10-M high ; 2) Take the grouping labels of each sample in each group obtained in 1) and the transcriptome expression data of the 97 genes of the corresponding sample as input data, train the support vector machine model, generate the myeloma classification score of each sample in the training set, and record it as the MCS value. 3) Using the transcriptome expression data of 97 genes in all samples of the training set as explanatory variables and the MCS value of all samples in the training set as response variables, a single-sample prediction algorithm based on a multiple linear regression model is constructed. S3. Subtract the MCS value of the mth test from the MCS value of the nth test of the same multiple myeloma patient subject and record the value as ΔMCS. Based on ΔMCS, the disease progression of the multiple myeloma patient subject is grouped. B2. A method for comparing disease progression in subjects with multiple myeloma, characterized in that the method comprises the following steps: S1. Obtain sample data, used to receive the m-th and n-th sample data of the same multiple myeloma patient subject, where n is greater than m; The m sample data refers to the CD138+ excised bone marrow samples from multiple myeloma patients during the m-th test. ACBD3, ADAR, ADSS, ALDH2, ANP32E, ANXA2, ATF3, ATP8B2, CACYBP, CAPN2, CCND1, CCT3, CDC42SE1, CERS2, CHSY3, CLIC1, CLMN, COPA, CSNK1G3, DAP3, DENND1B, ENSA, EPRS, EPSTI1, EV in plasma cell samples L, FAM13A, FAM49A, FLAD1, FRZB, GLRX2, HAX1, HDGF, HLA-A, HLA-B, HLA-C, HLA-F, HLA-G, IL6R , ISG20L2, JTB, KLF2, LAMTOR2, LDHA, MCL1, MOXD1, MRPL24, MRPL9, MVP, MYL6, NDUFS2, NOP58, Transcriptome expression data for 97 genes including NOTCH2NL, NTAN1, PAK1, PI4KB, PIEZO1, PIK3AP1, PIM2, PIP5K1B, PMVK, POGZ, PPIA, PRCC, PRKCA, PRRC2C, PSMB4, PSMD4, RAB29, RCBTB2, SCAMP3, SCAPER, SDHC, SEL1L3, SELPLG, SHC1, SIDT1, SSR2, STAP1, TAP1, TIMM17A, TLR10, TMCO1, TOR1AIP2, TOR3A, TP53INP1, TPM3, TRANK1, TROVE2, UAP1, UBE2Q1, UBQLN4, UHMK1, VPS45, YY1AP1, ZC3H11A, ZFP36, and ZNF593; The nth sample data refers to the transcriptome expression data of the 97 genes in the ex vivo bone marrow CD138+ plasma cell sample of the multiple myeloma patient subject during the nth test. S2. Input the m-th sample data and n-th sample data of the same multiple myeloma patient subject into the single-sample prediction algorithm, and output the myeloma classification score of the m-th test and the myeloma classification score of the n-th test of the multiple myeloma patient subject, which are recorded as the MCS value of the m-th test and the MCS value of the n-th test, respectively. The single-sample prediction algorithm is constructed according to the following steps: 1) Multiple bone marrow CD138+ plasma cell samples from multiple multiple myeloma patients were used as the training set. The transcriptome expression data of the 97 genes in each sample in the training set were used as the training data. Unsupervised clustering was used to divide the training set into two groups, labeled MCL1-M... high and TLR10-M high ; 2) Take the grouping labels of each sample in each group obtained in 1) and the transcriptome expression data of the 97 genes of the corresponding sample as input data, train the support vector machine model, generate the myeloma classification score of each sample in the training set, and record it as the MCS value. 3) Using the transcriptome expression data of 97 genes in all samples of the training set as explanatory variables and the MCS value of all samples in the training set as response variables, a single-sample prediction algorithm based on a multiple linear regression model is constructed. S3. The value obtained by subtracting the MCS value of the mth test from the MCS value of the nth test of the same multiple myeloma patient subject is denoted as ΔMCS. The disease progression of the multiple myeloma patient subject is compared based on ΔMCS. B3. A method for comparing the response of multiple myeloma patients to combination therapy regimens, characterized in that the method comprises the following steps: S1. Obtain sample data, used to receive the m-th and n-th sample data of the same multiple myeloma patient subject, where n is greater than m; The m sample data refers to the CD138+ excised bone marrow samples from multiple myeloma patients during the m-th test. ACBD3, ADAR, ADSS, ALDH2, ANP32E, ANXA2, ATF3, ATP8B2, CACYBP, CAPN2, CCND1, CCT3, CDC42SE1, CERS2, CHSY3, CLIC1, CLMN, COPA, CSNK1G3, DAP3, DENND1B, ENSA, EPRS, EPSTI1, EV in plasma cell samples L, FAM13A, FAM49A, FLAD1, FRZB, GLRX2, HAX1, HDGF, HLA-A, HLA-B, HLA-C, HLA-F, HLA-G, IL6R , ISG20L2, JTB, KLF2, LAMTOR2, LDHA, MCL1, MOXD1, MRPL24, MRPL9, MVP, MYL6, NDUFS2, NOP58, Transcriptome expression data for 97 genes including NOTCH2NL, NTAN1, PAK1, PI4KB, PIEZO1, PIK3AP1, PIM2, PIP5K1B, PMVK, POGZ, PPIA, PRCC, PRKCA, PRRC2C, PSMB4, PSMD4, RAB29, RCBTB2, SCAMP3, SCAPER, SDHC, SEL1L3, SELPLG, SHC1, SIDT1, SSR2, STAP1, TAP1, TIMM17A, TLR10, TMCO1, TOR1AIP2, TOR3A, TP53INP1, TPM3, TRANK1, TROVE2, UAP1, UBE2Q1, UBQLN4, UHMK1, VPS45, YY1AP1, ZC3H11A, ZFP36, and ZNF593; The nth sample data refers to the transcriptome expression data of the 97 genes in the ex vivo bone marrow CD138+ plasma cell sample of the multiple myeloma patient subject during the nth test. S2. Input the m-th sample data and n-th sample data of the same multiple myeloma patient subject into the single-sample prediction algorithm, and output the myeloma classification score of the m-th test and the myeloma classification score of the n-th test of the multiple myeloma patient subject, which are recorded as the MCS value of the m-th test and the MCS value of the n-th test, respectively. The single-sample prediction algorithm is constructed according to the following steps: 1) Multiple bone marrow CD138+ plasma cell samples from multiple multiple myeloma patients were used as the training set. The transcriptome expression data of the 97 genes in each sample in the training set were used as the training data. Unsupervised clustering was used to divide the training set into two groups, labeled MCL1-M... high and TLR10-M high ; 2) Take the grouping labels of each sample in each group obtained in 1) and the transcriptome expression data of the 97 genes of the corresponding sample as input data, train the support vector machine model, generate the myeloma classification score of each sample in the training set, and record it as the MCS value. 3) Using the transcriptome expression data of 97 genes in all samples of the training set as explanatory variables and the MCS value of all samples in the training set as response variables, a single-sample prediction algorithm based on a multiple linear regression model is constructed. S3. The value obtained by subtracting the MCS value of the mth test from the MCS value of the nth test of the same multiple myeloma patient subject is denoted as ΔMCS. The response of the multiple myeloma patient subject to the combination treatment regimen is compared based on ΔMCS.

8. A method for constructing a single-sample prediction algorithm, characterized in that: The method includes the following steps: 1) Multiple bone marrow CD138+ plasma cell samples from multiple multiple myeloma patients were used as the training set. Transcriptome expression data of 97 genes in each sample in the training set were used as training data. Unsupervised clustering was used to divide the training set into two groups, labeled MCL1-M... high and TLR10-M high ; The 97 genes are ACBD3, ADAR, ADSS, ALDH2, ANP32E, ANXA2, ATF3, ATP8B2, CACYBP, CAPN2, CCND1, CCT3, CDC42SE1, CERS2, CHSY3, CLIC1, CLMN, COPA, CSNK1G3, DAP3, DENND1B, ENSA, EPRS, EPST I1, EVL, FAM13A, FAM49A, FLAD1, FRZB, GLRX2, HAX1, HDGF, HLA-A, HLA-B, HLA-C, HLA-F, HLA-G, IL6R, ISG20L2, JTB, KLF2, LAMTOR2, LDHA, MCL1, MOXD1, MRPL24, MRPL9, MVP, MYL6, NDUFS 2. NOP58, NOTCH2NL, NTAN1, PAK1, PI4KB, PIEZO1, PIK3AP1, PIM2, PIP5K1B, PMVK, POGZ, PPIA, PRCC, PRKCA, PRRC2C, PSMB4, PSMD4, RAB29, RCBTB2, SCAMP3, SCAPER, SDHC, SEL1L3, SELPL G, SHC1, SIDT1, SSR2, STAP1, TAP1, TIMM17A, TLR10, TMCO1, TOR1AIP2, TOR3A, TP53INP1, TP M3, TRANK1, TROVE2, UAP1, UBE2Q1, UBQLN4, UHMK1, VPS45, YY1AP1, ZC3H11A, ZFP36 and ZNF593; 2) Take the grouping labels of each sample in each group obtained in 1) and the transcriptome expression data of the 97 genes of the corresponding sample as input data, train the support vector machine model, generate the myeloma classification score of each sample in the training set, and record it as the MCS value. 3) Using the transcriptome expression data of 97 genes in all samples of the training set as explanatory variables and the MCS value of all samples in the training set as response variables, a single-sample prediction algorithm based on a multiple linear regression model is constructed.