Gene combination for evaluating prognosis of multiple myeloma and prmm risk and evaluation method thereof

By combining biomarkers of MCM3, FABP5, MIR4435-2HG, FH, HELLS, CCDC34, RPS6KA1, FEN1, and OPN3 with single-cell transcriptome sequencing technology, the problem of accuracy in prognostic assessment for multiple myeloma patients has been solved, especially in risk identification of primary refractory multiple myeloma, providing precise prediction and treatment guidance.

CN120945052BActive Publication Date: 2026-05-29CHINESE PEOPLES LIBERATION ARMY NAVAL SPECIALTY MEDICAL CENT

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINESE PEOPLES LIBERATION ARMY NAVAL SPECIALTY MEDICAL CENT
Filing Date
2025-08-05
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing technologies are insufficient for accurately identifying and assessing the prognosis of patients with multiple myeloma, especially the risk of primary refractory multiple myeloma, leading to treatment ineffectiveness and poor prognosis.

Method used

The combination of genes MCM3, FABP5, MIR4435-2HG, FH, HELLS, CCDC34, RPS6KA1, FEN1 and OPN3 was used as biomarkers. Molecular differences were analyzed using single-cell transcriptome sequencing technology, and a risk scoring model was constructed to assess the prognostic risk of patients.

Benefits of technology

It enables high-precision and specific identification of patients with primary refractory multiple myeloma, predicts their sensitivity to standard therapy and mortality risk, and provides a basis for risk stratification and drug selection in clinical treatment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120945052B_ABST
    Figure CN120945052B_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of molecular diagnosis, and relates to a gene combination for evaluating prognosis of multiple myeloma and PRMM risk and a method for evaluating the same. The gene combination is a combination of MCM3, FABP5, MIR4435-2HG, FH, HELLS, CCDC34, RPS6KA1, FEN1 and OPN3 genes. Detection with the gene combination as a biomarker can specifically and accurately evaluate the prognosis of a multiple myeloma patient, especially the death risk, and the risk of the multiple myeloma patient being a primary refractory multiple myeloma according to the detection result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of molecular diagnostic technology and relates to gene combinations and assessment methods for evaluating the prognosis of multiple myeloma and the risk of PRMM. Background Technology

[0002] Multiple myeloma (MM) is a clonal plasma cell malignancy. Although significant progress has been made in its treatment, a considerable number of patients still do not respond to initial treatment regimens. Primary refractory multiple myeloma (PRMM) specifically refers to patients who do not achieve a minimum response (MR) or better with initial treatment regimens based on proteasome inhibitors (such as bortezomib), representing "de novo resistance," and these patients have an extremely poor prognosis. Therefore, accurate identification of PRMM patients before treatment initiation is of crucial clinical significance for avoiding ineffective treatment, guiding personalized medication, and improving patient outcomes.

[0003] Current technologies primarily rely on indicators such as high-risk cytogenetic abnormalities to predict the prognosis of multiple myeloma, but their accuracy in distinguishing PRMM is limited. While traditional whole-transcriptome sequencing (Bulk RNA-seq) can provide panoramic gene expression information, the true signals of malignant plasma cells are often averaged and masked due to the presence of a large number of stromal cells, immune cells, and other microenvironmental components in bone marrow samples. This makes it exceptionally difficult to discover the key molecular features that truly drive primary drug resistance. Summary of the Invention

[0004] The primary objective of this invention is to provide a gene combination that can be used as a biomarker to specifically and accurately assess the prognosis of patients with multiple myeloma, especially the risk of death, and the risk of multiple myeloma patients having primary refractory multiple myeloma.

[0005] To achieve this objective, in a basic implementation, the present invention provides a gene combination, which is a combination of the genes MCM3, FABP5, MIR4435-2HG, FH, HELLS, CCDC34, RPS6KA1, FEN1 and OPN3 (hereinafter referred to as "PRMM-Sig9").

[0006] In the above gene combinations:

[0007] MCM3 (Minichromosome Maintenance Complex Component 3) is located on human chromosome 6p12.2, and its mRNA reference sequence has the GenBank accession number NM_001270472.3. The protein encoded by this gene is a core member of the DNA replication permission complex, which is crucial for the initiation of DNA replication and is closely related to cell proliferation and genome stability.

[0008] FABP5 (Fatty Acid Binding Protein 5) is located on human chromosome 8q21.13, and its mRNA reference sequence has the GenBank accession number NM_001444.3. The protein encoded by this gene is mainly involved in the transport and metabolism of long-chain fatty acids in cells, and can regulate gene expression, affecting cell proliferation, differentiation, and inflammatory responses.

[0009] MIR4435-2HG (MIR4435-2 Host Gene) is located on human chromosome 2q13, and its mRNA reference sequence has the GenBank accession number NM_001013744.1. This gene is a long non-coding RNA (lncRNA), and its function is still under investigation, but it has been found to be associated with the progression and prognosis of various tumors.

[0010] FH (Fumarate Hydratase) is located on human chromosome 1q43, and its mRNA reference sequence has the GenBank accession number NM_000143.4. The protein encoded by this gene is a key enzyme in the tricarboxylic acid cycle, catalyzing the hydration of fumarate to malate. Its inactivation is associated with hereditary leiomyomatosis and renal cell carcinoma.

[0011] HELLS (Helicase, Lymphoid-Specific) is located on human chromosome 10q23.33, and its mRNA reference sequence has the GenBank accession number NM_001289067.2. The protein encoded by this gene is a DNA helicase that plays a crucial role in DNA methylation and chromatin remodeling, and is an important factor in epigenetic regulation and tumorigenesis.

[0012] CCDC34 (Coiled-Coil Domain Containing 34) is located on human chromosome 11p14.1, and its mRNA reference sequence has the GenBank accession number NM_030771.2. This gene encodes a protein containing a coiled-coil domain, the exact function of which is not fully understood, but studies have shown that it may be involved in cell proliferation and tumor progression.

[0013] RPS6KA1 (Ribosomal Protein S6 Kinase A1) is located on human chromosome 1p36.11, and its mRNA reference sequence has the GenBank accession number NM_001006665.2. The p90RSK protein encoded by this gene is a key kinase downstream of the RAS-MAPK signaling pathway, regulating various processes such as cell growth, survival, proliferation, and differentiation.

[0014] FEN1 (Flap Endonuclease 1) is located on human chromosome 11q12.2, and its mRNA reference sequence has the GenBank accession number NM_004111.6. The protein encoded by this gene is a structure-specific endonuclease that plays a central role in DNA replication and repair (especially the maturation process of the Okazaki segment) and is essential for maintaining genome stability.

[0015] OPN3 (Opsin 3) is located on human chromosome 1q43, and its mRNA reference sequence has the GenBank accession number NM_001030011.1. The protein encoded by this gene belongs to the opsin family and is a G protein-coupled receptor. It was initially thought to be related to photosensitivity, but recent studies have found that it is expressed in a variety of non-visual tissues and may be involved in regulating cellular signaling and physiological processes.

[0016] A second objective of this invention is to provide the use of the genomic collaboration described above as a biomarker in the preparation of reagents / reagent kits to specifically and accurately assess the prognosis of patients with multiple myeloma, particularly the risk of death, and the risk of multiple myeloma patients having primary refractory multiple myeloma.

[0017] To achieve this objective, in a basic implementation, the present invention provides the use of the genomic collaboration described above as a biomarker in the preparation of reagents / reagent kits, said use including assessing the prognosis of patients with multiple myeloma.

[0018] In a preferred embodiment, the present invention provides the use of the genomic collaboration described above as a biomarker in the preparation of reagents / reagent kits, wherein the prognosis is the risk of death.

[0019] In a preferred embodiment, the present invention provides the use of the genomic collaboration described above as a biomarker in the preparation of reagents / reagent kits, wherein said use further includes assessing the risk of a multiple myeloma patient having primary refractory multiple myeloma.

[0020] In a preferred embodiment, the present invention provides the use of the genomic collaboration described above as a biomarker in the preparation of reagents / reagent kits, wherein said use further includes assessing the likelihood of a patient’s natural resistance to a proteasome inhibitor.

[0021] In a preferred embodiment, the present invention provides the use of the genomic collaboration as a biomarker as described above in the preparation of reagents / reagent kits, wherein the proteasome inhibitor includes bortezomib.

[0022] A third objective of this invention is to provide an in vitro auxiliary method for assessing the prognosis of patients with multiple myeloma, so as to specifically and accurately assess the prognosis of patients with multiple myeloma, especially the risk of death, and the risk of multiple myeloma patients having primary refractory multiple myeloma.

[0023] To achieve this objective, in a basic implementation, the present invention provides a method for in vitro adjuvant assessment of the prognosis of patients with multiple myeloma, the method comprising the following steps:

[0024] (1) Collect ex vivo biological samples from patients with multiple myeloma;

[0025] (2) Detect the expression level of mRNA or cDNA of each gene in the gene combination as described above in the in vitro biological sample;

[0026] (3) Calculate the risk score using the risk model formula based on the detection results of the expression levels of each gene;

[0027] (4) The risk score is compared with a preset threshold to assess the mortality risk of patients with multiple myeloma, wherein a risk score above the preset threshold indicates a significantly increased mortality risk.

[0028] In a preferred embodiment, the present invention provides a method for in vitro auxiliary assessment of the prognosis of patients with multiple myeloma, wherein in step (1), the ex vivo biological sample is an original bone marrow sample.

[0029] In a preferred embodiment, the present invention provides a method for in vitro auxiliary assessment of the prognosis of patients with multiple myeloma, wherein in step (2), the expression level of mRNA of each gene in the gene combination as described above in the ex vivo biological sample is obtained by transcriptome sequencing.

[0030] The transcriptome sequencing used in this invention is preferably whole transcriptome sequencing or single-cell transcriptome sequencing.

[0031] Whole transcriptome sequencing (also known as RNA-seq) is the application of next-generation sequencing (NGS) technology in the field of transcriptomics. Its core objective is to comprehensively and quantitatively analyze the types and abundance of all RNA transcripts in a specific cell or tissue under a certain state. The basic process includes: first, extracting total RNA and enriching messenger RNA (mRNA) from the biological sample to be tested (such as CD138+ plasma cells in this invention); then, randomly fragmenting the purified RNA into short fragments, using them as templates to synthesize double-stranded cDNA through reverse transcription, and then repairing the ends of the cDNA fragments, adding A tails, and ligating sequencing adapters to construct a cDNA library that can be sequenced on a sequencer; then, loading the library onto a high-throughput sequencer (such as Illumina's NovaSeq or NextSeq platform) for massively parallel sequencing, reading the base sequences (reads) of hundreds of millions of short cDNA fragments; finally, using bioinformatics methods, aligning these reads to a reference genome, counting the number of reads mapped to each gene, and standardizing them (e.g., calculating TPM or FPKM values) to obtain the precise expression abundance information of each gene, thereby achieving precise quantification of the entire transcriptome. Those skilled in the art are familiar with the above-described standard RNA-seq procedure, which has become the gold standard for gene expression research and can provide reliable and reproducible technical support for the implementation of this invention.

[0032] Single-cell RNA sequencing (scRNA-seq) is a revolutionary application of high-throughput sequencing technology at the single-cell level. Its core objective is to comprehensively and quantitatively analyze gene expression profiles at single-cell resolution, thereby revealing cell population heterogeneity, discovering rare cell subpopulations, and tracking dynamic changes in cell state. The basic process includes: first, dissociating the biological sample to be tested (such as CD138+ plasma cells in this invention) into a single-cell suspension, and capturing and separating individual cells using microfluidics, droplets, or well plates; next, within each cell or its corresponding microreaction unit, the cell is lysed, and its mRNA is reverse transcribed using primers carrying cell-specific barcodes and unique molecular identifiers (UMIs) to synthesize cDNA with an identifier; subsequently, the cDNA from all cells is collected, amplified, fragmented, end-repaired, A-tailed, and ligated with sequencing adapters to construct a single-cell library that can be sequenced on a sequencer; then, the library is loaded into a high-throughput sequencer (such as a 10x sequencer). Massive parallel sequencing was performed on a platform jointly developed by Genomics and Illumina, reading hundreds of millions of reads containing cell barcodes and transcript sequences. Finally, using bioinformatics methods, the reads were assigned back to each original cell based on the cell barcodes, and UMI was used to correct amplification bias, accurately counting the transcript counts of each gene in each cell. Ultimately, through dimensionality reduction and clustering analyses, cell types and functional states could be identified with single-cell precision. Those skilled in the art are familiar with the core principles and procedures of scRNA-seq, which has become the gold standard for revealing cellular heterogeneity in complex biological systems, providing unprecedented and profound biological insights for the implementation of this invention.

[0033] The present invention preferably uses single-cell transcriptome sequencing technology in the discovery of PRMM-Sig9, while whole transcriptome sequencing technology is used in the actual application of PRMM-Sig9 as a biomarker because single-cell transcriptome sequencing technology is too expensive and not cost-effective for large clinical cohorts, while whole transcriptome sequencing technology is relatively economical and can be applied to large clinical cohorts.

[0034] In a preferred embodiment, the present invention provides a method for in vitro adjuvant assessment of the prognosis of patients with multiple myeloma, wherein:

[0035] In step (3), the risk model formula is: Risk score = (0.225 × MCM3 expression level) + (0.160 × FABP5 expression level) + (0.150 × MIR4435-2HG expression level) + (0.122 × FH expression level) + (0.072 × HELLS expression level) + (0.058 × CCDC34 expression level) + (0.036 × RPS6KA1 expression level) + (0.013 × FEN1 expression level) + (0.011 × OPN3 expression level), where the expression level of each gene is the standardized expression abundance value;

[0036] In step (4), the preset threshold is 4.8949.

[0037] In a preferred embodiment, the present invention provides a method for in vitro auxiliary assessment of the prognosis of patients with multiple myeloma, wherein in step (4), the risk score above the preset threshold indicating a significantly increased risk of death means that, based on the results of LASSOCox regression analysis, at p<0.001, at any time point after the patient is diagnosed with multiple myeloma, the instantaneous risk of death of the patient is 2.4-3.9 times that of the instantaneous risk of death when the patient's risk score is less than or equal to the preset threshold.

[0038] In a preferred embodiment, the present invention provides a method for in vitro auxiliary assessment of the prognosis of patients with multiple myeloma, wherein in step (4), the risk score is compared with a preset threshold to assess the risk of the multiple myeloma patient having primary refractory multiple myeloma, wherein a risk score higher than the preset threshold indicates that the patient is at high risk of primary refractory multiple myeloma.

[0039] The beneficial effect of this invention is that, by using the genomic collaboration of this invention as a biomarker for detection, the prognosis of multiple myeloma patients, especially the risk of death, and the risk of multiple myeloma patients having primary refractory multiple myeloma can be assessed with specificity and high precision based on the detection results.

[0040] Existing technologies have limited accuracy in differentiating PRMM, posing a significant technical challenge. To address this challenge, this invention analyzes molecular differences at the single-cell level, discovering a genomic synergy that can specifically and accurately identify PRMM as a biomarker. This gene synergy can accurately identify high-risk PRMM patients and predict their sensitivity to standard therapies, thus solving the problem of existing technologies' inability to effectively differentiate and guide the treatment of PRMM patients.

[0041] The beneficial effects of this invention are specifically reflected in:

[0042] (1) Advanced and precise: The biomarkers of this invention are obtained from the deep analysis of single-cell sequencing technology. Through "digital cell sorting", the unique molecular characteristics of PRMM malignant plasma cells are locked in the most faithful state, overcoming the problem of microenvironment signal interference in traditional Bulk RNA sequencing, making the basis for biomarker discovery more reliable and precise.

[0043] (2) Strong predictive ability and clinical consistency: The PRMM-Sig9 model of this invention has been validated in consistency in three large, independent clinical cohorts. All samples in these cohorts were enriched with CD138+ cells, representing the gold standard for clinical diagnosis, ensuring that the model validation process is highly consistent with clinical practice. The model can stably and accurately classify patients into high-risk and low-risk groups, and the risk of death is significantly increased in the high-risk group.

[0044] (3) Clear and reproducible methods: This invention clarifies the technical path based on high-throughput sequencing, including specific gene expression quantification and standardized methods (e.g., log2(TPM+1)), as well as a precise risk scoring formula and a fixed diagnostic threshold (4.8949). This enables any laboratory with a standard sequencing platform to reproduce the assessment results without creative effort based on the teachings of this invention, laying a solid foundation for clinical translation.

[0045] (4) Clear clinical guidance significance: This invention can identify high-risk patients for PRMM at the new diagnosis stage. Since PRMM is defined as ineffectiveness of initial treatment, this is directly equivalent to a very poor prognosis and a higher risk of death. Therefore, this invention achieves effective assessment of PRMM risk by predicting mortality risk, and can provide clinicians with a strong basis for decision-making in risk stratification, considering alternative treatment options, or recommending clinical trials before treatment. Attached Figure Description

[0046] Figure 1 This refers to the technical approach adopted in this invention and the results of single-cell atlas analysis of bone marrow samples from multiple myeloma patients, wherein:

[0047] Figure 1A is a schematic diagram of the technical route of this invention, illustrating the complete process from obtaining bone marrow samples from patients with newly diagnosed (NDMM, i.e., Newly Diagnosed Multiple Myeloma), primary refractory (PRMM, i.e., Primary Refractory Multiple Myeloma), and relapsed / refractory (RRMM, i.e., Relapsed / Refractory Multiple Myeloma), through single-cell transcriptome sequencing (scRNA-seq) and a series of bioinformatics analyses, to finally identifying the PRMM-specific gene marker (i.e., PRMM-Sig9 of this invention), and validating its prognostic value and predicting drug sensitivity. Specifically, it includes:

[0048] The left side (Patient Cohorts and scRNA-seq Platform) indicates the data source and core technology. Bone marrow samples were obtained from three patient cohorts: newly diagnosed NDMM, primary refractory MM, and relapsed / refractory MM. High-quality single-cell gene expression data were obtained through the single-cell transcriptome sequencing platform (scRNA-seq Platform).

[0049] The middle section (Data preprocessing and Comparative analysis) represents the data preprocessing and comparative analysis process, including data quality control, standardization, and integration to eliminate batch effects, followed by unsupervised clustering and cell type annotation. The core step is to perform comparative analysis of differentially expressed genes to identify gene expression characteristics specific to the PRMM group.

[0050] The right-hand side (In-Depth Analysis) represents in-depth analysis based on comparative analysis results. This includes plasma cell subclustering, analysis of cell-cell communications, construction of pseudotime trajectories to explore cell evolution, analysis of copy number variation (CNV) and in-depth exploration of the bone marrow microenvironment. By constructing PRMM-Sig9, this model effectively predicted the prognosis of myeloma patients and the risk of PRMM.

[0051] Figure 1B is a UMAP dimensionality-reduced clustering diagram after integrating single-cell sequencing data from all samples. Each point in the diagram represents a single cell. The left diagram annotates cells into 11 major cell types and colors them according to known marker genes, such as "Plasma cells" representing plasma cells and "CD8+T cells" representing CD8+T cells, demonstrating the cellular heterogeneity of the bone marrow microenvironment. The right diagram is colored according to the unsupervised clustering results, showing the natural clustering status of cells at the transcriptome level.

[0052] Figure 1 C shows the distribution of cells from different clinical states (NDMM, PRMM, RRMM) in UMAP space. Different colors in the figure represent patients in different clinical groups. It can be seen that the cell distribution of each group overlaps, but there are also their own enriched areas, suggesting differences in their transcriptomes.

[0053] Figure 1 D represents the quality control and cell composition of the samples. The bar chart in the top panel shows the total number of cells retained in each of the 12 samples after data quality control, demonstrating sufficient data volume. The stacked bar chart in the bottom panel shows the relative proportions of the 11 major cell types in each sample, proving the comparability of cell composition across different samples and the reliability of data collection quality.

[0054] Figure 2 The figure shows the results of differentially expressed genes in malignant plasma cells from patients with different clinical subtypes of multiple myeloma.

[0055] Figure 2 A and Figure 2 B represents Venn diagrams and Upset diagrams, used to illustrate the number, overlap, and specificity of differentially expressed genes (DEGs) in malignant plasma cells across three key comparisons (PRMM vs NDMM, RRMM vs NDMM, and RRMM vs PRMM). These diagrams demonstrate that PRMM possesses a unique gene expression profile that distinguishes it from NDMM and RRMM.

[0056] Figure 2 C represents the volcano plots for the three comparison groups. In each volcano plot, the horizontal axis (log2(Fold Change)) represents the fold change in gene expression levels, with positive values ​​indicating upregulation and negative values ​​indicating downregulation. The vertical axis (-log10(p-value)) represents the statistical significance of the difference, with higher values ​​indicating greater significance. Each point in the plot represents a gene, used to visually screen for genes that are specifically highly or poorly expressed in PRMM.

[0057] Figure 2D is a heatmap of key differentially expressed genes. Each row in the graph represents a gene, and each column represents a patient sample. The color from blue to red represents gene expression levels from low to high. This graph visually demonstrates that the PRMM group (the middle part of the graph) has a set of highly expressed genes that are significantly different from the patterns of the other two groups (NDMM and RRMM), providing direct evidence for subsequent biomarker screening.

[0058] Figure 3 The diagram shows the construction process of the core 9-gene prognostic biomarker (PRMM-Sig9) of this invention and its prognostic validation results in multiple cohorts, wherein:

[0059] Figure 3 A demonstrates the process of screening candidate prognostic risk genes that are significantly associated with overall survival (OS) or mortality risk from genes upregulated by PRMM using univariate Cox regression analysis.

[0060] Figure 3 B represents the final nine genes (MCM3, FABP5, MIR4435-2HG, FH, HELLS, CCDC34, RPS6KA1, FEN1, OPN3) determined by LASSOCox regression analysis, and their respective regression coefficients. These coefficients were used to construct the risk scoring model of this invention.

[0061] Figure 3 C represents the Kaplan-Meier survival curves, used to validate the model's prognostic discriminative ability. The horizontal axis represents time (days), and the vertical axis represents the overall survival probability. In three independent patient cohorts (left: MMRF training cohort; middle: GSE24080 validation cohort; right: GSE2658 validation cohort), patients were divided into high-risk (red curve) and low-risk (blue curve) groups based on the median risk score of this invention. P-values ​​in the figures were calculated using the log-rank test. The results clearly show that in all cohorts, the survival probability of patients in the high-risk group was significantly lower than that in the low-risk group, demonstrating the powerful and reproducible prognostic predictive ability of the biomarker of this invention.

[0062] Figure 4 The graph shows the results of validating the PRMM-Sig9 risk score as an independent prognostic factor and evaluating its predictive accuracy.

[0063] Figure 4A is a forest plot of multivariate Cox regression analysis. Each row in the plot represents a clinical variable or the risk score of this invention. The squares represent the hazard ratio (HR), and the horizontal lines represent the 95% confidence interval (CI). When the confidence interval does not cross the vertical dashed line 1, it indicates that the factor is an independent influencing factor. The results clearly show that after adjusting for age, gender, and ISS stage, the HR value of the risk score of this invention is much greater than 1, and the 95% CI is completely to the right of 1 (p<0.001), proving that it is an independent and powerful risk factor for predicting overall survival or mortality risk.

[0064] Figure 4 B is the time-dependent ROC curve, used to evaluate the model's predictive accuracy. The horizontal axis represents 1 - specificity, and the vertical axis represents sensitivity. The figure shows the predictive power of the PRMM-Sig9 model for 1-year, 3-year, and 5-year overall survival in three cohorts. The area under the curve (AUC) is a measure of accuracy; a value closer to 1 indicates a better model. The AUC values ​​in the figure are all above 0.6, indicating that the model has good and stable predictive accuracy.

[0065] Figure 5 This is a graph showing the results of a drug sensitivity prediction analysis based on PRMM-Sig9 risk stratification, where:

[0066] Figure 5 A and Figure 5 C represents the volcano plots showing the predicted drug sensitivity (represented by differences in IC50 values) between the high-risk and low-risk groups in the GSE24080 and GSE2658 cohorts. The horizontal axis represents the logarithm of the IC50 difference, and the vertical axis represents the significance of the difference. This plot is used to globally display drugs for which there are significant differences in sensitivity between the two groups.

[0067] Figure 5 B and Figure 5 D is a box plot comparing the predicted sensitivity of core treatment drugs for multiple myeloma. The horizontal axis represents the drug name, and the vertical axis represents the predicted half-maximal inhibitory concentration (IC50 value). A higher IC50 value indicates that the cells are less sensitive to the drug (i.e., more resistant). The results consistently showed in two independent cohorts that the high-risk group defined by the PRMM-Sig9 of this invention had significantly higher predicted IC50 values ​​for proteasome inhibitors (bortezomib, MLN2238 / ixazomib) and immunomodulators (thalidomide) than the low-risk group. This result strongly demonstrates that the biomarkers of this invention can effectively identify PRMM patients who are naturally resistant to first-line standard treatment regimens. Detailed Implementation

[0068] To better understand the technical solutions and advantages of the present invention, the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0069] Example 1: Discovery and Identification of PRMM-Specific Gene Markers

[0070] This embodiment details the complete discovery process of the core PRMM-Sig9 marker of the present invention.

[0071] 1. Data Acquisition and Processing

[0072] This study used publicly available single-cell transcriptome sequencing (scRNA-seq) data. Specifically, bone marrow mononuclear cell (BMMC) data were obtained from the NCBI GEO database from four newly diagnosed MM (NDMM) patients and four primary refractory MM (PRMM) patients with accession number GSE189460, and from four relapsed / refractory MM (RRMM) patients with accession number GSE161801. The technical approach of this study is attached. Figure 1 As shown in Figure A.

[0073] 2. Quality control and integration of scRNA-seq data

[0074] The raw data were processed using the Seurat software package (v5.3.0) in the R language environment. Quality control criteria included removing low-quality cells with fewer than 200 expressed genes or a mitochondrial gene expression rate higher than 20%, and removing genes with fewer than 3 expressed genes in the total number of cells. Data were standardized using the LogNormalize method. To eliminate technical differences (i.e., batch effects) between different patient samples, the FindIntegrationAnchors and IntegrateData functions of Seurat were used to integrate the data from 12 patients. Principal component analysis (PCA) was then performed, and dimensionality reduction visualization was achieved using the UMAP (Uniform Manifold Approximation and Projection) algorithm. The quality control results of this step and the integrated cell composition distribution are shown in the attached figure. Figure 1 As shown in D.

[0075] 3. Cell type “numerical sorting” and annotation

[0076] In the UMAP dimensionality reduction space, cell populations generated by unsupervised clustering are annotated using canonical marker genes of known cell types. For example, cells highly expressing SDC1 (CD138) and IGHG1 are annotated as malignant plasma cells; cells expressing CD8A and CD3E are annotated as CD8+ T cells, etc. The results of this step are attached. Figure 1 As shown in B, the figure illustrates the 11 main cell types as annotated. (Appendix) Figure 1 C shows the origin distribution of cells in the UMAP diagram for each clinical subtype (NDMM, PRMM, RRMM).

[0077] 4. Screening for differentially expressed genes

[0078] The malignant plasma cell subsets identified in the previous step's "digital sorting" were extracted. The gene expression profiles of the PRMM group versus the NDMM group, the RRMM group versus the NDMM group, and the RRMM group versus the PRMM group were compared in pairs using the FindMarkers function in the Seurat package and the Wilcoxon rank-sum test. The statistical criteria for screening differentially expressed genes (DEGs) were: a corrected p-value < 0.05 and |log2(fold change)| > 0.25. The results of this step are attached. Figure 2 As shown, the attached Figure 2 A and 2B show the intersection of DEG among the groups, with appendix. Figure 2 C's volcano map and appendix Figure 2 The heatmap of D visually illustrates the differential expression patterns of these genes.

[0079] 5. Construction of prognostic biomarker models

[0080] (a) Candidate Gene Screening: The DEGs obtained in step 4, which were significantly upregulated in PRMM malignant plasma cells, were integrated with a large-sample public cohort of 782 patients—the MMRF CoMMpass study (this cohort is an internationally recognized database of high-quality prospective MM studies. This invention screened 782 newly diagnosed patients with high-quality Bulk RNA sequencing data and complete clinical overall survival (OS) information from this cohort). In this cohort, univariate Cox proportional hazards regression analysis was performed on these candidate genes (for details on the Cox proportional hazards regression analysis, see: Liu C, Wang X, Genchev GZ, Lu H. Multi-omics facilitated variable selection in Cox-regression model for cancer prognosis prediction. Methods 2017; 124:100-107.). Genes that were both upregulated in PRMM and significantly associated with poor overall survival (OS) (hazard ratio HR>1, p<0.05) were screened, ultimately yielding 37 candidate prognostic risk genes. Part of the screening results are detailed in the appendix. Figure 3 A is illustrated in diagram A.

[0081] (b) Final Model Establishment: In the MMRF training cohort, LASSO (Least Absolute Shrinkage and Selection Operator) Cox regression analysis was applied to these 37 candidate genes (for details on LASSO Cox regression analysis, see: McEligot AJ, Poynor V, Sharma R, Panangadan A. Logistic LASSO Regression for Dietary Intakes and Breast Cancer. Nutrients 2020; 12:2652.). This algorithm, by introducing a penalty term, can compress the coefficients of genes with small contributions to zero, thereby screening out the optimal and most robust variable combination. The optimal penalty parameter lambda was determined through 10-fold cross-validation. Finally, this analysis screened out the optimal prognostic model consisting of 9 genes, namely PRMM-Sig9 of this invention. The 9 genes and their regression coefficients finally screened by this LASSO regression analysis are shown in the attached table. Figure 3 Detailed information is provided in section B.

[0082] Example 2: Construction and validation of the PRMM-Sig9 risk model based on whole transcriptome sequencing technology

[0083] This embodiment details the process of transforming the findings of Example 1 into a robust risk assessment model that can be applied to clinical sequencing data.

[0084] 1. Data sources, composition, and processing methods

[0085] This study used three large, independent, publicly available database cohorts, and the whole transcriptome sequencing data for all cohorts were derived from bone marrow plasma cell samples enriched with CD138+ magnetic beads, ensuring the clinical relevance and consistency of the data sources.

[0086] Training cohort: The Multiple Myeloma Research Foundation (MMRF) CoMMpass study cohort. This cohort is an internationally recognized database of high-quality prospective MM studies. This invention selected 782 newly diagnosed patients with high-quality Bulk RNA sequencing data and complete clinical overall survival (OS) information from this cohort as a dataset for model construction and initial training.

[0087] Validation cohort 1: GSE24080. This cohort was derived from the NCBI GEO database and contained gene expression profiles (Affymetrix U133 Plus 2.0 microarray data) and clinical follow-up data of 559 newly diagnosed MM patients.

[0088] Validation cohort 2: GSE2658. This cohort, also derived from the GEO database, contains gene expression profiles (Affymetrix U133 Plus 2.0 microarray data) and clinical follow-up data from 559 newly diagnosed MM patients.

[0089] The standardization methods for gene expression data are as follows:

[0090] To ensure the repeatability of this invention, the standardized process used in this invention is the mainstream and established method in the field.

[0091] For RNA sequencing data (such as MMRF cohorts): First, the sequencing reads are aligned to the human reference genome (such as GRCh38) using standard bioinformatics alignment software (such as STAR); then, the expression abundance of each gene is calculated using quantitative software (such as RSEM); finally, the expression abundance is uniformly quantified as "transcripts per million" (TPM). TPM is a standardized method that eliminates the influence of gene length and sequencing depth.

[0092] For gene chip data (such as GSE24080 and GSE2658): use standard chip data processing algorithms (such as RMA or MAS5) for background correction, normalization and summarization of probe set signal values.

[0093] For subsequent statistical analysis and model building, expression values ​​from all sources need to be transformed using log2. Specifically, this involves taking the logarithm to base 2 of the "TPM value + 1" or the equivalent chip signal value, i.e., log2(TPM + 1). Adding 1 is to avoid infinitesimal values ​​when taking the logarithm of genes with an expression value of 0.

[0094] 2. Determination of Risk Model Formula and Threshold

[0095] (a) Candidate gene screening: The PRMM upregulated genes found in Example 1 were integrated with the MMRF training cohort data, and genes significantly associated with poor overall survival (OS) were screened out by univariate Cox proportional hazards regression analysis.

[0096] (b) Final Model Establishment: In the MMRF training cohort, LASSOCox regression analysis was applied to the candidate genes. Through 10-fold cross-validation, the optimal prognostic model (i.e., PRMM-Sig9) consisting of 9 genes was finally selected. This analysis also provides the regression coefficients for each gene, thus establishing the risk model formula of this invention:

[0097] Risk score = (0.225 × MCM3 expression level) + (0.160 × FABP5 expression level) + (0.150 × MIR4435-2HG expression level) + (0.122 × FH expression level) + (0.072 × HELLS expression level) + (0.058 × CCDC34 expression level) + (0.036 × RPS6KA1 expression level) + (0.013 × FEN1 expression level) + (0.011 × OPN3 expression level)

[0098] The "expression level" of each gene is the value after TPM standardization and log2(x+1) transformation.

[0099] Determination of diagnostic threshold: In the MMRF training queue, the optimal risk score threshold for distinguishing between high and low risk groups by calculating the median risk score is 4.8949.

[0100] 3. Validation of the model's prognostic value

[0101] The fully defined model (containing 9 genes, fixed formula coefficients, and a unique threshold of 4.8949) was applied unchanged to the MMRF training queue and two independent validation queues, GSE24080 and GSE2658.

[0102] Prognostic stratification capability verification: as attached Figure 3As shown in C, in all three cohorts, patients with a risk score >4.8949 were defined as the high-risk group, and patients with a risk score ≤4.8949 were defined as the low-risk group. The transient risk of death for patients in the high-risk group was 2.4–3.9 times that for patients in the low-risk group (HR = 3.1, p < 0.001), demonstrating the model's strong and reproducible prognostic stratification ability.

[0103] Independent prognostic value validation: as attached Figure 4 As shown in Figure A, multivariate Cox regression analysis performed in the MMRF cohort indicated that, after adjusting for clinically recognized prognostic factors such as age and ISS stage, the PRMM-Sig9 risk score remained a potent independent prognostic risk factor (HR = 3.1, p < 0.001).

[0104] Model prediction accuracy assessment: as attached Figure 4 As shown in Figure B, the time-dependent ROC curve analysis shows that the model demonstrates good and stable accuracy in predicting 1-year, 3-year, and 5-year survival rates in all three cohorts (AUC values ​​are all above 0.6), proving the reliability of the model.

[0105] Example 3: Validation of the association between the PRMM-Sig9 model and drug sensitivity (computerized drug sensitivity prediction)

[0106] This embodiment aims to verify whether the "high-risk group" defined in this invention truly embodies the core characteristics of PRMM in biology, namely, natural resistance to first-line drugs.

[0107] 1. Verification Method

[0108] This study employs a widely accepted computer-based drug sensitivity prediction algorithm, such as the oncoPredict package in R. This method utilizes the relationship between gene expression profiles and drug sensitivity (IC50 value) established in the large drug screening database (The Cancer Therapeutics Response Portal, CTRP) to predict the sensitivity of new samples (patients in the GSE24080 and GSE2658 cohorts in this example) to a specific drug. A higher IC50 value indicates that the cells are less sensitive to the drug (i.e., more resistant).

[0109] 2. Verification process

[0110] Patients in the GSE24080 and GSE2658 cohorts were divided into high-risk and low-risk groups based on the risk score and threshold of 4.8949 calculated in Example 2. Then, oncoPredict was used to predict the IC50 values ​​of each patient for multiple drugs.

[0111] 3. Verification Results

[0112] As attached Figure 5 As shown, the results are very clear: in the two independent validation cohorts GSE24080 and GSE2658, the high-risk group, as defined by the model of this invention, had significantly higher predicted IC50 values ​​for the proteasome inhibitor bortezomib compared to the low-risk group (p<0.01).

[0113] 4. Conclusion

[0114] These results strongly confirm that the PRMM-Sig9 model of this invention can not only predict poor survival outcomes, but also that the high-risk population it identifies is biologically exhibiting predictive resistance to standard first-line treatments for MM. This perfectly links the prognostic value of the model with the clinical definition of PRMM, demonstrating the great potential of this invention in identifying PRMM patients.

[0115] A reasonable explanation for the relationship between "mortality risk" and "PRMM risk"

[0116] A core aspect of this invention lies in assessing the risk of primary refractory multiple myeloma (PRMM) by predicting mortality risk. The underlying logic and rationale are as follows:

[0117] The clinical definition of primary refractory multiple myeloma (PRMM) is ineffectiveness to initial treatment. This "de novo resistance" biological characteristic directly leads to extremely poor prognosis and a higher risk of death. The PRMM-Sig9 model of this invention captures the molecular characteristics driving this extremely poor prognosis, thereby accurately identifying high-risk individuals with PRMM biological characteristics before treatment begins. Furthermore, the drug sensitivity prediction results in Example 3 provide direct evidence supporting this logic. Therefore, in this invention, predicting mortality risk combined with drug sensitivity prediction is a direct, effective, and mechanistically validated method for assessing PRMM risk.

[0118] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims and their equivalents, this invention is also intended to include these modifications and variations. The above embodiments or implementations are merely illustrative examples of this invention, and it can also be implemented in other specific ways or forms without departing from its gist or essential characteristics. Therefore, the described embodiments should be considered illustrative rather than limiting in any respect. The scope of this invention should be defined by the appended claims, and any changes equivalent to the intent and scope of the claims should also be included within the scope of this invention.

Claims

1. A gene combination, characterized in that: The gene combination is a combination of MCM3, FABP5, MIR4435-2HG, FH, HELLS, CCDC34, RPS6KA1, FEN1 and OPN3 genes.

2. Use of the reagent for detecting the gene combination of claim 1 in the preparation of the reagent, the use of which includes assessing the prognosis of patients with multiple myeloma.

3. The use according to claim 2, characterized in that: The stated uses further include assessing the risk of a patient with multiple myeloma having primary refractory multiple myeloma.

4. The use according to claim 3, characterized in that: The described uses further include assessing the likelihood of a patient having natural resistance to a proteasome inhibitor, such as bortezomib.