Myeloid tumor prediction or diagnosis system and application

Raman spectroscopy technology screens biomarkers in serum, establishes a myeloid tumor prediction or diagnosis system, solves the problem of early diagnosis of myeloid tumors, and achieves rapid and non-invasive myeloid tumors and their subtypes, reducing detection costs and providing new treatment ideas.

CN120376103APending Publication Date: 2025-07-25INST OF HEMATOLOGY & BLOOD DISEASES HOSPITAL CHINESE ACADEMY OF MEDICAL SCI & PEKING UNION MEDICAL COLLEGE +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510457400.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The prior art is difficult to quickly and non-invasively diagnose myeloid tumors and their subtypes. The traditional methods are costly and highly invasive, and lack effective serological indicators for early identification of MPN, MDS/MPN and AML.

Method used

Raman spectroscopy combined with multivariate analysis was used to detect the Raman peak intensity of representative nucleic acids, proteins, lipids, β-carotene and collagen in serum, and biomarkers were screened out through the OPLS-DA model to establish a myeloid tumor prediction or diagnosis system.

Benefits of technology

It has achieved rapid and non-invasive prediction or diagnosis of myeloid tumors and their subtypes, reduced detection costs, improved diagnostic efficiency, and provided a new strategy for clinical treatment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120376103A_ABST
    Figure CN120376103A_ABST
Patent Text Reader

Abstract

The invention provides a prediction or diagnosis system for myeloid tumors and application, and belongs to the technical field of tumor detection. The research finds that the intensity levels of Raman peaks 786 and 1579 cm <-1 > representing nucleic acid, Raman peaks 643, 759, 103, 126, 1603 and 1616 cm <-1 > representing protein, Raman peaks 1437, 1443 and 1446 cm <-1 > representing lipid, Raman peak 957 cm <-1 > representing beta-carotene and Raman peak 1345 cm <-1 > representing collagen in serum are detected as biomarkers, and the serum can be used for predicting or diagnosing myeloid tumors and subtypes thereof. The myeloid tumor and the subtype thereof can be quickly and noninvasively predicted or diagnosed by utilizing the myeloid tumor prediction or diagnosis system, more choices are provided for quickly predicting or diagnosing the myeloid tumor and the subtype thereof, and the system has profound scientific research and clinical significance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of tumor detection, and particularly relates to a prediction or diagnosis system and application for myeloid neoplasms. Background Art

[0002] Myeloid neoplasms are a group of blood diseases caused by clonal mutations of hematopoietic stem cells, mainly including myelodysplastic syndromes (MDS), myeloproliferative neoplasms (MPN), myelodysplastic / myeloproliferative neoplasms (MDS / MPN), and acute myeloid leukemia (AML), etc. MPN is a group of clonal hematological malignancies caused by specific gene mutations, including chronic neutrophilic leukemia (CNL), primary myelofibrosis (PMF), chronic myeloid leukemia (CML), polycythemia vera (PV), essential thrombocythemia (ET), etc. Its main feature is the excessive proliferation of one or more myeloid cell lines in the bone marrow, resulting in a significant increase in the white blood cell, red blood cell, or platelet count in the peripheral blood.

[0003] MDS / MPN is a group of rare hematological malignancies with characteristics of both MDS and MPN. It has the characteristics of ineffective hematopoiesis and cytopenia in MDS, as well as the manifestations of excessive hematopoiesis of blood cells and organ enlargement in MPN. The diagnosis of MDS / MPN usually relies on peripheral blood differential count, cell morphology, flow cytometry, immunohistochemistry, and related gene detection. MDS / MPN mainly includes chronic myelomonocytic leukemia (CMML), etc. The current treatment is mainly guided by risk stratification and the specific clinical needs of patients, including symptomatic support, immunomodulatory therapy, demethylating drugs, and allogeneic hematopoietic stem cell transplantation.

[0004] AML is a malignant hematological disease caused by the clonal expansion of malignant hematopoietic precursor cells in the bone marrow. The occurrence of AML is closely related to gene mutations, abnormal signaling pathways, and immune imbalance. Common symptoms in AML patients include fatigue, shortness of breath, easy bruising and bleeding, and an increased risk of infection. In addition, due to impaired normal hematopoiesis, patients also experience anemia and thrombocytopenia. The treatment of AML usually includes high-dose chemotherapy, targeted therapy, and allogeneic hematopoietic stem cell transplantation, etc.

[0005] The early diagnosis of MPN, MDS / MPN, and AML is difficult. Therefore, exploring a diagnostic method that is inexpensive and helps to quickly identify MPN, MDS / MPN, and AML is of great significance in the diagnosis research of hematological diseases. In clinical practice, the traditional diagnostic methods for MPN, MDS / MPN, and AML require a long time, high economic costs, and are invasive. There is no "gold standard" for the diagnosis of some disease types. Moreover, the association between serological indicators related to glycolipid metabolism and MPN, MDS / MPN, and AML lacks sufficient evidence and fails to provide sufficient information for the rapid early differentiation of MPN, MDS / MPN, and AML.

[0006] In the field of scientific research, Raman spectroscopy has been used as a rapid, non-invasive, and label-free method to analyze and classify cells with position specificity. The advantages of using Raman spectroscopy for diagnosis have been proven: it can potentially replace the traditional diagnostic method based on visual inspection of histological samples. However, there is currently no report on the comprehensive application of Raman spectroscopy in the serum analysis of MPN, MDS / MPN, and AML, nor is there an in-depth discussion on combining serological indicators related to glycolipid metabolism in MPN, MDS / MPN, and AML patients. Therefore, developing a rapid identification method for MPN, MDS / MPN, and AML based on Raman spectroscopy without antibody labeling, and deeply exploring the massive serological examination data of MPN, MDS / MPN, and AML patients, is of great value for the early identification of diseases, will also improve the diagnostic efficiency of diseases, reduce the detection cost, and promote the further development of the rapid and accurate diagnosis of MPN, MDS / MPN, and AML. Summary of the Invention

[0007] In view of this, the purpose of the present invention is to provide a prediction or diagnosis system and application for myeloid tumors.

[0008] In order to achieve the above-mentioned invention purpose, the present invention provides the following technical solutions:

[0009] The present invention provides a prediction or diagnosis system for myeloid tumors. The prediction or diagnosis system includes a Raman spectroscopy detection module for a test sample and a Raman spectroscopy detection module for a negative control. Both the Raman spectroscopy detection module for the test sample and the Raman spectroscopy detection module for the negative control use the detection of Raman spectroscopy data of serum as an indicator. The Raman spectroscopy data includes Raman peaks at 786 and 1579 cm -1 representing nucleic acids, Raman peaks at 643, 759, 1031, 1260, 1603, and 1616 cm -1 representing proteins, Raman peaks at 1437, 1443, and 1446 cm -1 representing lipids, a Raman peak at 957 cm -1 representing β-carotene, and a Raman peak at 1345 cm -1 .

[0010] Preferably, the prediction or diagnosis system further includes a conclusion output module. Compared with the Raman spectroscopy data of the negative control, when the peak intensities of the protein Raman peaks, nucleic acid Raman peaks, β-carotene Raman peak, and lipid Raman peaks in the Raman spectroscopy data of the test sample are significantly reduced, and the peak intensity of the collagen Raman peak in the Raman spectroscopy data of the test sample is significantly increased, it is predicted or diagnosed as a myeloid tumor.

[0011] Preferably, the myeloid tumors include one or more of myelodysplastic / myeloproliferative neoplasms, myeloproliferative neoplasms, and acute myeloid leukemia.

[0012] Preferably, the myeloproliferative neoplasms are one or two of chronic neutrophilic leukemia and primary myelofibrosis.

[0013] Preferably, the negative control is healthy human serum.

[0014] The present invention provides the use of the intensity levels of Raman peaks at 786 and 1579 cm -1 representing nucleic acids, Raman peaks at 643, 759, 1031, 1260, 1603, and 1616 cm -1 representing proteins, Raman peaks at 1437, 1443, and 1446 cm -1 representing lipids, a Raman peak at 957 cm -1 representing β-carotene, and a Raman peak at 1345 cm -1 in serum as biomarkers in the preparation of products for predicting or diagnosing myeloid tumors.

[0015] Preferably, the myeloid tumors include one or more of myelodysplastic / myeloproliferative neoplasms, myeloproliferative neoplasms, and acute myeloid leukemia.

[0016] The present invention provides an application of the above-mentioned biomarker in the preparation of a product for predicting or diagnosing myeloproliferative neoplasm subtypes.

[0017] Preferably, the myeloproliferative neoplasm subtype is one or both of chronic neutrophilic leukemia and primary myelofibrosis.

[0018] The present invention provides a computer-readable storage medium for storing computer instructions, programs, code sets or instruction sets, which, when run on a computer, cause the computer to execute the prediction or diagnosis system as described above.

[0019] Compared with the prior art, the present invention has the following beneficial effects:

[0020] The present invention provides a prediction or diagnosis system and application for myeloid neoplasms. The present invention discovers through research that by detecting the intensity levels of Raman peaks 786 and 1579 cm -1 , which represent nucleic acids, Raman peaks 643, 759, 1031, 1260, 1603, and 1616 cm -1 , which represent proteins, Raman peaks 1437, 1443, and 1446 cm -1 , which represent lipids, Raman peak 957 cm -1 that represents β-carotene, and Raman peak 1345 cm -1 that represents collagen in serum as biomarkers, which can be used to predict or diagnose myeloid neoplasms and their subtypes. The present invention can quickly and non-invasively predict or diagnose myeloid neoplasms and their subtypes, providing more options for quickly predicting or diagnosing myeloid neoplasms and their subtypes, and having profound scientific research and clinical significance. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 Are the average serum Raman spectrograms of control group, MPN, MDS / MPN, and AML patients. From bottom to top in A are the average serum Raman spectrograms of Control, MPN, MDS / MPN, and AML respectively; from bottom to top in B are the average serum Raman spectrograms of Control, CNL, and PMF respectively.

[0022] Figure 2In A, it is the permutation plot of the Control, MPN, MDS / MPN, and AML groups analyzed by OPLS-DA; in B, it is the clustering analysis plot of the Control, MPN, MDS / MPN, and AML groups; in C, it is the ROC curve plot of the Control, MPN, MDS / MPN, and AML groups, AUC(control)=1, AUC(MPN)=1, AUC(MDS / MPN)=1, AUC(AML)=1; in D, it is the permutation plot of the Control and MPN groups; in E, it is the clustering analysis plot of the Control and MPN groups; in F, it is the ROC curve plot of the Control and MPN groups, AUC(control)=1, AUC(MPN)=1; in G, it is the permutation plot of the Control and MDS / MPN groups; in H, it is the clustering analysis plot of the Control and MDS / MPN groups; in I, it is the ROC curve plot of the Control and MDS / MPN groups, AUC(control)=1, AUC(MDS / MPN)=1; in J, it is the permutation plot of the Control and AML groups; in K, it is the clustering analysis plot of the Control and AML groups; in L, it is the ROC curve plot of the Control and AML groups, AUC(control)=1, AUC(AML)=1;

[0023] Figure 3 For the OPLS-DA discrimination of the Control and myeloid tumor serum samples, in A, it is the discrimination score plot of the Control, MPN, MDS / MPN, and AML by OPLS-DA drawn with Hotelling's 95% confidence ellipse; in B, it is the discrimination loading line plot of the Control, MPN, MDS / MPN, and AML; in C, it is the discrimination V+S plot of the Control, MPN, MDS / MPN, and AML; in D, it is the discrimination score plot of the Control and MPN by OPLS-DA drawn with Hotelling's 95% confidence ellipse; in E, it is the discrimination loading line plot of the Control and MPN; in F, it is the discrimination V+S plot of the Control and MPN; in G, it is the discrimination score plot of the Control and MDS / MPN by OPLS-DA drawn with Hotelling's 95% confidence ellipse; in H, it is the discrimination loading line plot of the Control and MDS / MPN; in I, it is the discrimination V+S plot of the Control and MDS / MPN; in J, it is the discrimination score plot of the Control and AML by OPLS-DA drawn with Hotelling's 95% confidence ellipse; in K, it is the discrimination loading line plot of the Control and AML; in L, it is the discrimination V+S plot of the Control and AML;

[0024] Figure 4In A, it is the permutation plot of the Control, CNL, and PMF groups analyzed by OPLS-DA; in B, it is the cluster analysis plot of the Control, CNL, and PMF groups; in C, it is the ROC curve plot of the Control, CNL, and PMF groups, AUC(control)=1, AUC(CNL)=1, AUC(PMF)=1; in D, it is the permutation plot of the Control and CNL groups; in E, it is the cluster analysis plot of the Control and CNL groups; in F, it is the ROC curve plot of the Control and CNL groups, AUC(control)=1, AUC(CNL)=1; in G, it is the permutation plot of the Control and PMF groups; in H, it is the cluster analysis plot of the Control and PMF groups; in I, it is the ROC curve plot of the Control and PMF groups, AUC(control)=1, AUC(PMF)=1;

[0025] Figure 5 For the OPLS-DA discrimination of the serum samples of Control and MPN subtypes, in A, it is the discrimination score plot of Control, CNL, and PMF by OPLS-DA drawn with the 95% confidence ellipse of Hotelling; in B, it is the discrimination loading line plot of Control, CNL, and PMF; in C, it is the discrimination V+S plot of Control, CNL, and PMF; in D, it is the discrimination score plot of Control and CNL by OPLS-DA drawn with the 95% confidence ellipse of Hotelling; in E, it is the discrimination loading line plot of Control and CNL; in F, it is the discrimination V+S plot of Control and CNL; in G, it is the discrimination score plot of Control and PMF by OPLS-DA drawn with the 95% confidence ellipse of Hotelling; in H, it is the discrimination loading line plot of Control and PMF; in I, it is the discrimination V+S plot of Control and PMF;

[0026] Figure 6 For the statistics of potential biomarkers and serum biochemical results of Control vs MPN vs MDS / MPN vs AML, in A, it is the potential biomarkers of Control vs MPN vs MDS / MPN vs AML; in B, it is the TP results of the Control, MPN, MDS / MPN, and AML groups; in C, it is the LDL results of the Control, MPN, MDS / MPN, and AML; in D, it is the ADA results of the Control, MPN, MDS / MPN, and AML;

[0027] Figure 7For the statistical results of potential biomarkers and serum biochemical results of Control vs MPN subtypes, A shows potential biomarkers of Control vs MPN subtypes; B shows TP results of Control, CNL, and PMF; C shows LDL results of Control, CNL, and PMF; D shows ADA results of Control, CNL, and PMF. Detailed implementation manners

[0028] The present invention combines Raman spectroscopy with multivariate analysis to establish a prediction or diagnosis system for myeloid tumors to distinguish Control, MPN, MDS / MPN, AML, and MPN subtypes. In addition, the present invention also screens Raman peak positions that make major contributions to disease classification as biomarkers for predicting or diagnosing MPN, MDS / MPN, and AML, laying a foundation for better using clinical serological test data for rapid early discrimination of MPN, MDS / MPN, and AML.

[0029] The present invention provides a prediction or diagnosis system for myeloid tumors. The prediction or diagnosis system includes a Raman spectroscopy detection module for a test sample and a Raman spectroscopy detection module for a negative control. Both the Raman spectroscopy detection module for the test sample and the Raman spectroscopy detection module for the negative control use the detection of Raman spectroscopy data of serum as an index. The Raman spectroscopy data are Raman peaks 786 and 1579 cm representing nucleic acids -1 , Raman peaks 643, 759, 1031, 1260, 1603, and 1616 cm representing proteins -1 , Raman peaks 1437, 1443, and 1446 cm representing lipids -1 , Raman peak 957 cm representing β-carotene -1 and Raman peak 1345 cm representing collagen -1 .

[0030] In the above prediction or diagnosis system, the negative control is healthy human serum. The prediction or diagnosis system further includes a data processing device, which includes a data input module, a data recording module, a data comparison module, and a conclusion output module. The data input module is configured to input the Raman spectral data of the sample to be tested and the negative control; the data recording module is configured to store the Raman spectral data of the sample to be tested and the negative control; the data comparison module is configured to receive the Raman spectral data of the sample to be tested and the negative control sent by the data input module, and then compare the Raman spectral data. The conclusion output module is configured to receive the comparison result sent by the data comparison module and determine the comparison result according to a predetermined determination condition. The determination condition is that compared with the Raman spectral data of the negative control, when the peak intensities of the protein Raman peak, nucleic acid Raman peak, β-carotene Raman peak, and lipid Raman peak in the Raman spectral data of the sample to be tested are significantly reduced, and the peak intensity of the collagen Raman peak in the Raman spectral data of the sample to be tested is significantly increased, it is predicted or diagnosed as myeloid neoplasm.

[0031] Through the loading plots of the three models of Control vs MPN, Control vs MDS / MPN, and Control vs AML, the present invention found that compared with the control group, the representative nucleic acids (786, 1579 cm -1 ) of patients with myeloid neoplasms, proteins (643, 759, 1031, 1260, 1603, 1616 cm -1 ), lipids (1437, 1443, 1446 cm -1 ), and β-carotene (957 cm -1 ) had significantly reduced Raman peak intensities, and the representative collagen (1345 cm -1 ) had a significantly increased Raman peak intensity.

[0032] In the present invention, the myeloid neoplasm includes one or more of myelodysplastic / myeloproliferative neoplasms, myeloproliferative neoplasms, and acute myeloid leukemia. Further, the myeloproliferative neoplasm is one or both of chronic neutrophilic leukemia and primary myelofibrosis. The negative control is healthy human serum. The present invention has no special limitation on the preparation method of the serum, and any known preparation method in the art can be used.

[0033] The present invention compares the Raman spectral data of the serum by the OPLS-DA statistical method, and screens out the Raman spectral data of the serum of the MPN, MDS / MPN, and AML patient groups and the control group as biomarkers. The biomarkers are the Raman peaks 786 and 1579 cm representing nucleic acids in the detected serum -1, the Raman peaks 643, 759, 1031, 1260, 1603, 1616 cm representing proteins -1 , the Raman peaks 1437, 1443, 1446 cm representing lipids -1 , the Raman peak 957 cm representing β-carotene -1 and the Raman peak 1345 cm representing collagen -1 of the intensity levels.

[0034] The present invention provides the use of the intensity levels of the Raman peaks 786, 1579 cm representing nucleic acids, -1 , the Raman peaks 643, 759, 1031, 1260, 1603, 1616 cm representing proteins -1 , the Raman peaks 1437, 1443, 1446 cm representing lipids -1 , the Raman peak 957 cm representing β-carotene -1 and the Raman peak 1345 cm representing collagen -1 in the preparation of products for predicting or diagnosing myeloid neoplasms as biomarkers.

[0035] In the above application, the myeloid neoplasms include one or more of myelodysplastic / myeloproliferative neoplasms, myeloproliferative neoplasms, and acute myeloid leukemia.

[0036] The present invention provides the use of the above biomarkers in the preparation of products for predicting or diagnosing subtypes of myeloproliferative neoplasms.

[0037] In the above application, the subtypes of myeloproliferative neoplasms are one or two of chronic neutrophilic leukemia and primary myelofibrosis.

[0038] In the above application, the products include kits or reagents. The present invention can accurately diagnose myeloid neoplasms and their subtypes by OPLS-DA using the above Raman spectral data, and can also provide new ideas and strategies for the treatment of diseases, and ultimately is expected to improve the treatment effect and quality of life of patients.

[0039] The present invention also provides a computer-readable storage medium for storing computer instructions, programs, code sets or instruction sets, which when run on a computer, cause the computer to execute the above prediction or diagnosis system.

[0040] In the present invention, unless otherwise specified, all raw material components are commercially available products well known to those skilled in the art.

[0041] The technical solutions provided by the present invention are described in detail below in conjunction with embodiments, but they should not be construed as limiting the protection scope of the present invention.

[0042] Example 1

[0043] 1. Sample collection

[0044] Fifteen patients who visited the Institute of Hematology & Blood Diseases Hospital, Chinese Academy of Medical Sciences in 2024 were recruited, including 7 males and 8 females, with the youngest age of 8 years old and the oldest age of 71 years old. They were divided into three groups: among them, there were 8 MPN patients, including 4 CNL patients and 4 PMF patients; 4 MDS / MPN patients; 3 AML patients; and 9 healthy control subjects.

[0045] All cases in the experimental group were diagnosed by experienced hematology experts through examinations such as blood phase, bone marrow phase, cytogenetics, immunophenotype analysis, and gene analysis. This study was approved by the Ethics Committee of the Institute of Hematology & Blood Diseases Hospital, Chinese Academy of Medical Sciences. The serum samples used in the study were from the remaining samples of clinical tests of MPN, MDS / MPN, and AML patients and the control group. All MPN, MDS / MPN, and AML patients and the control group underwent routine serum biochemical tests. The serum biochemical test data were from the Clinical Testing Center of the Institute of Hematology & Blood Diseases Hospital, Chinese Academy of Medical Sciences.

[0046] 2. General clinical data

[0047] After fasting for 10 hours, serum was collected from the subjects, and total protein (TP), glucose, triglyceride (TG), total cholesterol (TC), high-density lipoprotein (HDL), low-density lipoprotein (LDL), and adenosine deaminase (ADA) in peripheral blood were measured using an automatic biochemical analyzer.

[0048] 3. Raman spectroscopy analysis of peripheral serum

[0049] 5 μL of serum was dropped on a quartz glass slide, and a confocal Raman spectrometer XploRA Raman microscope was used to measure the quantity. 785 nm laser was selected as the excitation light, with an output power of 40 mW. A 40-fold objective lens was selected, and the specimen was fixed on an XYZ three-dimensional platform. During the shooting process, a ×400.75NA Nikon lens was used, and the sample received laser beam irradiation with an output power of 40 mW in a spot size range of approximately 2×2 μm. The single integration time was 250 s, and the integration times were once. The measurement range was 600 - 1800 cm -1 , 5 - 10 sites were measured for each group, and the resolution was 1 cm -1Meanwhile, the Raman spectrum of the quartz glass slide was measured as the background. The Labspec6 software was used for data processing such as smoothing, background removal, and baseline correction. All spectra were intensity-normalized with their respective 1450 cm -1 Raman peak as the internal standard.

[0050] 4. Establishment of the OPLS-DA diagnostic model based on Raman spectral data analysis

[0051] The SIMCA14.1 software was used to perform supervised orthogonal partial least-squares discrimination analysis (OPLS-DA) on the serum Raman spectral data of patients with myeloproliferative neoplasms and the control group. The goodness-of-fit parameters R 2 and Q 2 were used to evaluate the performance of the OPLS model respectively. Under the null hypothesis condition, the model was resampled 200 times by randomly changing the y matrix for model validation. To find the Raman peak positions with statistically significant differences in the classification model as potential biological markers, cluster analysis and V+S analysis were used. Based on the comprehensive consideration of parameters such as the correlation coefficient, loading, and the distance from the center in the V+S plot, the peak positions with Variable Importance (VIP) > 0.5 were selected as potential biological markers. The obtained potential biological markers were subjected to a significance test, and the potential biological markers with P < 0.05 were considered to be statistically significant. The Origin software was used for related data processing. The IBM SPSS Statistics 27.0 was used for statistical analysis, and the Graphpad Prism 9.0 was used to draw statistical-related graphs.

[0052] 5. Statistical analysis

[0053] SPSS 27.0 and Graphpad Prism 9.0 were used for data analysis and graph generation. The statistical spectral data were independent of each other. The normally distributed data were expressed as M ± SE (mean ± standard error of the mean). The one-way analysis of variance (ANOVE) and Bonferroni multiple comparison tests were used for the homogeneous variance group. The Welch’s one-way analysis of variance (Welch’s ANOVE) and Tamhane’s T2 multiple comparison tests were used for the heterogeneous variance group. The non-normally distributed data were expressed as m(Q1-Q3) (median (25-75 percentile)), and the Kruskal Wallis test and Dunn multiple comparison test were used for comparison.

[0054] P < 0.05 is the significance threshold. The collected clinical data were represented by violin plots and analyzed by the above methods. The chi-square test and Fisher's exact test were used for the comparison of frequency data.

[0055] 6. Results

[0056] 6.1 Raman spectroscopic analysis results of sera from MPN, MDS / MPN, and AML patient groups and the control group

[0057] To study the Raman spectral characteristics of sera from MPN, MDS / MPN, and AML patients and the control group, a total of 54, 45, 28, and 15 Raman characteristic spectra of sera from the control, MPN, MDS / MPN, and AML groups were obtained respectively in this study. Among them, the MPN group included 23 spectra of CNL patients and 22 spectra of PMF patients. The control group (control) was sera from healthy individuals.

[0058] Figure 1 Respectively showed the Raman spectra of sera from the control, MPN, MDS / MPN, and AML groups, and the control and MPN subtypes in the range of 600 - 1800 cm -1 . The peak assignments of the relevant Raman spectra of sera were referred to Table 1. Figure 1 A in Figure 1 gave the Raman spectra of the four groups of samples: control, MPN, MDS / MPN, and AML, -1 and B in -1 gave the Raman spectra of control, CNL, and PMF. The spectral patterns presented similar morphologies. The vertical lines in red, blue, green, purple, orange, and black in the figure represented the peak positions related to nucleic acids (726, 781, 1579 cm -1 ), proteins (621, 643, 759, 1603 cm -1 ), β-carotene (957 cm -1 ), lipids (1437, 1446 cm -1 ), carbohydrates (920, 1123 cm -1 ), and collagen (1345 cm -1 ), respectively. It can be seen that it is difficult to distinguish the differences in serum substances between the MPN, MDS / MPN, and AML patient groups and the control group based on the above spectral patterns and peak positions. It is necessary to further screen out the peak positions that can effectively distinguish the control group and the MPN, MDS / MPN, and AML groups as potential biomarkers by combining the classification model established by the OPLS-DA method.

[0059] Table 1 Peak assignments of serum Raman spectra

[0060]

[0061] 6.2 Initial screening of potential biomarkers using the OPLS-DA model

[0062] Control vs MPN vs MDS / MPN vs AML

[0063] To establish a method for differentiating the Control group, MPN group, MDS / MPN group, and AML group based on multi-parameter analysis, all Raman spectral data of the Control group, MPN group, MDS / MPN group, and AML were respectively formed into four groups of data. The supervised orthogonal partial least squares discriminant analysis (OPLS-DA) was applied to the sample data using SIMCA-P software and detailed analysis and comparison were carried out. The effectiveness of the supervised OPLS-DA model established based on multi-parameter analysis can be evaluated by permutation plots ( Figure 2 A, D, G, and J in Figure 2 ), cluster analysis plots ( Figure 2 B, E, H, and K in

[0064] To determine whether the OPLS-DA model holds, permutation was used for the determination of model establishment. Permutation analysis showed that the intercept of Q 2 on the Y-axis was negative, indicating that the OPLS-DA model holds and is not overfitted (see Figure 2 A, D, G, and J in Figure 2 ). The results of cluster analysis showed that the model had a good discrimination effect on each group of samples (see Figure 2 B, E, H, and K in Figure 2 ). On this basis, ROC plots were used to evaluate the authenticity of the discrimination method. The closer the area under the curve (AUC) is to 1, the higher the authenticity of the discrimination method. The ROC curve showed that in the Control vs MPN vs MDS / MPN vs AML model, AUC(control) = 1, AUC(MPN) = 1, AUC(MDS / MPN) = 1, AUC(AML) = 1 (see Figure 2 C in Figure 2in L), indicating a high accuracy of the discriminant analysis results.

[0065] To further study the relationship between the Raman spectral data of the Control group, MPN group, MDS / MPN group, and AML group and hematological malignancies, and to screen for potential biomarkers, the OPLS-DA model was used for further analysis. Figure 3 A in it is the OPLS-DA score plot. In the plot, the four groups of samples are clearly clustered, indicating that the control, MPN, MDS / MPN, and AML groups are distinguished. This discrimination result shows that OPLS-DA can well distinguish the serum spectral data of patients in the control, MPN, MDS / MPN, and AML groups, laying a foundation for analyzing the material characteristics of the four groups. Figure 3 B in it is the OPLS-DA loading plot, which is used to preliminarily screen the Raman peak positions that contribute to the Control vs MPN vs MDS / MPN vs AML discrimination model. The peak position numbers in red, blue, green, purple, and black in the plot are related to nucleic acids, proteins, β-carotene, lipids, and collagen, respectively. There is a correlation between the loading plot and the score plot when representing the relationship of the sample material content, that is, the substances represented by the peak positions on the positive half-axis of the ordinate of the loading plot have relatively higher contents in the group on the positive half-axis of the abscissa of the score plot, and there is a similar corresponding relationship between the negative half-axis of the ordinate of the loading plot and the negative half-axis of the abscissa of the score plot. The characteristic peak positions representing nucleic acids (786, 1579 cm -1 ), proteins (643, 759, 1031, 1260, 1603, 1616 cm -1 ), lipids (1437, 1443, 1446 cm -1 ), β-carotene (957 cm -1 ), and collagen (1345 cm -1 ) play important roles in the discrimination of the four groups of samples, reflecting that the contents of nucleic acids, proteins, lipids, and β-carotene in the control group are higher than those in the myeloid tumor group (see Figure 3 B in it). Figure 3 C in it is the OPLS-DA V+S plot, which integrates indicators such as VIP and correlation coefficients that characterize the contribution degree of peak positions to the classification model. It is used to further screen for potential biomarkers and conduct a significance test on the characteristic peak positions in subsequent analysis to determine the peak positions that can effectively distinguish the Control group and the myeloid tumor group as potential biomarkers. The V+S plot can provide a list of Raman peak positions in descending order of VIP values, and the peak positions with VIP>0.5 and biological significance are screened as potential markers.

[0066] Based on the Control vs MPN vs MDS / MPN vs AML model, the Control group was combined with the samples of the MPN group, MDS / MPN group, and AML group respectively for OPLS-DA analysis (D-L in Figure 3 ). Figure 3 D, G, and J in Figure 3 are the OPLS-DA score plots of the three models of Control vs MPN, Control vs MDS / MPN, and Control vs AML respectively. The two groups of samples in the three plots are located on the positive and negative semi-axes of the X-axis respectively, and the samples are significantly clustered in the scatter plot, indicating that OPLS-DA can better extract the differential information in the spectra, and the established discrimination method can identify the differences in the metabolic components of serum samples. The three models can well distinguish the two groups of samples in the models (see D, G, and J in Figure 3 ). Figure 3 E, H, and K in -1 are the loading plots of the three models of Control vs MPN, Control vs MDS / MPN, and Control vs AML respectively. The results show that the peak intensities representing proteins, β-carotene, nucleic acids, and lipids in the control group are generally higher than those in the myeloid tumor group, and the peak intensity representing collagen is lower than that in the myeloid tumor group (E, H, and K in -1 ). Specifically, compared with the control group, the intensities of the Raman peaks representing nucleic acids (786, 1579 cm -1 ), proteins (643, 759, 1031, 1260, 1603, 1616 cm -1 ), lipids (1437, 1443, 1446 cm -1 ), and β-carotene (957 cm Figure 3 ). F, I, and L in Figure 3 are the V+S plots of the three models of Control vs MPN, Control vs MDS / MPN, and Control vs AML respectively. The V+S plot can provide a list of Raman peak positions in the order of decreasing VIP values, which provides the main basis for determining the potential biomarkers in the control, MPN, MDS / MPN, and AML models (see Figure 3C, F, I, and L in it). Peaks with VIP>0.5 and biological significance were screened as potential biomarkers. In the four models of Control vs MPN vs MDS / MPN vs AML, Control vs MPN, Control vs MDS / MPN, and Control vs AML, relevant parameters such as VIP (VIP>0.5), correlation coefficient, loading, and distance from the center in the V+S plot were comprehensively considered, and important Raman peak positions affecting sample classification were found. During the subsequent biomarker verification process, Raman peak positions with no statistical differences were excluded from the biomarker range.

[0067] 6.3 Control vs MPN subtypes

[0068] Based on the screening of potential biomarkers for MPN, MDS / MPN, and AML, the MPN subtype model was used to analyze the differences in sample substance content and screen potential biomarkers for MPN subtypes. The effectiveness of the supervised OPLS-DA model established based on multi-parameter analysis can be evaluated by permutation plots ( Figure 4 A, D, G in it), cluster analysis plots ( Figure 4 B, E, H in it), and ROC plots ( Figure 4 C, F, I in it). Permutation analysis showed that the intercept of Q 2 on the Y-axis was negative, indicating that the OPLS-DA model was established and not overfitted (see Figure 4 A, D, G in it). The cluster analysis results showed that the model had a good discrimination effect on each group of samples. The ROC curve showed that in the Control vs CNL vs PMF model, AUC(control)=1, AUC(CNL)=1, AUC(PMF)=1 (see Figure 4 C in it), in the control vs CNL model, AUC(control)=1, AUC(CNL)=1 (see Figure 4 F in it), and in the control vs PMF model, AUC(control)=1, AUC(PMF)=1 (see Figure 4 I in it), indicating that the discriminant analysis results had high accuracy. Figure 5In it, A is the OPLS-DA scoreplot. In the figure, the three groups of samples are clearly clustered. The control group is located on the positive half-axis of the X-axis, and the CNL and PMF groups are located on the negative half-axis of the X-axis, reflecting that the control group is distinguished from the MPN subtype groups. This discrimination result indicates that OPLS-DA can well distinguish the serum spectral data of control, CNL, and PMF patients, providing conditions for analyzing the material characteristics of the three groups. Figure 5 In it, B is the OPLS-DA loading plot, which is used to preliminarily screen the Raman peak positions that contribute to the Control vs MPN subtypes discrimination model. The peak position numbers in red, blue, green, purple, and black in the figure are related to nucleic acids, proteins, β-carotene, lipids, and collagen, respectively. Figure 5 In it, C is the OPLS-DA V+S plot, which is used to further screen potential biomarkers, perform a significance test on the characteristic peak positions in subsequent analyses, and determine the peak positions that can effectively distinguish the control group and MPN subtypes as potential biomarkers.

[0069] Based on the Control vs MPN subtypes model, the control group is combined pairwise with the two groups of MPN subtype samples for OPLS-DA analysis ( Figure 5 in D-I of it). Figure 5 In it, D and G are the OPLS-DA scoreplots of the Control vs CNL and Control vs PMF models, respectively. The two groups of samples in the figure are located on the positive and negative half-axes of the X-axis, and the samples in the scatter plot are clearly clustered, reflecting that the two models have good discrimination ability for the two groups of samples in the models. Figure 5 In it, E and H are the loading plots of the Control vs CNL and Control vs PMF models, respectively. The results show that compared with the control group, the intensities of the Raman peak positions representing nucleic acids (786, 1579 cm -1 ), proteins (643, 759, 1031, 1260, 1603, 1616 cm -1 ), lipids (1437, 1443, 1446 cm -1 ), β-carotene (957 cm -1 ) in MPN tumor patients are significantly reduced, and the intensity of the Raman peak position representing collagen (1345 cm -1 ) is significantly increased. Figure 5In which, F and I are respectively the V+S plots of the Control vs CNL and Control vs PMF models. The V+S plot provides the main basis for determining potential biomarkers in the Control and MPN subtypes models ( Figure 5 in C, F, and I). According to the V+S plot, a list of peak positions with VIP>0.5 can be derived, and peak positions with biological significance are selected as the peak position range of potential biomarkers based on previous literature. In the Control vs MPN subtypes, Control vs CNL, and Control vs PMF models, the screening and validation of biomarkers were carried out using the same method as above.

[0070] 6.4 Results of validating biomarkers using statistical analysis

[0071] Control vs MPN vs MDS / MPN vs AML

[0072] Figure 6 In which, A is the statistical analysis of Raman characteristic peak positions with VIP>0.5 in the four models of Control vs MPN vs MDS / MPN vs AML, Control vs MPN, Control vs MDS / MPN, and Control vs AML. Figure 6 In which, B-D show the three peripheral blood biochemical indexes of total protein (TP), low-density lipoprotein (LDL), and adenosine deaminase (ADA) in the four groups of control, MPN, MDS / MPN, and AML. Through statistical analysis, it can be seen that the peak intensity representing proteins in the Control group is higher than that in other groups, among which 643, 759, 1031, 1260, 1603, 1616 cm -1 showed significant statistical differences (P<0.05) when compared with the disease groups. 759 cm -1 showed significant statistical differences (P<0.001) when compared with the MPN, MDS / MPN, and AML groups, which was consistent with the serological results (see Figure 6 in B). The peak positions representing lipids in the Control group (1437, 1443, 1446 cm -1 ) had higher intensities than those in other groups. 1443 and 1446 cm -1 showed significant statistical differences (P<0.01) when compared with the MPN, MDS / MPN, and AML groups, which was consistent with the serological results (see Figure 6in C). The peak intensity of the Control group representing collagen (1345 cm -1 ) was significantly lower than that of the MPN and MDS / MPN groups (P < 0.01). The peak intensity of the Control group representing nucleic acid was higher than that of other groups. Among them, 786 and 1579 cm -1 showed significant statistical differences in comparison with the disease groups (P < 0.05). 1579 cm -1 showed significant statistical differences in comparison with the MPN, MDS / MPN, and AML groups (P < 0.001), which was consistent with the serological results (see Figure 6 in D). The peak intensity of the peak position (957 cm -1 ) of the Control group representing β-carotene was significantly higher than that of the MPN, MDS / MPN, and AML groups (P < 0.01).

[0073] 6.5 Control vs MPN subtypes

[0074] Figure 7 A in shows the statistical analysis of Raman characteristic peak positions with VIP > 0.5 for the three models of Control vs MPN subtypes, Control vs CNL, and Control vs PMF. Figure 7 B - D in show the three peripheral blood biochemical indexes of TP, LDL, and ADA for the three groups of control, CNL, and PMF. Through statistical analysis, it can be seen that the peak intensity of the Control group representing protein was higher than that of MPN subtypes. Among them, 759, 1260, and 1603 cm -1 showed significant statistical differences in comparison with MPN subtypes (P < 0.05). 759 cm -1 showed significant statistical differences in comparison with the CNL and PMF groups (P < 0.01), which was consistent with the serological results (see Figure 7 in B). The intensity of the peak positions (1443, 1446 cm -1 ) of the Control group representing lipids was significantly higher than that of the CNL and PMF groups (P < 0.05), which was consistent with the serological results (see Figure 7 in C). The peak intensity of the peak position (1345 cm -1 ) of the Control group representing collagen was significantly lower than that of the CNL and PMF groups (P < 0.01). The peak intensity of the Control group representing nucleic acid was higher than that of MPN subtypes. Among them, 786 cm -1 showed statistical differences in comparison with the PMF group, and 1579 cm -1 showed statistical differences in comparison with the CNL and PMF groups (P < 0.05), which was consistent with the serological results (Figure 7 in D). The Control group represents the peak position of β-carotene (957 cm -1 ), and the peak intensity is significantly higher than that of the CNL and PMF groups (P < 0.05).

[0075] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A prediction or diagnosis system for myeloid tumors, characterized in that, The prediction or diagnosis system includes a Raman spectrum detection module for the sample to be tested and a Raman spectrum detection module for the negative control; both the Raman spectrum detection module for the sample to be tested and the Raman spectrum detection module for the negative control use the detection of Raman spectrum data of serum as an indicator; the Raman spectrum data are Raman peaks 786 and 1579 cm representing nucleic acids -1 , Raman peaks 643, 759, 1031, 1260, 1603, and 1616 cm representing proteins -1 , Raman peaks 1437, 1443, and 1446 cm representing lipids -1 , Raman peak 957 cm representing β-carotene -1 and Raman peak 1345 cm representing collagen -1 .

2. The prediction or diagnosis system according to claim 1, wherein The prediction or diagnosis system further includes a conclusion output module. Compared with the Raman spectroscopy data of the negative control, when the peak intensities of the protein Raman peak, nucleic acid Raman peak, β-carotene Raman peak, and lipid Raman peak in the Raman spectroscopy data of the sample to be tested are significantly reduced, and the peak intensity of the collagen Raman peak in the Raman spectroscopy data of the sample to be tested is significantly increased, it is predicted or diagnosed as myeloid neoplasm.

3. The prediction or diagnosis system according to claim 1 or 2, characterized in that The myeloid neoplasm includes one or more of myelodysplastic / myeloproliferative neoplasms, myeloproliferative neoplasms, and acute myeloid leukemia.

4. The prediction or diagnosis system according to claim 1 or 2, characterized in that The myeloproliferative neoplasm is one or both of chronic neutrophilic leukemia and primary myelofibrosis.

5. The prediction or diagnosis system according to claim 1 or 2, characterized in that, The negative control is healthy human serum.

6. Application of the intensity levels of Raman peaks 786 and 1579 cm representing nucleic acids, Raman peaks 643, 759, 1031, 1260, 1603, and 1616 cm representing proteins, Raman peaks 1437, 1443, and 1446 cm representing lipids, Raman peak 957 cm representing β-carotene, and Raman peak 1345 cm representing collagen as biomarkers in the preparation of products for predicting or diagnosing myeloid neoplasms or their subtypes. -1 -1 -1 -1 -1 ​​​​​ 7. The application according to claim 6, wherein The myeloid neoplasm includes one or more of myelodysplastic / myeloproliferative neoplasms, myeloproliferative neoplasms, and acute myeloid leukemia.

8. Use of the biomarker according to claim 6 in the preparation of a product for predicting or diagnosing a subtype of myeloproliferative neoplasm.

9. The application according to claim 8, characterized in that, The subtype of myeloproliferative neoplasm is one or both of chronic neutrophilic leukemia and primary myelofibrosis.

10. A computer-readable storage medium for storing computer instructions, programs, code sets, or instruction sets, which, when run on a computer, cause the computer to execute the prediction or diagnosis system according to any one of claims 1 to 5.

Citation Information

Cited By

  • Application of serum Raman spectroscopy in detection of lymphoplasmal cell lymphoma and diffuse large B-cell lymphoma

    CN121577608A