Method and kit for constructing a differential model for MDS and AML based on Raman spectroscopy and serum metabolism
By using Raman spectroscopy technology to detect the intensity of specific peaks in serum, a model for distinguishing MDS and AML was constructed, which solved the problems of high cost and invasiveness of traditional detection methods, as well as difficulty in early identification of the transformation of MDS to AML, and achieved low-cost and rapid disease identification.
Patent Information
- Application Number
- CN202210717142.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-23
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2042-06-23
AI Technical Summary
There is a lack of effective methods for diagnosing the transformation of MDS to AML in existing technologies. Traditional detection methods are costly and invasive, and serological indicators related to glycolipid metabolism lack sufficient basis, making it difficult to identify MDS and AML early.
Raman spectroscopy combined with the peak intensity of serum metabolites was used as a biomarker to construct a model for distinguishing MDS and AML. The peak intensities representing collagen, carbohydrates, proteins, lipids and nucleic acids in serum were detected by Raman spectroscopy, and an OPLS-DA diagnostic model was established for identification.
It achieves non-invasive and low-cost differentiation of MDS and AML, can quickly identify the transformation trend of MDS to AML, and improves the efficiency and accuracy of disease diagnosis.
Smart Images

Figure CN115165838B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of medical diagnosis, and in particular relates to a method and a kit for constructing a myelodysplastic syndrome (MDS) and acute myeloid leukemia (AML) identification model based on Raman spectroscopy and serum metabolism. Background Art
[0002] Acute leukemia (AML) is a malignant clonal disease originating from hematopoietic stem cells. Based on cell morphology, it can be divided into two major categories: acute myeloid leukemia and acute lymphoblastic leukemia, each of which is further divided into several subtypes. Acute myeloid leukemia (AML) is the most common leukemia in adults. It results from abnormal proliferation and differentiation of hematopoietic cells in the bone marrow. Clinically, AML can be divided into primary AML and secondary AML. Secondary AML is defined as a disease that does not meet the criteria for leukemia at diagnosis but transforms into leukemia during disease progression. Compared with primary AML, secondary AML is more challenging to treat and has a worse prognosis. Some patients with myelodysplastic syndromes (MDS), myeloproliferative neoplasms (MPN), lymphoma, paroxysmalnocturnal hemoglobinuria (PNH), multiple myeloma, and chronic lymphocytic leukemia may develop secondary AML.
[0003] Myelodysplastic syndromes (MDS) typically present with a combination of hematopoietic cell proliferation and apoptosis. They remain considered the most difficult chronic myeloid neoplasms to diagnose and classify correctly, carrying the risk of transformation to acute myeloid leukemia (AML). Clinically, MDS is categorized into MDS with single-lineage dysplasia (MDS-SLD), MDS with multilineage dysplasia (MDS-MLD), MDS with excess blasts (MDS-EB), and unclassifiable MDS (MDS-U). MDS-EB is further subdivided into MDS-EB1 and MDS-EB2, and patients are at risk of bone marrow failure and transformation to AML. Traditional testing for MDS primarily includes blood and bone marrow examinations, cytogenetic testing, immunophenotyping, and genetic analysis. There is currently no "gold standard" for the diagnosis of MDS; it is a diagnosis of exclusion, requiring the support of test results from multiple platforms to differentiate it from AML. Refractory anemic MDS is easily confused with AML, particularly the MDS-EB subtype. Therefore, developing a low-cost diagnostic method that facilitates early identification of MDS and AML is of great significance in the diagnosis of hematologic diseases.
[0004] Traditional diagnostic methods are time-consuming, costly, and invasive, and there is no "gold standard" for AML and MDS. Furthermore, existing technologies lack sufficient evidence to link serological markers related to glycolipid metabolism with the transformation of MDS to AML, failing to provide sufficient information for rapid and early differentiation of MDS and AML. Therefore, developing a rapid identification method for MDS and AML based on Raman spectroscopy that does not require antibody labeling, and further exploring serological markers that are meaningful for the diagnosis of MDS and AML patients to construct disease diagnostic models, are of great value for early identification of the disease. This will also improve diagnostic efficiency, reduce testing costs, and more quickly and accurately identify the trend of MDS to AML transformation. Summary of the Invention
[0005] In view of this, the present invention aims to propose a method for constructing a differentiation model between myelodysplastic syndrome (MDS) and acute myeloid leukemia (AML) based on Raman spectroscopy and serum metabolism, so as to overcome the defects of existing technologies, which are expensive, invasive and unable to effectively identify the trend of MDS to AML transformation.
[0006] To achieve the above object, the technical solution of the present invention is achieved as follows:
[0007] Provided is a myelodysplastic syndrome (MDS) and acute myeloid leukemia (AML) identification model, comprising the following modules:
[0008] 1) MDS and AML Raman spectroscopy identification module: serum Raman spectroscopy data is used to represent collagen (859 and 1345 cm -1 ) and carbohydrates (920 and 1123cm -1 ) as an indicator;
[0009] 2) Differentiation module between MDS-SLD / MLD group and MDS-EB1 / MDS-EB2 / primary AML / secondary AML group: serum Raman spectroscopy data were used to represent proteins (853, 1003, 1206 and 1616 cm -1 ) and the peak intensities of lipids (1437, 1443 and 1446 cm -1 ) as an indicator;
[0010] 3) Differentiation module between MDS-EB1 / MDS-EB2 group and primary AML group: serum Raman spectroscopy data represents nucleic acid (781, 786 and 1485 cm -1 ) and carbohydrates (920cm -1 ) as an indicator;
[0011] Also included is a comparison module: Representative collagen levels in AML patients compared to MDS (859 and 1345 cm -1 ) and carbohydrates (920 and 1123cm -1 ) peak intensities were significantly increased; compared with the MDS-SLD / MLD group, the peak intensities of representative proteins (853, 1003, 1206, and 1616 cm) in the MDS-EB1, MDS-EB2, primary AML, and secondary AML groups were significantly increased. -1 ) decreased significantly, while the peak intensities of lipids (1437, 1443 and 1446 cm -1 ) increased significantly; the peak intensities of nucleic acids (781, 786, and 1485 cm -1 ) and carbohydrates (920cm -1 ) was significantly lower than that in the primary AML group.
[0012] Another object of the present invention is to provide a Raman spectroscopic identification method for different MDS subtypes and different AML subtypes. To achieve the above object, the technical solution of the present invention is implemented as follows:
[0013] Provided is a method for detecting representative collagen (859 and 1345 cm) in serum based on Raman spectroscopy. -1 ) and carbohydrates (920 and 1123cm -1 ) as a biomarker in the preparation of a kit for distinguishing MDS and AML.
[0014] Provided is a method for detecting representative proteins (853, 1003, 1206 and 1616 cm) in serum based on Raman spectroscopy. -1 ) and the peak intensities of lipids (1437, 1443 and 1446 cm -1 ) peak intensities represent proteins (853, 1003, 1206, and 1616 cm -1 ) and the peak intensities of lipids (1437, 1443 and 1446 cm -1 ) as a biomarker in the preparation of kits for the MDS-SLD / MLD group and the MDS-EB1 / MDS-EB2 / primary AML / secondary AML group.
[0015] Provided is a method for detecting representative nucleic acids (781, 786 and 1485 cm) in serum based on Raman spectroscopy. -1 ) and carbohydrates (920cm -1 ) as a biomarker in the preparation of MDS-EB1 / MDS-EB2 group and primary AML group kits.
[0016] Furthermore, in the model of the present invention and any of the above-described applications, the Raman spectroscopy measurement conditions are as follows: a 785 nm laser is selected as the excitation light, the output power is 40 mW, the objective lens is selected at 40 times, the specimen is fixed on an XYZ three-dimensional platform, a ×40 0.75 NA Nikon lens is used during the shooting process, a spot size range of approximately 2 × 2 μm on the sample is irradiated by a laser beam with an output power of 40 mW, a single integration time is 250 s, the number of integrations is one, and the measurement range is 600-1800 cm -1 , 5-10 sites were measured in each group with a resolution of 1cm -1 .
[0017] Compared with the existing technology, the present invention has the following beneficial effects: using Raman peak positions as biomarkers for MDS and AML, this non-invasive method makes a major contribution to disease classification, and based on Raman spectroscopy, MDS and AML subtypes can also be identified and classified, laying the foundation for better use of clinical serological test data for rapid and early identification of MDS and AML. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 Schematic diagram of Raman measurement of serum;
[0019] Figure 2A Raman spectra of healthy controls, MDS, and AML samples; Figure 2B Raman spectra of healthy controls, MDS-SLD / MLD, MDS-EB1, and MDS-EB2; Figure 2C Raman spectra of healthy controls, primary AML, and secondary AML; Figure 2D Permutation, cluster, and ROCplots analysis of the OPLS model for healthy controls, MDS, and AML samples; Figure 2E OPLS discrimination validation diagrams for healthy controls vs MDS vs AML, different MDS subtypes vs primary AML, and different MDS subtypes vs secondary AML models;
[0020] Figure 3A Score plot, loading plot, and V+Splot were drawn using Hotelling's 95% confidence ellipse for the OPLS-DA of the four models: healthy controls vs MDS vs AML, healthy controls vs MDS, healthy controls vs AML, and MDS vs AML; Figure 3B The statistical results of potential biomarkers for the four models of healthy controls vs MDS vs AML, healthy controls vs MDS, healthy controls vs AML, and MDS vs AML are shown; Figure 3C Serum biochemical results for healthy controls, MDS, and AML; Figure 3D Volcano plot of significantly differentially expressed genes between MDS and AML groups; Figure 3E Heat map of significantly differentially expressed genes between MDS and AML groups; Figure 3F Bubble plot for functional enrichment analysis of significantly differentially expressed genes between MDS and AML groups; Figure 3G The biological processes, cellular components, molecular functions, and KEGG enrichment results of significantly differentially expressed genes between MDS and AML groups; Figure 3HTo identify genes that are most likely to be involved in the functional and regulatory signaling pathways of MDS and AML;
[0021] Figure 4A The score plot, loading plot, and V+S plot were drawn using Hotelling's 95% confidence ellipse for the OPLS-DA of the four models: MDS subtypes vs primary AML, MDS-SLD / MLD vs primary AML, MDS-EB1 vs primary AML, and MDS-EB2 vs primary AML; Figure 4B The statistical results of potential biomarkers for four models: different subtypes of MDS vs primary AML, MDS-SLD / MLD vs primary AML, MDS-EB1 vs primary AML, and MDS-EB2 vs primary AML; Figure 4C Serum biochemical results for MDS-SLD / MLD, MDS-EB1, MDS-EB2, and primary AML; Figure 4D The score plot, loading plot, and V+S plot of the OPLS-DA using Hotelling's 95% confidence ellipse for the four models of MDS different subtypes vs secondary AML, MDS-SLD / MLD vs secondary AML, MDS-EB1 vs secondary AML, and MDS-EB2 vs secondary AML were drawn. Figure 4E The statistical results of potential biomarkers for four models: different subtypes of MDS vs secondary AML, MDS-SLD / MLD vs secondary AML, MDS-EB1 vs secondary AML, and MDS-EB2 vs secondary AML; Figure 4F Serum biochemical results for MDS-SLD / MLD, MDS-EB1, MDS-EB2, and secondary AML;
[0022] Figure 5A The Permutation, Cluster, ROC, and Validation plots of OPLS discrimination for the healthy control group and MDS-SLD / MLD group; Figure 5B The Permutation, Cluster, ROC and OPLS discrimination validation plots of the healthy control group and MDS-EB1 group; Figure 5CThe Permutation, Cluster, ROC and Validation plots of OPLS discrimination for the healthy control group and MDS-EB2 group; Figure 5D The Permutation, Cluster, ROC and OPLS discrimination validation plots for healthy controls and primary AML groups; Figure 5E The Permutation, Cluster, ROC and OPLS discrimination validation plots for healthy controls and secondary AML groups;
[0023] Figure 6A Volcano plot of significantly differentially expressed genes between healthy controls and MDS groups; Figure 6B Heat map of significantly differentially expressed genes between healthy controls and MDS groups; Figure 6C Bubble plot for functional enrichment analysis of significantly differentially expressed genes between healthy controls and MDS groups; Figure 6D The biological processes, cellular components, molecular functions and KEGG enrichment results of significantly differentially expressed genes between healthy controls and MDS groups; Figure 6E To identify genes that are most likely to be involved in the functional and regulatory signaling pathways identified in healthy controls and MDS; Figure 6F Volcano plot of significantly differentially expressed genes between healthy controls and AML groups; Figure 6G Heat map of significantly differentially expressed genes between healthy controls and AML groups; Figure 6H Bubble plot for functional enrichment analysis of significantly differentially expressed genes between healthy controls and AML groups Figure 6I The biological processes, cellular components, molecular functions, and KEGG enrichment results of significantly differentially expressed genes between healthy controls and AML groups; Figure 6J To identify genes that are most likely involved in the functional and regulatory signaling pathways identified in healthy controls and AML;
[0024] Figure 7 A is the degree value of the key differential candidate genes screened by MDS vs AMLL; Figure 7 B is the degree value of the key differential candidate genes screened by the healthy control group vs. MDSL; Figure 7 C is the degree value of the key differential candidate genes screened between the healthy control group and AML. DETAILED DESCRIPTION
[0025] Unless otherwise defined, technical terms used in the following examples have the same meanings as commonly understood by those skilled in the art to which this invention belongs. Unless otherwise specified, the test agents used in the following examples are all conventional biochemical reagents; and the experimental methods described are all conventional methods unless otherwise specified.
[0026] The present invention will be described in detail below with reference to the embodiments and accompanying drawings.
[0027] Example 1 Construction of MDS and AML identification models
[0028] 1. Experimental Methods
[0029] 1. Collection of samples and clinical information of MDS, AML patients and healthy controls
[0030] Fifty-eight patients who visited the Hematology Hospital of the Chinese Academy of Medical Sciences (Institute of Hematology, Chinese Academy of Medical Sciences) in 2021 were recruited, including 27 males and 31 females, with an age range of 5 to 77 years. They were divided into three groups: 33 patients in the MDS group (7 with MDS-SLD / MLD, 10 with MDS-EB1, and 16 with MDS-EB2); 25 patients with AML (20 with primary AML and 5 with secondary AML); and 29 healthy controls. All patients in the experimental group underwent blood and bone marrow phenotype analysis, cytogenetic analysis, immunophenotyping, and genetic analysis, and were diagnosed by experienced hematologists. This study was approved by the Ethics Committee of the Hematology Hospital of the Chinese Academy of Medical Sciences (Institute of Hematology, Chinese Academy of Medical Sciences). Serum samples used in the study were obtained from residual samples collected from clinical testing of patients with MDS, AML, and controls. Routine serum biochemistry testing was performed on all patients and controls. Serum biochemistry data were obtained from the Clinical Testing Center of the Hematology Hospital of the Chinese Academy of Medical Sciences (Institute of Hematology, Chinese Academy of Medical Sciences).
[0031] Methods: Serum was collected from the subjects after fasting for 10 hours. Peripheral blood total protein (TP), glucose, triglyceride (TG), total cholesterol (TC), high-density lipoprotein (HDL), and low-density lipoprotein (LDL) were measured using an automatic biochemical analyzer. The data of the MDS group, AML group, and healthy control group, as well as the data of the MDS subtype, different AML subtypes, and healthy control group are shown in Table 1.
[0032] 2. Raman spectroscopy analysis of peripheral blood of MDS, AML patients and healthy controls
[0033] To investigate the Raman spectral characteristics of serum from MDS and AML patients and controls, 173, 197, and 144 serum Raman spectra were collected from healthy controls, MDS, and AML groups, respectively. The MDS group included 42 spectra from patients with the MDS-SLD / MLD subtype, 60 spectra from patients with the MDS-EB1 subtype, and 95 spectra from patients with the MDS-EB2 subtype. The AML group included 114 spectra from patients with primary AML subtypes and 30 spectra from patients with secondary AML subtypes.
[0034] Methods: 5 μL of serum was dropped onto a quartz glass slide and measured using a confocal Raman spectrometer, an XploRA Raman microscope. Excitation light was a 785 nm laser with a 40 mW output power and a 40x objective lens. The specimen was mounted on an XYZ three-dimensional stage. A ×40 0.75 NA Nikon lens was used for imaging. A spot size of approximately 2 × 2 μm was illuminated by a 40 mW laser beam. The integration time was 250 s, with one integration attempt. The measurement range was 600–1800 cm⁻¹, with 5–10 sites measured per group at a resolution of 1 cm⁻¹. Raman spectra of the quartz glass slide were also measured as background. Data processing, including smoothing, background removal, and baseline correction, was performed using Labspec6 software. All spectra were intensity normalized using the respective 1450 cm⁻¹ Raman peak as an internal standard. Peak assignments for serum Raman spectra are shown in Table 1. Figure 1 Schematic diagram of serum Raman measurements. Raman spectra were obtained using Raman spectroscopy, and the dataset was analyzed using multivariate statistical analysis using SIMCA 14.1 to screen for potential biomarkers for MDS, AML, and MDS and AML subtypes.
[0035]
[0036]
[0037] Table 1. Peak assignments of serum Raman spectra
[0038] 3. Establishment of OPLS-DA diagnostic model based on Raman spectroscopy data analysis
[0039] 3.1 Establishment of OPLS-DA diagnostic model
[0040] SIMCA-P software was used to apply OPLS-DA to the sample data and conduct detailed analysis and comparison.
[0041] Eighteen characteristic spectra (six per category) were randomly selected from serum Raman spectroscopy data for healthy controls, MDS, and AML, forming three datasets. Twenty-four characteristic spectra (six per category) were randomly selected from serum Raman spectroscopy data for MDS-SLD / MLD, MDS-EB1, MDS-EB2, and primary / secondary AML, forming four datasets. Supervised orthogonal partial least-squares discrimination analysis (OPLS-DA) was performed on the serum Raman spectral data from MDS, AML, and controls using SIMCA14.1 software. The performance of the OPLS model was evaluated using the goodness-of-fit parameters R² and Q², respectively. Under the null hypothesis, the model was resampled 200 times by randomly varying the y matrix for model validation. Cluster analysis and V+S analysis were used to identify statistically significant Raman peak positions in the classification model as potential biomarkers. Based on a comprehensive consideration of parameters such as correlation coefficient, loading, and distance from the center of the V+S plot, peaks with a variable importance (VIP) greater than 1.0 were selected as potential biomarkers. Significance testing of these potential biomarkers was performed, and potential biomarkers with a P < 0.05 were considered statistically significant. Data were processed using Origin software. Statistical analysis was performed using IBM SPSS Statistics 20, and statistical correlation graphs were generated using Graphpad Prism 5.
[0042] 3.2 Verification of model validity
[0043] Permutation analysis, Cluster analysis, ROC (receiver operating characteristic curve) and OPLS discrimination validation plots were used to verify the validity of the model and to construct a validation model to evaluate the specificity of the constructed model.
[0044] 3.2.1 Permutation Analysis, Cluster, and ROC Analysis
[0045] This part of the method refers to the procedures in the literature Liang H, Cheng X, Dong S, et al. Rapid and non-invasive discrimination of acute leukemia bone marrow supernatants by Ramanspectroscopy and multivariate statistical analysis[J]. Journal of Pharmaceutical and Biomedical Analysis, 2022, 210: 114560.
[0046] 3.2.2 Establishment of a specificity evaluation model for validation identification methods
[0047] Based on the established discrimination models for healthy controls, MDS, and AML based on differences in serum substance levels, a validation model was developed to verify the validity of the discrimination method. The validation model consisted of a training set and a prediction set. The training set included 12 spectra from patients in the same group and 12 spectra from patients in a different group. The spectral data were explicitly grouped when the validation model was established. The prediction set included 8 spectra from patients in the same group and 8 spectra from patients in a different group. The spectral data were not grouped when the validation model was established. Using this method, nine validation models were constructed for healthy controls vs. MDS, healthy controls vs. AML, MDS vs. AML, MDS-SLD / MLD vs. primary AML, MDS-EB1 vs. primary AML, MDS-EB2 vs. primary AML, MDS-SLD / MLD vs. secondary AML, MDS-EB1 vs. secondary AML, and MDS-EB2 vs. secondary AML. The sensitivity and specificity of the validation models were determined. SIMCA-P software generated classification scores for each spectrum in the training and prediction sets based on the Raman spectral data. The software assigns classification scores to the two groups in the training set based on the spectral data groupings. The software also assigns scores to the two groups in the prediction set based on the similarity of their spectral data to the spectral data in the training set. A classification is considered correct if the first group in the training and prediction sets is assigned a positive score, and the second group is assigned a negative score. Otherwise, it is considered incorrect.
[0048] To more intuitively represent the sensitivity and specificity of the diagnostic model, a dot plot is drawn showing both the training set and the prediction set samples based on their classification scores. Setting the predicted value to 0 as the cutoff value gives the sensitivity and specificity of the classification and discrimination model.
[0049] 2. Experimental Results
[0050] 1. Clinical data on serum metabolites in the MDS, AML, and healthy control groups. Data that conformed to a normal distribution were expressed as mean ± standard deviation (SD). Means were compared between groups using one-way analysis of variance. Data that did not conform to a normal distribution were expressed as median (25th–75th percentile) using M(Q1-Q3). Intergroup comparisons were performed using the Kruskal-Wallis test. P values were used to compare differences between the healthy control group, MDS, and AML groups; P < 0.05 was considered statistically significant.
[0051]
[0052] Table 2 Clinical data of MDS, AML and healthy controls
[0053] 2. Clinical data on serum metabolites in different MDS and AML subtypes and healthy controls. Normally distributed data were expressed as mean ± standard deviations (SD). Intergroup means were compared using one-way analysis of variance. Data not normally distributed were expressed as median (25th–75th percentile) using M(Q1-Q3). Intergroup comparisons were performed using the Kruskal-Wallis test. P values were used to compare differences between healthy controls, different MDS subtypes, and different AML subtypes; P < 0.05 was considered statistically significant.
[0054]
[0055]
[0056] Table 3 Clinical data of different subtypes of MDS, AML and healthy controls
[0057] 3. Raman spectroscopy analysis of peripheral blood of MDS, AML patients and healthy controls
[0058] The results are shown in Figures 2A-2C middle, Figures 2A-2C Serum Raman spectra of healthy controls, MDS and AML, healthy controls and MDS subtypes in the range of 600-1800 cm-1 are shown respectively. The peak positions of the relevant serum Raman spectra are assigned to Table 1. Figure 2AFrom top to bottom are the average serum spectra of the healthy control group, MDS group and AML group. Figure 2B From bottom to top are the average serum spectra of the healthy control group, MDS-SLD / MLD subtype, MDS-EB1 subtype, and MDS-EB2 subtype. Figure 2C From bottom to top: average serum spectra of healthy controls, primary AML subtypes, and secondary AML subtypes; Figure 2A The Raman spectra of three groups of samples, namely healthy control, MDS and AML, are shown. The pink, yellow and blue vertical lines in the figure represent the Raman spectra of proteins (643, 759, 1003, 1260, 1603, 1654 cm -1 ), nucleic acid (826, 1579cm -1 ) and lipids (1446cm -1 ) related peak position; Figure 2B Raman spectra of healthy controls, MDS-SLD / MLD, MDS-EB1, and MDS-EB2 are given; Figure 2C The Raman spectra of healthy controls, primary AML, and secondary AML are shown, showing similar spectral patterns. This suggests that it is difficult to distinguish serum substances between healthy controls, MDS, and AML groups based solely on the Raman spectra and peak positions. A classification model established using the OPLS-DA method is needed to further identify peak positions that can effectively distinguish between MDS and AML groups as potential biomarkers.
[0059] 4. Establishment and validation of the OPLS-DA model for the identification of MDS and AML based on Raman spectroscopy
[0060] 4.1 Identification Model for MDS, AML, and Healthy Control Groups
[0061] Healthy controls vs MDS vs AML, healthy controls vs MDS, healthy controls vs AML, and MDS vs AML
[0062] The OPLS-DA score plot, loading plot and V+S plot of the four models are shown in Figure 3A middle.
[0063] The three sample groups are clearly separated in the score plot, with the healthy control group and the MDS group located on the positive x-axis, and the AML group on the negative x-axis, reflecting the distinction between the healthy control group, the MDS group, and the AML group. This differentiation demonstrates that OPLS-DA can effectively discriminate serum spectral data from healthy controls, MDS patients, and AML patients, laying the foundation for analyzing the material characteristics of these three groups.
[0064] The loading plot is used to preliminarily screen the Raman peaks that contribute to the identification model of healthy control vs AA vs MDS. The red, yellow, blue, green and purple peak numbers in the figure are related to proteins, nucleic acids, lipids, collagen and carbohydrates, respectively. There is a correlation between the loading plot and the score plot when representing the relationship between the sample substance content, that is, the substance represented by the peak located on the positive semi-axis of the vertical coordinate of the loading plot is relatively higher in the group on the positive semi-axis of the horizontal coordinate of the score plot, and the negative semi-axis of the vertical coordinate of the loading plot and the negative semi-axis of the horizontal coordinate of the score plot also have a similar corresponding relationship. Protein (897, 1003, 1260 and 1660 cm -1 ), nucleic acid (726cm -1 ), cholesterol / carotenoids (957cm -1 ), collagen (859, 1345cm -1 ) and carbohydrates (920, 1123cm -1 ) played an important role in the identification of the three groups of samples, reflecting that the collagen and carbohydrate contents of the AML group were higher than those of the healthy control and MDS groups, while the cholesterol content was lower than that of the healthy control and MDS groups.
[0065] The V+S plot integrates metrics such as VIP and correlation coefficient, which characterize the contribution of peak positions to the classification model, to further screen potential biomarkers. In subsequent analyses, significance testing of characteristic peak positions is performed to identify peaks that effectively differentiate healthy controls, MDS, and AML groups as potential biomarkers. The V+S plot provides a list of Raman peaks in descending order of VIP values. Peaks with a VIP greater than 1.0 and biological significance are selected as potential biomarkers. Figure 3B Statistical analysis of the Raman characteristic peak positions with VIP>1.0 for the four models: healthy control vs MDS vs AML, healthy control vs MDS, healthy control vs AML, and MDS vs AML. Figure 3C Six types of peripheral blood biochemical indicators, including total protein (TP), glucose, triglyceride (TG), total cholesterol (TC), high-density lipoprotein (HDL) and low-density lipoprotein (LDL), were displayed for three groups: healthy controls, MDS and AML.
[0066] Statistical analysis showed that the healthy control group represented proteins (897, 1003, 1206 and 1660 cm -1 ), cholesterol / carotenoids (957cm -1 ), nucleic acid (726cm -1 ) and collagen (1345cm -1 ) was significantly higher than that of the MDS group. The healthy control group represented cholesterol / carotenoids (957 cm -1 ), nucleic acid (726cm -1 ) was significantly higher than that of the AML group, while the peak intensity of collagen (1345 cm -1 ) and carbohydrates (920 and 1123cm -1 ) were significantly lower than those in the AML group. The MDS group represented collagen (859 and 1345 cm -1 ) and carbohydrates (920 and 1123cm -1 ) was significantly lower than that in the AML group. The above results were consistent with the serological results of TP, glucose and TC, which showed statistical differences between the two groups ( Figure 3C ).
[0067] 4.2 Establishment of a model for different MDS subtypes, different AML subtypes, and healthy controls
[0068] 4.2.1 Model for Differentiating MDS Subtypes from Primary AML
[0069] Based on the potential biomarkers for MDS and AML screened in Section 4.1, the MDS subtype model was used to analyze differences in sample material content and screen potential classification markers for MDS subtypes and primary AML. Figure 4A OPLS-DA score plot, loading plot and V+S plot of the four models: MDS subtype vs primary AML, MDS-SLD / MLD vs primary AML, MDS-EB1 vs primary AML, and MDS-EB2 vs primary AML.
[0070] The score plot shows a clear separation between MDS subtypes and primary AML samples, with the MDS subtype group located on the positive x-axis and the primary AML group on the negative x-axis, reflecting the distinction between the MDS subtype and primary AML groups. This differentiation demonstrates that OPLS-DA can effectively discriminate serum spectral data from patients with MDS-SLD / MLD, MDS-EB1, MDS-EB2, and primary AML, providing the foundation for analyzing the material characteristics of these four groups.
[0071] The loading plot was used to preliminarily screen the Raman peaks that contribute to the differentiation model of MDS subtypes vs. primary AML. The red, yellow, blue, green, and purple peak numbers in the figure are associated with proteins, nucleic acids, lipids, collagen, and carbohydrates, respectively. Proteins (853, 1003, 1206, 1616 cm -1 ), lipids (1437, 1443 and 1446 cm -1 ), nucleic acids (781, 786 and 1485 cm -1 ) and carbohydrates (920cm -1 ) played an important role in the identification of the five groups of samples, reflecting that the nucleic acid and carbohydrate contents of the primary AML group were higher than those of the MDS subtype group.
[0072] V+S plot was used to further screen potential biomarkers. In subsequent analysis, the characteristic peak positions were tested for significance, and the peak positions that could effectively distinguish MDS subtypes and primary AML were identified as potential biomarkers ( Figure 4A ).
[0073] Based on the MDS subtype vs primary AML model, primary AML was combined with the three groups of MDS subtype samples for OPLS-DA analysis, and the OPLS-DA score plot, loading plot, and V+S plot of the three models ( MDS-SLD / MLD vs primary AML, MDS-EB1 vs primary AML, and MDS-EB2 vs primary AML) were generated. Figure 4A The two groups of samples in the three score plots are located on the positive and negative half axes of the X-axis respectively. The samples are clearly clustered in the scatter plots, which reflects that the three models have good discrimination capabilities for the two groups of samples in the model.
[0074] The loading plot shows that the peak intensities of nucleic acid (781, 786, and 1485 cm-1) and carbohydrate (920 cm-1) peaks in the primary AML group are generally higher than those in the MDS subtype group. The peaks representing lipids (1437, 1443, and 1446 cm-1) are higher than those in the MDS-SLD / MLD group and lower than those in the MDS-EB1 and MDS-EB2 groups. The V+S plot provides the main basis for determining potential classification markers in MDS subtypes and primary AML models. Figure 4AA list of peaks with VIP>1.0 was derived from the V+S plot. Biomarker screening and validation were performed using the same method in the MDS subtype vs primary AML, MDS-SLD / MLD vs primary AML, MDS-EB1 vs primary AML, and MDS-EB2 vs primary AML models.
[0075] Figure 4B Statistical analysis of the Raman characteristic peak positions with VIP>1.0 of the four models: MDS subtype vs primary AML, MDS-SLD / MLD vs primary AML, MDS-EB1 vs primary AML, and MDS-EB2 vs primary AML. Figure 4C The six peripheral blood biochemical indicators of MDS-SLD / MLD, MDS-EB1, MDS-EB2 and primary AML groups are shown. Statistical analysis showed that the representative proteins (853, 1003, 1206 and 1616 cm -1 ) were significantly higher than those in MDS-EB1, MDS-EB2 and primary AML groups, while the peak intensities of lipids (1437, 1443 and 1446 cm -1 ) were significantly lower than those in the MDS-EB1, MDS-EB2, and primary AML groups; the MDS-EB1 and MDS-EB2 groups represented nucleic acids (781, 786, and 1485 cm -1 ) and carbohydrates (920cm -1 ) were significantly lower than those in the primary AML group. These results were consistent with the serological indicators such as TG, TC, HDL and LDL that showed statistical differences between the two groups ( Figure 4C ).
[0076] 4.2.2 Model for Differentiating MDS Subtypes from Secondary AML
[0077] Based on the screening of potential classification markers for MDS subtypes and primary AML, the MDS subtype model was used to screen potential classification markers for MDS subtypes and secondary AML. Figure 4DFigures 2 and 3 show the OPLS-DA score plot, loading plot, and V+S plot. The score plot shows clear clustering of the four sample groups, distributed across four quadrants, reflecting the distinction between the MDS subtypes and the secondary AML group. This differentiation demonstrates that OPLS-DA can effectively discriminate serum spectral data from patients with MDS-SLD / MLD, MDS-EB1, MDS-EB2, and secondary AML, providing the foundation for analyzing the material characteristics of these four groups.
[0078] The loading plot was used to preliminarily screen the Raman peaks that contribute to the differentiation model of MDS subtypes vs. secondary AML. The red, yellow, blue, green, and purple peak numbers in the figure are associated with proteins, nucleic acids, lipids, collagen, and carbohydrates, respectively. Proteins (1003, 1616 cm -1 ) and lipids (1437, 1443cm -1 ) played an important role in the identification of the four groups of samples, reflecting that the lipid content of the secondary AML group was higher than that of the MDS subtype group. V+S plot was used to further screen potential biomarkers. In subsequent analysis, the significance test of the characteristic peaks was performed to determine the peaks that could effectively distinguish the healthy control group and the MDS subtype as potential biomarkers ( Figure 4D ).
[0079] Based on the MDS subtype vs secondary AML model, secondary AML was combined with the three groups of MDS subtype samples for OPLS-DA analysis, and the OPLS-DA score plot, loading plot, and V+S plot of the three models (MDS-SLD / MLD vs secondary AML, MDS-EB1 vs secondary AML, and MDS-EB2 vs secondary AML) were generated. Figure 4D ), the two groups of samples in the three score plots are located on the positive and negative half axes of the X-axis respectively. The sample clusters in the scatter plots are obvious, reflecting that the two models have good discrimination capabilities for the two groups of samples in the model. The loading plot shows that the peak position of the protein representative in the secondary AML group (1003cm -1 ) intensity was higher than that of MDS-EB1 and MDS-EB2 groups, but lower than that of MDS-SLD / MLD group. V+S plot provides the main basis for determining potential classification markers in MDS subtypes and secondary AML models ( Figure 4DA list of peaks with VIP>1.0 was derived from the V+S plot. Biomarker screening and validation were performed using the same method in the MDS subtype vs. secondary AML, MDS-SLD / MLD vs. secondary AML, MDS-EB1 vs. secondary AML, and MDS-EB2 vs. secondary AML models.
[0080] Figure 4E Statistical analysis of the Raman characteristic peak positions with VIP>1.0 of the four models: MDS subtype vs secondary AML, MDS-SLD / MLD vs secondary AML, MDS-EB1 vs secondary AML, and MDS-EB2 vs secondary AML. Figure 4F The six peripheral blood biochemical indicators of MDS-SLD / MLD, MDS-EB1, MDS-EB2 and secondary AML groups are shown. Statistical analysis showed that the representative proteins (1003 and 1616 cm -1 ) were significantly higher than those in MDS-EB1, MDS-EB2 and secondary AML groups, while the peak intensities of lipids (1437 and 1443 cm -1 ) were significantly lower than those in the MDS-EB1, MDS-EB2, and secondary AML groups. These results were consistent with the serological indicators such as TG, TC, HDL, and LDL that showed statistical differences between the groups ( Figure 4F ).
[0081] 4.3 The identification model of MDS, AML and healthy control group was constructed.
[0082] Verification of the validity of the discrimination model between the same subtype and healthy control group
[0083] The validity of the constructed model was verified by using Permutation analysis, Cluster analysis, ROC (Receiver Operating Characteristic) curve and Validation plots of OPLS discrimination. The results of this part are shown in Figure 5: Figure 5A Permutation analysis, Cluster analysis, ROC receiver operating characteristic curve and OPLS discrimination validation plot for the healthy control group and MDS-SLD / MLD group;
[0084] Figure 5B Permutation analysis, Cluster analysis, ROC receiver operating characteristic curve and OPLS discrimination validation diagram for the healthy control group and MDS-EB1 group; Figure 5C Permutation analysis, Cluster analysis, ROC receiver operating characteristic curve and OPLS discrimination validation plot for the healthy control group and MDS-EB2 group; Figure 5D Permutation analysis, Cluster analysis, ROC receiver operating characteristic curve and OPLS discrimination validation diagram for the healthy control group and primary AML group; Figure 5E Permutation analysis, Cluster analysis, ROC receiver operating characteristic curve and OPLS discrimination validation diagram for the healthy control group and secondary AML group;
[0085] 4.3.1 Permutation Analysis
[0086] Permutation analysis showed that Q 2 The intercept on the Y-axis is negative, indicating that the OPLS-DA model is valid and not overfitting ( Figure 5A -E).
[0087] 4.3.2 Cluster Analysis
[0088] Figure 5A The Cluster analysis results show that the OPLS-DA model can effectively distinguish the healthy control group from MDS-SLD / MLD. Figure 5B ), healthy controls and MDS-EB2 ( Figure 5C ), healthy controls and primary AML ( Figure 5D ) and healthy control group and secondary AML group ( Figure 5E ) Using Raman spectroscopy to identify can achieve 100% identification accuracy.
[0089] 4.3.3 ROC (Receiver Operating Characteristic) Analysis
[0090] (1) Figure 2D The ROC curves of different identification models are shown, indicating that the discriminant analysis results are highly accurate:
[0091] Healthy control vs MDS vs AML discrimination model: AUC(healthy control)=1, AUC(MDS)=1, AUC(AML)=1;
[0092] MDS subtype vs primary AML discrimination model: AUC(MDS-SLD / MLD)=0.75, AUC(MDS-EB1)=0.842593, AUC(MDS-EB2)=0.805556, AUC(primary AML)=1;
[0093] MDS subtype vs secondary AML differentiation model: AUC(MDS-SLD / MLD)=1,AUC(MDS-EB1)=1,AUC(MDS-EB2)=1,AUC(secondary AML)=1.
[0094] (2) Figure 5A -E shows the ROC plot curve for the identification of different MDS subtypes and AML subtypes using the discrimination model. The data shows that the discriminant analysis results are highly accurate:
[0095] A. Healthy control and MDS-SLD / MLD groups, AUC(healthy control)=1, AUC(MDS-SLD / MLD)=1;
[0096] B. Healthy control and MDS-EB1 groups, AUC(healthy control)=1, AUC(MDS-EB1)=1;
[0097] C. Healthy control and MDS-EB2 groups, AUC(healthy control)=1, AUC(MDS-EB2)=1;
[0098] D. Healthy control and primary AML groups, AUC(healthy control)=1, AUC(primary AML)=1;
[0099] E. Healthy control and secondary AML groups, AUC(healthy control)=1, AUC(secondary AML)=1.
[0100] 4.3.4 Validation plots of OPLS
[0101] Figure 5A -E shows that the sensitivity range of validation plots is 81%-100% and the specificity is 100%, reflecting that the OPLS-DA model can better distinguish the serum sample data of the healthy control group and the MDS / AML subtypes group, and the OPLS-DA model is evaluated from the perspective of practical validation.
[0102] Based on the above validity analysis, the established OPLS-DA model (1) can distinguish the Raman spectra of three types of serum samples, namely the healthy control group, the MDS patient group, and the AML patient group, with an accuracy rate of 100%; (2) can distinguish the Raman spectra of MDS-SLD / MLD, MDS-EB1 / 2 and primary AML serum samples with an accuracy rate of 100%; (3) can distinguish the Raman spectra of MDS-SLD / MLD, MDS-EB1, MDS-EB2 and secondary AML serum samples with an accuracy rate of 100%.
[0103] 4.4 Sensitivity and specificity of diagnostic models
[0104] In order to more intuitively represent the sensitivity and specificity of the diagnostic model, a dot plot showing the samples of the training set and the prediction set is drawn according to the classification scores of the training set and the prediction set ( Figure 2E By setting the predicted value to 0 as the cutoff, the sensitivity and specificity of the classification and discrimination models were obtained (Tables 4-5). The results show that the sensitivity of the nine validation models ranged from 75% to 100%, and the specificity ranged from 92% to 100%. This indicates that the OPLS-DA model can effectively discriminate between the healthy control group and the MDS / AML group, between different MDS subtypes and primary AML, and between different MDS subtypes and secondary AML serum sample data. The OPLS-DA model was evaluated from a practical validation perspective.
[0105]
[0106] Table 5 Sensitivity and specificity of diagnostic models for healthy controls, different MDS subtypes, and different AML subtypes
[0107] The results showed that the sensitivity range of the nine validation models was 75%-100%, and the specificity range was 92%-100%, indicating that the OPLS-DA model can well discriminate the serum sample data of the healthy control group and the MDS / AML group, MDS subtypes vs primary AML, and MDS subtypes vs secondary AML. The OPLS-DA model was evaluated from the perspective of practical validation.
[0108] Example 2 Bioinformatics Analysis
[0109] Based on Raman spectroscopy and multi-parameter analysis, bioinformatics was used to screen differentially expressed genes between MDS and AML, healthy controls and MDS, and healthy controls and AML, and to analyze the biological functions and regulatory pathways of the differentially expressed genes, in order to identify key genes that play a vital role in the biological processes associated with changes in healthy controls, MDS, and AML.
[0110] 1. Experimental Methods
[0111] Gene chip GSE15061 was retrieved and downloaded from the GEO (gene expression omnibus) database. R software and the bioconductor software package were used to compare and identify differentially expressed genes (DEGs) between healthy controls and MDS, healthy controls and AML, and MDS and AML. Differentially expressed genes were screened based on P < 0.01 and |logFC| > 0.6. GO functional annotation and KEGG signaling pathway enrichment analysis were performed on the identified differentially expressed genes using DAVID 6.8. Protein-protein interaction (PPI) network analysis was performed using the STRING online analysis tool and Cytoscape software.
[0112] 1. Data Source
[0113] The data are from the NCBI.GEO (Gene.Expression.Omnibus, GEO, http: / / www.ncbi.nlm.nih.gov / geo / ) database, data series number GSE15061 (species: Homo sapiens). The data contains 716 sample data, including 164 MDS, 202 AML, and 69 control bone marrow samples, all of which are expression profile data. All samples were analyzed using the GPL570.Affymetrix HG-U133_Plus_2 Array platform.
[0114] 2. Data Preprocessing and Differential Expression Analysis
[0115] For the downloaded original CEL format expression spectrum data, the R (.version.3.6.2) software package affy[2] was used to perform expression value background correction and expression spectrum data normalization, including conversion of the original data format, missing value supplementation, background correction (MAS method), and data normalization using quantile method (quantiles).
[0116] The gene expression matrix was divided into three groups: healthy controls vs. MDS, healthy controls vs. AML, and MDS vs. AML. Differentially expressed genes were screened. Significant P values for gene expression differences were calculated using the unpaired T-test provided by limma, with Benjamini & Hochberg (BH) correction applied. For each significantly differentially expressed gene, a P value less than 0.01 and |logFC| greater than 0.6 were required. Heatmaps of differentially expressed genes were generated using the R package pheatmap.
[0117] 3. Functional Analysis of Differentially Expressed Genes
[0118] The Database for Annotation, Visualization, and Integration Discovery (DAVID) online analysis tool (https: / / david.ncifcrf.gov / ) was used to perform GO functional annotation and KEGG signaling pathway enrichment analysis on differentially expressed genes. GO functions mainly include three aspects: biological process (BP), cellular component (CC), and molecular function (MF).
[0119] 4. PPI network construction and key gene analysis
[0120] We used the protein interaction database STRING11.0 (https: / / sring-db.org / ) to explore the relationships between proteins encoded by differentially expressed genes and construct a protein-protein interaction network (PPI) network. The PPI results were then opened and edited in Cytoscape software, and the network topology property metric, Degree Centrality, was used to analyze the scores of nodes in the network. A higher node score indicates a more important position in the network and a higher likelihood of being a key node. We used the genes corresponding to the proteins ranked in the top 10 of the PPI network nodes (degree.TOP10) as hub genes with high connectivity in the network.
[0121] 5. Statistical Analysis
[0122] SPSS 26.0 statistical software package was used to process the data. The frequency data were compared between groups using the chi-square test (Fisher's exact test); the data that met the normal distribution were analyzed using the [mean ± SD] and standard deviations (SD) are expressed. Data that do not conform to the normal distribution are expressed as M(Q1-Q3) [median (25th–75th percentile)]. One-way analysis of variance was used to compare the normally distributed data between groups. The LSD method was used for pairwise comparisons between groups with equal variances. The Tamhane's T2 method was used for pairwise comparisons between groups with unequal variances. The Kruskal-Wallis test was used for nonparametric analysis to compare the data that do not conform to the normal distribution between groups. P < 0.05 was considered statistically significant.
[0123] 3. Experimental results:
[0124] According to the screening criteria of differentially expressed genes, 456, 26, and 418 upregulated and 1003, 107, and 981 downregulated DEGs were identified in the comparison between MDS and AML, healthy controls and MDS, and healthy controls and AML, respectively. Figure 3D 、 Figure 6A and Figure 6F The red or blue dots in the figure represent genes that are significantly up-regulated or down-regulated, respectively. Figure 3E 、 Figure 6B and Figure 6G Functional analysis of MDS and AML showed that these DEGs were associated with Th17 cell differentiation, Pertussis, and Cytokine-cytokine receptor interaction pathways.
[0125] DAVID was used to perform gene ontology and pathway enrichment analysis of common DEGs:
[0126] (1) The differentially expressed genes between MDS and AML were significantly enriched in three KEGG pathways (Th17 cell differentiation, Pertussis, Cytokine-cytokine receptor interaction), and were significantly enriched in one GO-BP (mitotic spindle organization), three GO-CCs (kinetochore, condensed chromosome kinetochore, kinesin complex), and two GO-MFs (microtubule binding, microtubule motor activity). The results were arranged in ascending order of P value, and the top five GO functional results were plotted ( Figure 3F ,3G).
[0127] (2) The differentially expressed genes between healthy controls and MDS were significantly enriched in four KEGG pathways (Hematopoietic cell lineage, T cell receptor signaling pathway, Th1 and Th2 cell differentiation, PD-L1 expression and PD-1 checkpoint pathway in cancer), and were significantly enriched in four GO-BPs (T cell receptor signaling pathway, T cell activation, cell surface receptor signaling pathway, positive thymic T cell selection), two GO-CCs (alpha-beta T cell receptor complex, T cell receptor complex), and five GO-MFs (transmembrane signaling receptor activity, T cell receptor binding, non-membrane spanning protein tyrosine kinase activity, protein tyrosine kinase activity, transmembrane receptor protein tyrosinekinase activity transmembrane receptor protein tyrosine kinase activity) results. Arrange them in ascending order according to the P value, and take the top five GO function results for plotting ( Figure 6C ,6D).
[0128] (3) The differentially expressed genes between healthy controls and AML were significantly enriched in two KEGG pathways (T cell receptor signaling pathway, Th17 cell differentiation), and were significantly enriched in 0 GO-BP, 2 GO-CC (kinetochore, chromosome, centromeric region), and 1 GO-MF (microtubule binding). The top five GO function results were arranged in ascending order of P value and plotted ( Figure 6H ,6I). Combining the results of the three groups of differentially expressed genes, the main processes in the biological process (BP) were concentrated in the mitotic spindle and T cell receptor signal transduction; the cellular component (CC) was mainly present in the kinetochore centromere and chromosome; the molecular function (MF) mainly involved the microtubule; in the KEGG analysis, the regulated signals mainly involved the T cell signaling pathway, especially the Th17 cell differentiation pathway.
[0129] Based on the signaling pathway analysis of DEGs, the PPI network was used to identify key candidate genes.
[0130] (1) The PPI network of MDS and AML has a total of 1323 nodes and 27566 interaction pairs. The topological scores are high and can be regarded as key nodes of the network. The degree values of the top ten genes ( Figure 7 A, Table 6). Hub genes were confirmed using the cytohubba plug-in. The results showed that the hub genes were CDK1, CCNB1, IL1B, CCNA2, ITGAM, AURKB, TOP2A, KIF11, TLR4, and MAD2L1. The data showed that there may be strong interactions between them.
[0131]
[0132] Table 6 List of TOP10 differentially expressed genes in MDS vs AML PPI network degree values
[0133] Combined with the top five GO function results, it was suggested that among the top 10 genes, KIF11, AURKB, CCNB1, and MAD2L1 genes were involved in mitotic spindle organization; AURKB and MAD2L1 genes were mainly present in the kinetochores, MAD2L1 genes were mainly present in the condensed chromosome kinetochore, and KIF11 gene was mainly present in the kinesin complex; KIF11 gene mainly played a role in microtubule binding and microtubule motor activity; IL1B gene was involved in Th17 cell differentiation and cytokine-cytokinereceptor interaction signaling pathways, and ITGAM, IL1B, and TLR4 genes were involved in the pertussis signaling pathway ( Figure 3H ).
[0134] (2) The PPI network of healthy controls and MDS has a total of 83 nodes and 702 interaction pairs. The topological scores are high and can be regarded as key nodes of the network. The degree values of the top ten genes ( Figure 7 , Table 7).
[0135]
[0136]
[0137] Table 7 List of TOP10 differentially expressed genes in healthy controls vs MDS PPI network degree values
[0138] The Hub genes were confirmed using the cytohubba plug-in, and the results showed that the Hub genes were CD8A, CD19, IL7R, CD79A, CD2, CCR7, LCK, PAX5, CD247 and CD3D. The data showed that there may be strong interactions between them.Combined with the top five GO function results, it was suggested that among the top 10 genes, CD8A, CD247, and CD3D genes were involved in the T cell receptor signaling pathway, CD2 and CD8A genes were involved in T cell activation, CD2, CD8A, CD247, IL7R, and CD3D genes were involved in the cell surface receptor signaling pathway, and CD3D gene was involved in positive thymic T cell selection; CD3D, CD2, CD79A, CD8A, CD19, CCR7, and IL7R genes were mainly present on the external side of plasma membrane, and CD247 and CD3D genes were mainly present in the alpha-beta T cell receptor complex; CD79A, CD247, and CD3D genes mainly played a role in transmembrane signaling receptor activity, and LCK gene mainly played a role in T cell receptor binding. LCK gene mainly played a role in non-membrane spanning protein tyrosine kinase activity, protein tyrosine kinase activity, and transmembrane receptor protein tyrosine. Kinase activity molecular function; CD79A, LCK, CD8A, CD19, IL7R, CD3D genes are involved in the Primary immunodeficiency signaling pathway, CD2, CD8A, CD19, IL7R, CD3D genes are involved in the Hematopoietic cell lineage signaling pathway, LCK, CD8A, CD247, CD3D genes are involved in regulating the T cell receptor signaling pathway, LCK, CD247, CD3D genes are involved in regulating the Th1 and Th2 cell differentiation signaling pathways, LCK, CD247, CD3D genes are involved in regulating PD-L1 expression and PD-1 checkpoint pathway incancer pathway (. Figure 6E ).
[0139] (3) The PPI network of healthy controls and AML has a total of 1251 nodes and 21942 interaction pairs. The topological scores are high and can be regarded as key nodes of the network. The degree values of the top ten genes ( Figure 7 C, Table 8).
[0140]
[0141] Table 8 List of TOP10 differentially expressed genes in healthy controls vs AML PPI network degree values
[0142] The cytohubba plug-in was used to confirm the Hub genes. The results showed that the Hub genes were ITGAM, CD8A, CDK1, JUN, CCNB1, CCNA2, CD44, MMP9, KIF11, and AURKB. The data showed that there may be strong interactions between them. Combined with the top five GO function results, it was suggested that among the top 10 genes, the AURKB gene was mainly present in the kinetochore, chromosome, and centromeric regions; the KIF11 gene mainly played a microtubule binding molecular function; the JUN gene was involved in the T cell receptor signaling pathway and the Th17 cell differentiation signaling pathway ( Figure 6J ).
[0143] Bioinformatics analysis conclusions:
[0144] Based on the GEO database, we screened the genes related to MDS and AML and found that ITGAM, IL1B and TLR4 genes were involved in the pertussis signaling pathway among the key differentially expressed genes. Pertussis is a type of disease that can trigger a leukemic reaction in lymphocytes. Pertussis toxin (PTX) can inhibit T cell movement and displacement by inhibiting G protein activation. G protein is the abbreviation of "guanylate binding protein", which is a special protein in the cell membrane that is related to transmembrane signal transduction. In the Raman spectroscopy detection model of Example 1, it was verified that primary AML in the AML subtype and MDS-EB1 in the MDS subtype were at 1485 cm -1 The difference in peak intensity is statistically significant, indicating that the bioinformatics results based on patient big data analysis further validate the Raman spectroscopy diagnostic model. The multiple diagnostic models constructed in Example 1 are consistent with the results of key genes obtained by bioinformatics analysis, and they validate each other. -1The correlation between peak position and guanine metabolism is also consistent with the known conjecture in the prior art, namely: MDS-EB1 patients have T cell dysfunction and a high risk of transformation to AML, and abnormal nucleic acid metabolism in the MDS-EB1 metabolic microenvironment may be involved in the progression of MDS to AML.
[0145] Using bioinformatics analysis, we also found that Th17 cell differentiation is one of the pathways of differential gene action in MDS and AML. Th17 is a type of T cell subset that has received widespread attention in recent years. Th17 cells show heterogeneity and plasticity in different immune environments and are closely related to the occurrence of autoimmune diseases, tumors, etc. The metabolic microenvironment has an important influence on the differentiation and function of Th17. The serine / threonine kinase Akt signaling pathway is necessary for the occurrence of peripheral induction of Th17 (inducible Th17, iTh17). Serine and threonine are hydroxy aliphatic side chain amino acids. It can be seen that the lipids (1437, 1443 and 1446 cm) in the blood microenvironment of MDS and AML verified in the Raman spectroscopy detection model in Example 1 -1 ) peak position difference, which is consistent with the above conclusions in the prior art and verifies each other.
[0146] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
[0147] References:
[0148] [1] N. Stone, C. Kendall, J. Smith, P. Crow, H. Barr, Raman spectroscopy for identification of epithelial cancers, Faraday Discuss. 126 (2004) 141-57; discussion 169-83.
[0149] [2]VVPully, A.Lenferink, C.Otto, Time﹍apse Raman imaging of singlelive lymphocytes, J Raman Spectrosc.42(2015).
[0150] [3] J.L. González-Solís, J.C. Martínez-Espinosa, J.M. Salgado-Román, P. Palomares-和a, Monitoring of chemotherapy leukemia treatment using Raman spectroscopy and principal component analysis, Lasers Med Sci. 29(3)(2014)1241-9.
[0151] [4] A. Silva, F.S.A.d.S.e. Oliveira, P. Brito, L. Silveira, Spectral model for diagnosis of acute leukemias in whole blood and plasma through Raman spectroscopy, J Biomed Opt. 23(2018).
[0152] [5] H.F. Nargis, H. Nawaz, A. Ditta, T. Mahmood, M.I. Majeed, N. Rashid, M. Muddassar, H.N. Bhatti, M. Saleem, K. Jilani, F. Bonnier, H.J. Byrne, Raman spectroscopy of blood plasma samples from breast cancer patients at different stages, Spectrochim Acta A Mol Biomol Spectrosc. 222(2019)117210.
[0153] [6] S. Managò, C. Valente, P. Mirabelli, D. Circolo, F. Basile, D. Corda, A.C. De Luca, A reliable Raman-spectroscopy-based approach for diagnosis, classification and follow-up of B-cell acute lymphoblastic leukemia, Sci Rep. 6(2016)24821.
Claims
1. A model for differentiating myelodysplastic syndromes (MDS) and acute myeloid leukemia (AML), characterized in that: The identification model includes the following modules: 1) MDS and AML Raman spectroscopy identification module: serum Raman spectroscopy data representing collagen at 859 and 1345 cm -1 The peak intensities at 920 and 1123 cm representing carbohydrates -1 The peak intensity of is used as an indicator; 2) Differentiation module between MDS-SLD / MLD group and MDS-EB1 / MDS-EB2 / primary AML / secondary AML group: serum Raman spectroscopy data representing protein at 853, 1003, 1206 and 1616 cm -1 The peak intensities of 1437, 1443, and 1446 cm representing lipids -1 The peak intensity of is used as an indicator; 3) Differentiation module between MDS-EB1 / MDS-EB2 group and primary AML group: serum Raman spectroscopy data were used to identify the peak intensities of 781, 786 and 1485 cm-1 representing nucleic acids and 920 cm-1 representing carbohydrates. -1 The peak intensity was used as an indicator.
2. Detection of collagen-representing 859 and 1345 cm-1 in serum based on Raman spectroscopy -1 The peak intensities at 920 and 1123 cm representing carbohydrates -1 The application of the peak intensity level as a biomarker in the preparation of a kit for distinguishing MDS and AML.
3. Detection of 853, 1003, 1206 and 1616 cm-1 representing proteins in serum based on Raman spectroscopy -1 The peak intensities of 1437, 1443, and 1446 cm representing lipids -1 The application of peak intensity as a biomarker in the preparation of MDS-SLD / MLD group and MDS-EB1 / MDS-EB2 / primary AML / secondary AML group kits.
4. Detection of 781, 786 and 1485 cm-1 representing nucleic acids in serum based on Raman spectroscopy -1 The peak intensity of 920 cm representing carbohydrates -1 The application of peak intensity as a biomarker in the preparation of MDS-EB1 / MDS-EB2 group and primary AML group kits.
5. The model according to claim 1 or the use according to any one of claims 2 to 4, characterized in that: The Raman spectroscopy measurement conditions were as follows: a 785 nm laser was selected as the excitation light, the output power was 40 mW, the objective lens was selected at 40 times, the specimen was fixed on an XYZ three-dimensional stage, and a ×40 0.75 NA Nikon lens was used for the imaging process. The sample was illuminated by a laser beam with an output power of 40 mW within a 2 × 2 μm spot size range. The single integration time was 250 s, the number of integrations was one, and the measurement range was 600–1800 cm. -1 , 5-10 sites were measured in each group with a resolution of 1cm -1 .
Citation Information
Patent Citations
Biomarker for predicting or detecting acute leukemia
CN112763474A