Biomarkers for diagnosing or warning of acute coronary syndrome, screening methods thereof, and applications
Through machine learning algorithms and mass spectrometry data, ACS-related biomarkers were screened out, and predictive models were constructed, which solved the problems of low accuracy of ACS diagnosis and high time consumption, and achieved accurate identification and early diagnosis of ACS.
Patent Information
- Application Number
- CN202410455310.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-16
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2044-04-16
AI Technical Summary
The prior art has problems with low diagnostic accuracy, high time consumption and low patient acceptance in the diagnosis of acute coronary syndrome (ACS), and has limited understanding of metabolic changes in ACS.
11 biomarkers were screened through machine learning algorithms and mass spectrometry data, including D-erythroid-sphingosine-1-phosphate, glucose reducing ketone, etc., and a variety of machine learning algorithms were used to build prediction models to identify ACS patients and healthy subjects.
Accurate identification and early diagnosis of ACS are achieved, higher diagnostic accuracy and shorter diagnosis time are provided, and better decision support is provided to clinicians.
Smart Images

Figure CN118518860B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biotechnology, and particularly to biomarkers for diagnosing or warning of acute coronary syndrome, and screening methods and applications thereof. Background Art
[0002] Acute coronary syndrome (ACS) is acute myocardial ischemic necrosis caused by the sudden occlusion of the coronary artery. ACS is caused by atherosclerotic plaque rupture and is affected by genetic susceptibility factors exposed to the atherosclerotic environment. Although the research on the pathogenesis of ACS has been carried out for several years, its incidence is still increasing year by year. Improving the prediction rate of ACS patients and reducing the diagnosis time are of great significance for giving patients timely treatment to reduce the mortality rate.
[0003] Currently, clinically, the main basis for diagnosing ACS is the examination results such as coronary angiography, echocardiogram, and CT. However, these detections rely on professional equipment and personnel, have poor mobility, and have certain limitations for rapid detection and judgment in the early and emergency states. We urgently need to study the pathogenesis of ACS more comprehensively from a new perspective, and there is an urgent need to identify and develop diagnostic technologies and molecular biomarkers with higher diagnostic accuracy, shorter time consumption, and more convenient for patients to accept. Plasma metabolome analysis can simultaneously detect hundreds to thousands of metabolites, providing comprehensive information about the in-vivo metabolic state. This "metabolic fingerprint" method helps to more comprehensively understand the biological background of diseases and may reveal new pathological mechanisms. The changes in metabolites in plasma can reflect the subtle changes in physiological and pathological processes, which means that even in the early stage of acute coronary syndrome, accurate diagnosis may be possible through the changes in specific metabolites. Such changes in metabolites can provide a more sensitive or earlier disease indication than traditional biochemical markers. Currently, there have been studies exploring the metabolic characteristics related to coronary artery disease, but only a few studies with limited sample sizes have explored the metabolic changes in ACS. Since not all patients with coronary atherosclerosis will have acute cardiovascular events, our understanding of the overall metabolic changes in acute coronary syndrome is still limited, and further research is still needed.
[0004] Based on the progress and wide application of bioinformatics analysis and high-throughput sequencing technology, as well as the gradual maturity of machine learning in bioinformatics applications, this provides important methods and means for exploring the potential mechanisms, potential biomarkers, and therapeutic targets of various diseases. Screening biomarkers through machine learning is of great significance for the diagnosis or warning of ACS. Summary of the Invention
[0005] Aiming at the above-mentioned existing technologies, the purpose of the present invention is to provide biomarkers for diagnosing or warning of acute coronary syndrome, and screening methods and applications thereof.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] In a first aspect of the present invention, there is provided the use of one or more of the following biomarkers in the preparation of a product for diagnosing or warning of acute coronary syndrome:
[0008] (1) D-erythro-sphingosine-1-phosphate;
[0009] (2) Glucosone;
[0010] (3) 5-Hydroxyflavone;
[0011] (4) 2-cis-4-trans-Abscisic acid;
[0012] (5) Decanoyl-L-carnitine;
[0013] (6) Linoleoyl carnitine;
[0014] (7) D-erythro-sphingosine-1-phosphate;
[0015] (8) Acyl carnitine.
[0016] Preferably, the acyl carnitine is selected from acyl carnitine-20:4, acyl carnitine-20:3, acyl carnitine-9:0 or acyl carnitine-11:0.
[0017] The above 11 biomarkers can all accurately identify acute coronary syndrome, and the AUC value of each biomarker meets the diagnostic significance and has high clinical diagnostic application value; moreover, when multiple biomarkers are used in combination, the AUC is closer to 1 than that of a single biomarker, and the diagnostic effect is better.
[0018] Furthermore, using the following five biomarkers alone or in combination can achieve a better diagnostic effect:
[0019] (1) D-erythro-sphingosine-1-phosphate;
[0020] (2) Linoleoyl carnitine;
[0021] (3) D-erythro-sphingosine-1-phosphate;
[0022] (4) Acyl carnitine-20:4;
[0023] (5) Glucosone.
[0024] In a second aspect of the present invention, there is provided the use of a reagent for detecting at least one of the above biomarkers in the preparation of a product for diagnosing or warning of acute coronary syndrome.
[0025] Further, the reagent is a reagent for detecting the content of metabolic biomarkers in plasma samples.
[0026] In the third aspect of the present invention, a method for screening metabolic biomarkers for diagnosing or warning acute coronary syndrome is provided, including the following steps:
[0027] S1: Data acquisition and processing: including collecting plasma samples of a certain number of acute coronary syndrome patients and healthy people to construct an experimental group and a control group; extracting relevant metabolites from the collected plasma samples to further obtain non-targeted metabolomics data;
[0028] S2: Feature selection: performing differential analysis on the plasma metabolomics of acute coronary syndrome patients and healthy control groups, and integrating metabolomics data to obtain relevant features of acute coronary syndrome;
[0029] S3: Constructing a training model: using five interpretable machine learning algorithms, and constructing a prediction model based on the features selected from comprehensive metabolites;
[0030] S4: Model evaluation: using a five-fold cross strategy to evaluate the prediction performance of different machine learning algorithms;
[0031] S5: Importance ranking and interpretability analysis: based on the logistic regression model classifier, performing importance ranking and interpretability analysis on the selected metabolites.
[0032] The specific content of each step is further introduced below.
[0033] (1) The acquisition and processing of the metabolomics data in step S1, the specific process is as follows:
[0034] The subjects are divided into an experimental group and a control group. The patients all collect fasting venous blood on the morning of the second day after percutaneous coronary intervention, and the venous blood of the control group is collected during the health check; the collected plasma samples are thawed at 4°C, mixed by the vertex method, methanol is used as the extraction solvent for the plasma samples to extract the plasma samples; using a liquid chromatography system and a mass spectrometer to obtain the total ion chromatogram of all samples; subsequently, liquid chromatography and tandem mass spectrometry (LC-MS / MS) are used to qualitatively or quantitatively analyze the extracted plasma samples; to ensure data quality and improve the effect of model training, it is necessary to process missing data and identify and process or delete outliers by methods such as interpolation, deleting samples or features with missing values, and scaling all numerical features to a unified range or distribution to eliminate the influence of different dimensions.
[0035] (2) The feature selection described in step S2, the specific content is:
[0036] An in-depth analysis of the metabolic characteristics of ACS patients was conducted. Principal coordinate analysis (PCoA) and heatmap visualization clearly delineated the boundaries between HC patients and ACS patients; using Benjamini-Hochberg-corrected t-tests, 89 metabolites were higher in ACS patients compared to HC (false discovery rate FDR-P < 0.05, fold change (FC) > 1.5), and another 79 metabolites were lower (FDR-corrected P < 0.05, FC < 0.67). Additionally, 65 differential metabolites between ACS and HC patients (including 44 upregulated and 21 downregulated) were highlighted with FC > 2 or < 0.5 (FDR-p < 0.05).
[0037] (3) In step S3, a training model is constructed, and the specific method is as described below:
[0038] Five highly interpretable machine learning algorithms are adopted: logistic regression, support vector machine (SVM), decision tree, random forest, and XGBoost, and combined with traditional statistical methods (such as t-tests and partial least squares discriminant analysis (PLS-DA)) for biomarker screening. The "Shapley Additive Explanations (SHAP)" algorithm is used to calculate the contribution of each feature to the model prediction to identify the most influential metabolic biomarkers, and a prediction model is constructed based on the selected features to diagnose the feasibility of identifying ACS patients and HC subjects.
[0039] Logistic regression predicts the probability of a binary classification problem by fitting a logistic function. After the model training is completed, the coefficient (weight) corresponding to each feature represents the magnitude and direction of the influence of that feature on the prediction result. The statistical significance of the coefficient is usually evaluated through hypothesis testing, and features with a lower significance level (p-value) are considered to have an important statistical impact on the prediction result. A higher absolute value of the coefficient means that the feature has more predictive power in classification.
[0040] Feature selection in SVM usually combines recursive feature elimination (RFE) technology to identify the most influential features by recursively reducing the size of the feature set.
[0041] A decision tree constructs a tree structure by recursively splitting the dataset. Each internal node represents a decision rule on an attribute, each branch represents a decision result, and the final leaf node represents a class. When constructing a decision tree, the algorithm selects the optimal feature for each split, divides the dataset into subsets according to the feature values, so that each subset is as "pure" as possible (i.e., belonging to the same class).
[0042] Random forest is an ensemble learning algorithm that improves the prediction performance of the overall model by constructing multiple decision trees and aggregating their prediction results. When training each tree, random forest uses a randomly selected subset of features and a subset of data (bootstrap samples) to increase the diversity of the model. The importance of features is evaluated based on the average impurity reduction or average accuracy decrease of the features across multiple decision trees.
[0043] XGBoost (Extreme Gradient Boosting) is an efficient implementation of Gradient Boosting Decision Trees (GBDT) that helps identify the features most contributive to the prediction.
[0044] By comparing the performance of different models and the feature importance scores, the SHAP values are used to further understand the impact of each feature on the model prediction. Considering the consistency determination and importance ranking of each model for specific metabolites comprehensively, a group of the most influential metabolic markers is selected to achieve the purpose of identifying ACS patients and HC subjects.
[0045] (4) The specific method for model evaluation in step S4 is as follows:
[0046] A comprehensive comparison is made on the correlations of different machine learning methods for identifying ACS-related biomarkers; a five-fold cross-validation strategy is used for the entire metabolomics dataset (five-fold cross-validation is a commonly used evaluation method that divides the dataset into five parts, taking turns using four of them as training data and the remaining one as test data, and this process is repeated five times) to evaluate the prediction performance of five machine learning techniques (logistic regression, support vector machine, decision tree, random forest, and Xgboost). All the methods adopted can effectively distinguish ACS patients and HC subjects.
[0047] (5) The specific process for the interpretable analysis of the model to screen metabolic markers in step S5 is as follows:
[0048] Based on the "Shapley Additive Explanations (SHAP)" feature selection algorithm, the false discovery rate (FDR) is used for t-tests and the variable importance in projection (VIP) of partial least squares discriminant analysis (PLS-DA) to determine the metabolites that have the greatest impact on the discriminatory results and rank their importance, effectively facilitating the determination of more effective biological indicators.
[0049] All machine learning models unanimously recommend the described metabolic markers, which can effectively distinguish between the ACS and HC groups; eleven metabolic markers are significantly different among different groups, showing upregulated or downregulated states in ACS patients; the eleven metabolites have higher interpretability scores and significant odds ratios; we evaluated each metabolite separately as a potential biomarker for distinguishing ACS and HC subjects, and the results showed that all these metabolites demonstrated obvious discriminatory abilities, with the area under the receiver operating characteristic curve (ROC-AUC) results ranging from 0.84 to 0.97; to achieve the best balance between prediction performance and computational efficiency, an optimal combination of five downregulated biomarkers was finally selected, including D-erythro-phenylalanine-1-phosphate, linoleoyl carnitine, D-erythro-sphingosine-1-phosphate, acyl carnitine-20:4, and gluconerone (the ROC-AUC of this biomarker combination is 0.989, and the PR-AUC is 0.998).
[0050] Advantages of the present invention:
[0051] The present invention uses machine learning algorithms and mass spectrometry data to discover ACS-related metabolic biomarkers, which can accurately identify ACS and healthy subjects, providing effective tools and methods for the early diagnosis and prediction of ACS. At the same time, through interpretability analysis, the reliability and credibility of the model are enhanced, providing better decision support for clinicians. For diagnosis and prognosis, it not only has low cost and simple operation, but also is a rapid and non-invasive method, showing higher application value for the diagnosis of acute coronary syndrome, and can be widely applied to the prediction and diagnosis of acute coronary syndrome in the future. Description of the drawings
[0052] Figure 1 It is a principal coordinate analysis diagram in the analysis of metabolite characteristic differences between acute coronary syndrome and healthy control groups;
[0053] Figure 2 It is a volcano diagram in the analysis of metabolite characteristic differences between acute coronary syndrome and healthy control groups;
[0054] Figure 3 It is a heat map in the analysis of metabolite characteristic differences between acute coronary syndrome and healthy control groups;
[0055] Figure 4 Using five machine learning algorithms to construct a training model, which shows the number of important metabolites. This part shows the number of significant metabolites selected when using different machine learning methods to distinguish ACS patients from HC subjects, indicating the efficacy of various machine learning techniques in identifying metabolites that are important for the diagnosis of ACS;
[0056] Figure 5To construct a training model using five machine learning algorithms and demonstrate the effectiveness correlation of different machine learning methods;
[0057] Figure 6 To construct a training model using five machine learning algorithms and show the statistical data of significant metabolites using a specific machine learning method;
[0058] Figure 7 To construct a training model using five machine learning algorithms. The left panel shows the importance scores and normalized intensity levels of 11 selected significant metabolites; the right panel represents the association of these metabolites with ACS based on logistic regression analysis, expressed as odds ratio (OR) and 95% confidence interval (CI); the OR is at the center of the error bar, and the width of the 95% CI is represented by the line width;
[0059] Figure 8 To construct a training model using five machine learning algorithms and show the area under the receiver operating characteristic curve (ROC-AUC) of the metabolite panel-based logistic regression model for detecting ACS patients in five-fold cross-validation;
[0060] Figure 9 To construct a training model using five machine learning algorithms and show the area under the precision-recall curve (PR-AUC) of the same logistic regression model in five-fold cross-validation;
[0061] Figure 10 To construct a training model using five machine learning algorithms and independently evaluate the accuracy of the top five metabolites in differentiating HC subjects;
[0062] Figure 11 For the process and results of applying interpretive machine learning methods to select metabolite biomarkers for diagnosing acute coronary syndrome (ACS), the correlation between the importance scores obtained by machine learning and traditional statistical measurements (such as p-value or fold change FC value) is described;
[0063] Figure 12 For the process and results of applying interpretive machine learning methods to select metabolite biomarkers for diagnosing acute coronary syndrome (ACS), independently evaluate the predictive accuracy of using different biomarker combinations to distinguish patients from healthy control (HC) subjects;
[0064] Figure 13 For the process and results of applying interpretive machine learning methods to select metabolite biomarkers for diagnosing acute coronary syndrome (ACS), show the importance scores of the selected biomarkers obtained using the Xgboost algorithm;
[0065] Figure 14Receiver operating characteristic (ROC) curves for identifying ACS patients and validating the 11 biomarkers selected in the model in the validation cohort. Detailed implementation
[0066] It should be noted that the following detailed description is illustrative and aims to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.
[0067] In order to enable those skilled in the art to more clearly understand the technical solution of the present application, the technical solution of the present application will be described in detail below in conjunction with specific embodiments.
[0068] The test materials used in the embodiments of the present invention are all conventional test materials in the art and can be obtained through commercial channels.
[0069] Example 1: Screening of biomarkers:
[0070] 1. Data acquisition and processing:
[0071] In this embodiment, patients with acute coronary syndrome are selected as the experimental group, and healthy subjects are selected as the control group.
[0072] Recruitment criteria for patients with acute coronary syndrome (ACS): ① New ischemic signs on electrocardiogram, including new ST-T segment changes and new left bundle branch block. ② A significant upward or downward trend in the level of cardiac troponin I (TnI) measured by TnI-Ultra. ③ Coronary angiography shows at least one culprit lesion in a major coronary artery that requires interventional treatment; ④ Using TnI-Ultra detection, the level of cardiac troponin I does not exceed the 99th percentile standard (0.04 ng / mL at baseline) and shows negative for troponin in blood draws at 3 hours or 6 hours.
[0073] Recruitment criteria for the healthy control group (HC): ① No symptoms or signs of atherosclerotic vascular disease; ② No history of angina pectoris; ③ No history of severe liver or kidney disease, bleeding disorder, and malignant disease.
[0074] After obtaining the consent of the patients, 342 patients with acute coronary syndrome in the Chest Pain Center and Department of Cardiology of Jinan Central Hospital from January 2022 to November 2022 were recruited, including 253 males and 89 females, with an average age range of 63.13 ± 12.13; 100 healthy control subjects were recruited, including 77 males and 23 females, with an average age range of 60.84 ± 10.03 (the baseline characteristics of the experimental group and the control group are shown in Table 1).
[0075] All patients were diagnosed by clinicians and underwent coronary angiography and percutaneous coronary intervention. Fasting venous blood samples were collected from all patients on the morning of the second day after percutaneous coronary intervention, and venous blood samples from the control group were collected during the health examination.
[0076] The collected plasma samples were thawed at 4°C, and the metabolites in the plasma samples were extracted using pre-cooled 50% methanol by the peak-point method and mixed. The total ion chromatograms of all samples were obtained using a liquid chromatography system and a mass spectrometer. Subsequently, liquid chromatography-tandem mass spectrometry (LC-MS / MS) was used to qualitatively or quantitatively analyze the extracted plasma samples. The collected plasma samples were analyzed using a non-targeted metabolomics analysis method, and the UHPLC-MS data collected were pre-processed using XCMS software for peak picking, peak grouping, retention time correction, secondary peak grouping, annotation of isotopes and adducts, etc. The LC-MS raw data files were converted to the mzXML format and then processed using the XCMS, CAMERA, and metaX toolboxes and implemented using R software.
[0077] Table 1: Baseline characteristics table
[0078]
[0079] 2. Differential metabolite characteristics:
[0080] Differential analysis was performed on the plasma metabolomics of patients with acute coronary syndrome and healthy controls, and the metabolomic data were integrated to obtain the relevant characteristics of acute coronary syndrome;
[0081] In this case, to study the comprehensive metabolic changes in ACS patients, the plasma metabolomics of 342 ACS patients and 100 HC patients were analyzed. Principal coordinate analysis (PCoA) and heatmap visualization clearly delineated the significant differences between healthy control (HC) subjects and ACS patients ( Figures 1 - 3 ). Using the Benjamini-Hochberg corrected t-test, our study found that the levels of 89 metabolites were higher in ACS patients (false discovery rate (FDR)-adjusted p < 0.05, fold change (FC) > 1.5), and the levels of another 79 metabolites were lower (FDR-adjusted p < 0.05, FC < 0.67). Compared with the healthy control group, the fold changes of 65 different metabolites (including 44 up-regulated and 21 down-regulated metabolites) exceeded 2 or were less than 0.5 (FDR-p < 0.05), highlighting the differences between ACS and HC subjects ( Figure 3 ).
[0082] 3. Model training:
[0083] Five highly interpretable machine learning algorithms are adopted: logistic regression, support vector machine (SVM), decision tree, random tree forest, and XGBoost, and combined with traditional statistical methods (such as t-test and partial least squares discriminant analysis (PLSDA)) for biomarker screening ( Figure 4 ). The "Shapley Additive Explanations (SHAP)" algorithm is used to calculate the contribution of each feature to the model prediction to identify the most influential metabolic biomarkers. Compared with traditional statistical techniques such as the T-test, the interpretability of the importance scores derived from machine learning not only provides the same utility as traditional p-values and volcano plots. Two volcano plots show metabolites with high importance scores and significantly differential expression between the two groups ( Figure 11 ), indicating that these metabolites are crucial for distinguishing healthy and ACS states. Multidimensional information can also be considered when combined with non-linear methods ( Figure 12 、 Figure 13 ). Figure 12 Each point in represents a metabolite, and its size represents the average absolute SHAP value of the metabolite, that is, its average contribution to the model prediction. Figure 13 In each scatter plot, the red points represent the ACS group and the blue points represent the HC group. These plots provide detailed information at the single metabolite level, showing that the metabolites are significantly different between different groups and have high predictive contributions in the model.
[0084] Based on the selected features, a prediction model is constructed, and a five-fold cross-validation strategy is applied to comprehensively evaluate the model, verifying the ability of the model to distinguish ACS patients from healthy controls ( Figure 5 、 Figure 6 ). The results show that these models all show good performance in distinguishing ACS patients from healthy controls, and a set of 11 metabolites is consistently recommended by five machine models to distinguish ACS patients from healthy controls ( Figure 4 ). The 11 metabolic biomarkers are significantly different in different populations ( Figure 7) are in up-regulated or down-regulated states in ACS patients; the 11 metabolic markers are respectively ① D-erythro-Sphinganine-1-phosphate; ② Linoleoyicarnitine; ③ D-erythro-Sphingosine-1-phosphate; ④ Acyicarnitine20:4; ⑤ Glucosereductone; ⑥ AcyicACStine 11:0; ⑦ Acyicarnitine 20:3; ⑧ Acyicarnitine 9:0; ⑨ 5-Hydroxyflavone; ⑩ 2-cis-4-trans-Abscisic acid; Decanoyi-L-carnitine. The 11 metabolites obtained higher interpretability scores and significant odds ratios, and each metabolite was evaluated separately, and all metabolites showed significant discriminatory ability.
[0085] We evaluated each metabolite separately as a potential biomarker for differentiating ACS and HC subjects. The results showed that all these metabolites showed obvious discriminatory ability, and the area under the receiver operating characteristic curve (ROC-AUC) was between 0.84 and 0.97 ( Figure 14 ).
[0086] To achieve the best balance between prediction performance and computational efficiency, a combination of five down-regulated biomarkers was selected, including D-erythro-Sphinganine-1-phosphate, Linoleoyicarnitine, D-erythro-Sphingosine-1-phosphate, Acyicarnitine 20:4, and Glucosereductone. The combination of these five down-regulated metabolites had the best diagnostic effect on ACS, with an ROC-AUC of 0.989 and a PR-AUC of 0.998 ( Figure 8 、 Figure 9 ).
[0087] Example 2: Validation of Biomarkers
[0088] To evaluate and validate the performance of biomarkers composed of metabolites, an independent validation set was established, and the validation samples were 108 ACS patients and 50 healthy controls. The ACS patients and healthy controls were also recruited from the Chest Pain Center and the Department of Cardiology of Jinan Central Hospital.
[0089] Using the same model to distinguish ACS patients, 11 metabolites also showed strong discrimination ability, with the prediction accuracy between 0.68 and 0.90 ( Figure 12 , Figure 14 ), while the biomarker formed by the combination of five metabolites still showed excellent results in screening ACS patients, with the prediction accuracy between 0.899 and 0.956 ( Figure 10 ).
[0090] The above are only the preferred embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. Use of a metabolic biomarker combination in the preparation of a product for diagnosing or warning of acute coronary syndrome, characterized in that: The metabolic biomarker panel consists of D-erythritol-sphingosine-1-phosphate, linoleoylcarnitine, D-erythritol-sphingosine-1-phosphate, acylcarnitine-20:4, and glucose-reductone.
2. Use of a reagent for detecting the metabolic biomarker combination according to claim 1 in the preparation of a product for diagnosing or warning of acute coronary syndrome.
3. The use according to claim 2, characterized in that: The reagent is a reagent for detecting the content of metabolic biomarkers in plasma samples.
Citation Information
Patent Citations
Detection of risk of pre-eclampsia
CN102893156A
Application and method of biomarker in preparation of lung cancer detection reagent
CN116381073A