Blood biomarkers for early coronary heart disease diagnosis and screening methods and applications thereof
By combining blood clinical indicators and metabolomics characteristics, multidimensional blood biomarkers were screened, which solved the problem of low accuracy in existing coronary heart disease diagnosis methods and achieved high sensitivity and high specificity in the early diagnosis of coronary heart disease, making it suitable for primary healthcare institutions.
Patent Information
- Application Number
- CN202411278865.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-12
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2044-09-12
AI Technical Summary
Current methods for diagnosing coronary heart disease have low accuracy, especially in the early stages where myocardial biomarkers are not sensitive enough to effectively screen for and predict the risk of coronary heart disease.
Combining blood clinical indicators and plasma metabolomics characteristics, multidimensional blood biomarkers, including direct bilirubin, aspartate aminotransferase, and high-density lipoprotein, were used for metabolomics analysis via chromatography-mass spectrometry to screen out blood biomarkers with high specificity and sensitivity.
It improves the accuracy and sensitivity of early diagnosis of coronary heart disease, reduces result bias, and provides a non-invasive diagnostic method that is easy to obtain samples, making it suitable for primary healthcare institutions.
Smart Images

Figure CN119335076B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of biomedical technology, and in particular to blood biomarkers for early diagnosis of coronary heart disease, as well as their screening methods and applications. Background Technology
[0002] Coronary heart disease (CHD), commonly known as coronary angina, is the most common type of cardiovascular disease (CVD). It is primarily caused by atherosclerotic lesions in the coronary arteries (the main blood vessels supplying the heart), leading to obstruction or complete blockage of blood flow. This can result in serious consequences such as myocardial ischemia and hypoxia, angina pectoris, myocardial infarction, and even heart failure. With the gradual improvement of people's living standards, the incidence and mortality rates of CHD are on the rise, and the disease is increasingly affecting younger people. According to the latest statistics from the "China Health Statistics Yearbook 2021," the number of CHD patients in my country reached 11.39 million in 2020. The mortality rates for urban and rural residents were 126.01 / 100,000 and 135.88 / 100,000, respectively, with higher mortality rates in men than women in both urban and rural areas. The economic burden of CHD on residents and society is increasingly heavy, making it one of my country's major public health problems.
[0003] Currently, the main clinical diagnostic methods for coronary heart disease include electrocardiogram (ECG), echocardiography, coronary CT angiography (CCTA), coronary angiography (CAG), and myocardial biomarkers. Among them, (1) ECG is the most commonly used method for diagnosing coronary heart disease. It is simple to operate and is particularly suitable for diagnosing patients with arrhythmias. ECG can reflect the patient's cardiac status in a short period of time. If the patient is diagnosed without experiencing angina or arrhythmia, its specificity is low. (2) Echocardiography can examine the morphology, structure, wall motion, and left atrium of the heart. It is one of the commonly used examination methods. Although it has important diagnostic value for ventricular aneurysm, intracardiac thrombosis, cardiac rupture, and papillary muscle function, its accuracy is greatly affected by the subjective experience of the ultrasound examiner. (3) CCTA, as a non-invasive imaging technique, is mainly used to diagnose the degree of epicardial coronary artery stenosis. By measuring the amount of calcium deposits around the coronary arteries, it assesses whether there is stenosis in the coronary artery lumen and has important clinical diagnostic value for suspected coronary heart disease patients. Although CCTA examination has the advantages of being non-invasive, convenient, fast scanning speed and high spatial resolution, it is not suitable for all patients, especially those who are allergic to iodine or contrast agents, have hyperthyroidism or severe heart, liver and kidney dysfunction; (4) CAG is considered the "gold standard" for the diagnosis of coronary heart disease. It can confirm coronary heart disease and patients with complex coronary artery lesions, and provide a definite diagnostic basis for interventional treatment and bypass surgery. Although it has high accuracy and reliable results, and can clearly identify the stenosis of the coronary arteries and its location, degree and extent, it is an invasive examination, which will bring certain examination risks and high costs to patients. Therefore, it is not suitable as a routine means of screening for coronary heart disease; (5) Myocardial biomarkers are important indicators reflecting the biochemical changes of heart disease. For example, C-reactive protein (CRP), troponin (cTn) and myocardial enzyme spectrum examination are the most commonly used and most effective examination indicators for the diagnosis of coronary heart disease. They have important application value for the early diagnosis, disease monitoring and prognosis assessment of coronary heart disease. However, these myocardial biomarkers have relatively low sensitivity, especially for the diagnosis of chronic coronary heart disease.
[0004] Therefore, existing technologies still need to be improved and developed. Summary of the Invention
[0005] In view of the shortcomings of the prior art, the purpose of this invention is to provide blood biomarkers for early diagnosis of coronary heart disease, as well as screening methods and applications thereof, in order to solve the problem of low accuracy in existing methods for diagnosing coronary heart disease using myocardial biomarkers.
[0006] In clinical practice, it is often necessary to combine multiple biomarkers for comprehensive evaluation, and combined indicator analysis has advantages over single-marker testing. Currently, combined blood biochemical indicator testing is a common disease screening item for patients in clinical practice, and it has a certain predictive ability for the diagnosis, screening, and prognosis of coronary heart disease. Blood sample collection is simple, low-cost, provides a large amount of information, and can be repeated, making it suitable for primary healthcare institutions. If blood indicators are combined with metabolomics technology for comprehensive diagnosis and evaluation, it will not only help to determine the optimal treatment time for patients in a timely manner, but also alleviate their physical and psychological suffering. Metabolomics, through qualitative and quantitative analysis of small molecule metabolites in disease states, studies the relationship between metabolites and physiological and pathological changes in diseases, and can provide direct evidence for the pathogenesis of metabolic diseases. Metabolomics can comprehensively analyze multiple small molecule metabolites in the blood samples of patients with coronary heart disease, study the changes in endogenous small molecule metabolites under the influence of internal and external environments, establish metabolic fingerprint profiles between control and disease groups, and find specific metabolic marker combinations related to the disease for early diagnosis and screening. Therefore, this invention aims to provide a set of blood biomarkers based on combined clinical indicators and metabolomics characteristics. Using these blood biomarkers to diagnose early coronary heart disease has advantages such as being non-invasive, highly accurate, and non-invasive.
[0007] The technical solution of the present invention is as follows:
[0008] In a first aspect, the present invention provides a blood biomarker for the early diagnosis of coronary heart disease, wherein the blood biomarker includes blood clinical indicators and plasma metabolomic characteristics, the blood clinical indicators include direct bilirubin, aspartate aminotransferase and high-density lipoprotein, and the plasma metabolomic characteristics include pyroglutamic acid, N-formyl-L-methionine, hypoxanthine, menobiose and L-valine-alanine.
[0009] Optionally, the blood clinical indicators may also include total protein, globulin, urine specific gravity and hematocrit, and the plasma metabolomic characteristics may also include choline.
[0010] Further optionally, the blood clinical indicators also include albumin, red blood cell count, and mean corpuscular hemoglobin concentration, and the plasma metabolomic characteristics also include allantoic acid and xanthine.
[0011] Optionally, the blood biomarker consists of blood clinical indicators and plasma metabolomic characteristics. The blood clinical indicators consist of direct bilirubin, aspartate aminotransferase, and high-density lipoprotein. The plasma metabolomic characteristics consist of pyroglutamic acid, N-formyl-L-methionine, hypoxanthine, menobiose, and L-valine-alanine.
[0012] A second aspect of the present invention provides a method for screening blood biomarkers for early diagnosis of coronary heart disease, comprising the following steps:
[0013] Clinical indicator data of blood routine and blood biochemistry tests were obtained for the coronary heart disease group and the healthy group, respectively;
[0014] Plasma was extracted from both the coronary heart disease group and the healthy group to obtain extracts;
[0015] The aqueous phase in the extract was detected using chromatography-mass spectrometry to obtain detection data;
[0016] The detection data is preprocessed and then metabolites are identified to obtain metabolomics data.
[0017] The clinical indicator data and metabolomics data were then standardized, subjected to t-tests, and correlation analyses to obtain the remaining clinical indicators and metabolomics characteristics.
[0018] Univariate ROC analysis was performed on the remaining clinical indicators and metabolomic characteristics to screen for blood biomarkers for early diagnosis of coronary heart disease.
[0019] A third aspect of the present invention provides the application of the blood biomarker for early diagnosis of coronary heart disease described herein in the preparation of products for diagnosing early coronary heart disease.
[0020] Optionally, the product includes a test reagent or a kit.
[0021] A fourth aspect of the present invention provides a product for diagnosing early coronary heart disease, comprising the blood biomarkers for early coronary heart disease diagnosis described in the present invention.
[0022] Optionally, the product includes a test reagent or a kit.
[0023] Beneficial Effects: This invention not only utilizes plasma metabolomics characteristics but also incorporates clinical blood indicators through routine blood tests and blood biochemistry. Compared to other methods, it more comprehensively reflects the body's physiological and pathological state, reducing the possibility of biased results. Therefore, the use of multidimensional blood biomarkers can more accurately predict the risk of early coronary heart disease. The use of multidimensional blood biomarkers in this invention also improves the specificity and sensitivity of screening and diagnosing early coronary heart disease. Furthermore, the method of diagnosing early coronary heart disease by combining clinical indicators and metabolomics characteristics has advantages such as being non-invasive, easy to obtain samples, and non-invasive. Attached Figure Description
[0024] Figure 1This is a multivariate ROC curve analysis diagram for distinguishing between the healthy group and the coronary heart disease group based on eight blood biomarkers in the modeling group.
[0025] Figure 2 To validate the multivariate ROC curve analysis of the healthy group and the coronary heart disease group based on eight blood biomarkers in the validation group.
[0026] Figure 3 To validate the multivariate ROC curve analysis of the healthy group and the coronary heart disease group based on 13 blood biomarkers in the validation group.
[0027] Figure 4 To validate the multivariate ROC curve analysis of the healthy group and the coronary heart disease group based on 18 blood biomarkers in the validation group. Detailed Implementation
[0028] This invention provides blood biomarkers for early diagnosis of coronary heart disease, their screening methods, and applications. To make the objectives, technical solutions, and effects of this invention clearer and more explicit, the invention is further described in detail below. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0029] This invention provides a blood biomarker for the early diagnosis of coronary heart disease. The blood biomarker includes blood clinical indicators and plasma metabolomic characteristics. The blood clinical indicators include direct bilirubin, aspartate aminotransferase, and high-density lipoprotein. The plasma metabolomic characteristics include pyroglutamic acid, N-formyl-L-methionine, hypoxanthine, menobiose, and L-valine-alanine.
[0030] To address the problems existing in current coronary heart disease (CHD) diagnosis and screening, this invention provides a set of blood biomarkers. These biomarkers not only contain plasma metabolomic characteristics but also incorporate clinical blood indicators through routine blood tests and blood biochemistry analysis. Compared to other methods, this provides a more comprehensive reflection of the body's physiological and pathological state, reducing the possibility of biased results. Therefore, using multidimensional blood biomarkers can more accurately predict the risk of early CHD. The use of multidimensional blood biomarkers for screening and diagnosing early CHD also has high specificity and sensitivity. Furthermore, the method combining clinical indicators and metabolomic characteristics in this invention has advantages such as being non-invasive, easy to obtain samples, and providing effective clinical treatment strategies and guidance for early detection, early diagnosis, and early intervention.
[0031] In one embodiment, the blood biomarker comprises blood clinical indicators and plasma metabolomic characteristics, wherein the blood clinical indicators consist of direct bilirubin, aspartate aminotransferase, and high-density lipoprotein, and the plasma metabolomic characteristics consist of pyroglutamic acid, N-formyl-L-methionine, hypoxanthine, menobiose, and L-valine-alanine.
[0032] Using the aforementioned blood biomarkers, the risk of early coronary heart disease can be accurately predicted.
[0033] In one embodiment, the blood clinical indicators further include total protein, globulin, urine specific gravity, and hematocrit, and the plasma metabolomic characteristics further include choline.
[0034] In other words, the specific blood clinical indicators include direct bilirubin, aspartate aminotransferase, high-density lipoprotein, total protein, globulin, urine specific gravity, and hematocrit, and the specific plasma metabolome characteristics include pyroglutamic acid, N-formyl-L-methionine, hypoxanthine, menobiose, L-valine-alanine, and choline.
[0035] Using the above 13 blood biomarkers, the risk of early coronary heart disease can be accurately predicted.
[0036] In one embodiment, the blood clinical indicators further include albumin, red blood cell count, and mean corpuscular hemoglobin concentration, and the plasma metabolomic characteristics further include allantoic acid and xanthine.
[0037] In other words, the specific blood clinical indicators include direct bilirubin, aspartate aminotransferase, high-density lipoprotein, total protein, globulin, urine specific gravity, hematocrit, albumin, red blood cell count, and mean corpuscular hemoglobin concentration. The specific plasma metabolomics characteristics include pyroglutamic acid, N-formyl-L-methionine, hypoxanthine, menobiose, L-valine-alanine, choline, allantoic acid, and xanthine.
[0038] Using the above 18 blood biomarkers, the risk of early coronary heart disease can be accurately predicted.
[0039] This invention provides a method for screening blood biomarkers for early diagnosis of coronary heart disease, comprising the following steps:
[0040] Clinical indicator data of blood routine and blood biochemistry tests were obtained for the coronary heart disease group and the healthy group, respectively;
[0041] Plasma was extracted from both the coronary heart disease group and the healthy group to obtain extracts;
[0042] The aqueous phase in the extract was detected using chromatography-mass spectrometry to obtain detection data;
[0043] The detection data is preprocessed and then metabolites are identified to obtain metabolomics data.
[0044] The clinical indicator data and metabolomics data were then standardized, subjected to t-tests, and correlation analyses to obtain the remaining clinical indicators and metabolomics characteristics.
[0045] Univariate ROC analysis was performed on the remaining clinical indicators and metabolomic characteristics to screen for blood biomarkers for early diagnosis of coronary heart disease.
[0046] In one implementation, the t-test specifically includes: performing a Student's t-test between the standardized clinical indicator data and metabolomics data, between the healthy group and the coronary heart disease group. t -test), all P value( P Clinical indicator data and metabolomics data with a value (-value) < 0.05 were included in subsequent correlation analysis.
[0047] In one embodiment, the correlation analysis specifically includes: performing Pearson correlation analysis on the clinical indicators and metabolomics features included in the correlation analysis, and then extracting all indicator pairs and feature pairs with a correlation coefficient R value (Pearson correlation coefficient) greater than or equal to the threshold of 0.9, starting from the clinical indicators and metabolomics features that participate the most, until the correlation of any pair of clinical indicators and metabolomics features is less than the threshold of 0.9.
[0048] In one implementation, univariate ROC analysis is performed on the remaining clinical indicators and metabolomic characteristics to screen for blood biomarkers used for early diagnosis of coronary heart disease, specifically including:
[0049] For the remaining clinical indicators and metabolomic characteristics, univariate ROC analysis was performed between the healthy group and the coronary heart disease group to obtain the area under the curve (AUC) values of the clinical indicators and metabolomic characteristics.
[0050] Based on an AUC > 0.73, 10 clinical indicators and 8 metabolomics characteristics were selected, totaling 18 characteristic variables related to coronary artery disease (CAD), which are the blood biomarkers for early CAD diagnosis. Using these 18 blood biomarkers, the risk of early CAD can be accurately predicted.
[0051] In one implementation, from the initial 10 clinical indicators and 8 metabolomics features, 7 more clinical indicators and 6 more metabolomics features are further selected, resulting in a total of 13 features more relevant to coronary heart disease. These are the blood biomarkers used for early diagnosis of coronary heart disease. Using these 13 blood biomarkers, the risk of early coronary heart disease can be accurately predicted.
[0052] In one implementation, univariate ROC analysis is performed on the remaining clinical indicators and metabolomic characteristics to screen for blood biomarkers used for early diagnosis of coronary heart disease, specifically including:
[0053] For the remaining clinical indicators and metabolomic characteristics, univariate ROC analysis was performed between the healthy group and the coronary heart disease group to obtain the area under the curve (AUC) values of the clinical indicators and metabolomic characteristics.
[0054] Based on AUC>0.73, 10 clinical indicators and 8 metabolomics features were selected.
[0055] After combining the data from the selected 10 clinical indicators and 8 metabolomic features into a single dataset, the Lasso regression feature selection method was used to identify the 8 most relevant features for coronary artery disease (CAD), which are then used as blood biomarkers for early CAD diagnosis. These 8 blood biomarkers can accurately predict the risk of early CAD.
[0056] This invention provides an application of the blood biomarker for early coronary heart disease diagnosis described in this invention in the preparation of products for diagnosing early coronary heart disease.
[0057] In one embodiment, the product includes a detection reagent or a kit.
[0058] This invention provides a product for diagnosing early coronary heart disease, comprising the blood biomarkers for early coronary heart disease diagnosis described in this invention.
[0059] In one embodiment, the product includes a detection reagent or a kit.
[0060] The present invention will now be described in detail through specific embodiments.
[0061] Step 1: Information on the test subjects and samples
[0062] Subject Information 1) Inclusion Criteria for Healthy Subjects
[0063] (1) Male or female aged ≥18; (2) Read and fully understand the patient instructions and sign the informed consent form; (3) Exclude subjects with common chronic diseases (hypertension, diabetes, coronary heart disease), history of tumor treatment and major surgery according to the questionnaire; (4) Report no obvious clinical symptoms; (5) Physical examination results are within the normal range and no obvious abnormalities are found; (6) Physical examination results are outside the normal range or have abnormalities, but are judged by the doctor to be of no clinical significance; (7) Subjects who meet all 6 conditions above are included as healthy subjects.
[0064] 2) Inclusion criteria for subjects with coronary heart disease
[0065] (1) Male or female aged ≥18 years; (2) Read and fully understand the patient instructions and sign the informed consent form; (3) After coronary angiography, the degree of coronary artery stenosis is ≥50%; (4) Subjects who meet the above three conditions are included as coronary heart disease subjects.
[0066] 3) Acquisition of clinical data from subjects
[0067] When collecting plasma samples from the subjects, a questionnaire was used to collect information on their basic information and health status, including basic demographic characteristics, lifestyle, medical history, and medication history. Simultaneously, data on complete blood count and blood biochemistry indicators (Table 1) were collected for subsequent statistical analysis and model construction.
[0068] Table 1. Clinical Indicators of Complete Blood Count and Blood Biochemistry Tests
[0069]
[0070] Step 2: Plasma Sample Pretreatment and Polar Metabolite Detection
[0071] 1. Plasma sample pretreatment
[0072] After removing the plasma sample from the -80°C freezer, it was thawed on ice and then vortexed for 10 seconds. Subsequently, 100 μL of the plasma sample was placed in a 1000 μL mixture of pre-chilled methyl tert-butyl ether and methanol, and vortexed to obtain the sample extract. Next, 500 μL of a methanol-water mixture (methanol to water volume ratio 3:1) was added to the sample extract, and the mixture was centrifuged at 12700 rpm for 5 min. After centrifugation and separation, 400 μL of the lower layer was transferred to a centrifuge tube, and 1100 μL of ice-cold methanol was added, followed by vortexing to precipitate proteins. After protein precipitation, the mixture was incubated at 4°C for 1 h, then centrifuged at 12700 rpm for 10 min at 4°C. 1000 μL of the supernatant was transferred to the corresponding centrifuge tube and then evaporated to dryness using a Speed-Vac concentrator. After drying, add 200 μL of water to the centrifuge tube to reconstitute the solution and incubate at room temperature for 15 min. After incubation, vortex the centrifuge tube to mix well, sonicate for 5 min, and then centrifuge at 12700 rpm for 5 min at room temperature. Take 180 μL of the supernatant from the centrifuge tube into a 2 mL glass sample bottle. This is the test solution containing polar metabolites and is then analyzed by LC-MS.
[0073] 2. Detection of polar metabolites
[0074] (1) Liquid chromatography parameters
[0075] Polar metabolites were separated into small molecules using a Waters ACQUTTY UPLC® HSS T3 column (1.8µm, 2.1*100mm). In the column parameters, 1.8µm represents the particle size of the packing material, 2.1*100mm represents the column specifications, 2.1mm is the column inner diameter, and 100mm is the column length. Both liquid chromatography and mass spectrometry were performed using an ACQUITY UPLC I-Class liquid chromatography system (Waters) and a Q-Exactive mass spectrometry system (Thermo Fisher Scientific).
[0076] The composition of the mobile phase is as follows:
[0077] Mobile phase A was an aqueous solution containing 0.1% formic acid; mobile phase B was an acetonitrile solution containing 0.1% formic acid, with a flow rate of 0.4 mL / min. The separation and elution gradient was as follows: 0–1 min, 1% mobile phase B; 1–11 min, 1%–40% mobile phase B; 11–13 min, 40%–70% mobile phase B; 13–15 min, 70%–99% mobile phase B; 15–19 min, 99% mobile phase B; 19–22 min, 1% mobile phase B, where % are volume percentages.
[0078] (2) LC-MS mass spectrometry parameters
[0079] Polar metabolites were acquired using full scan and data-dependent acquisition (DDA) for both MS1 and MS2 mass spectrometry. The full scan mass spectrometry range was 100–1500 Da, while the MS2 mass spectrometry ranges were 100–310 Da, 300–710 Da, and 700–1500 Da, respectively. The Orbitrap high-resolution mass spectrometer (Thermo Fisher) is equipped with an electrospray ionization (ESI) source, acquiring data in both positive and negative ionization modes. The automatic gain control (AGC) target is 3E+6, the maximum IT is 200 ms, the full scan resolution is 70,000 FWHM (@200 m / z), and in the full MS / dd-MS2 mode, the resolution of the secondary mass spectrometer is 17,500 FWHM, the quadrupole window is 1.5 m / z, the AGC is 1E+5, the maximum ion implantation time is 50 ms, the relative collision energy (HCD) is 30 eV, the ion spray voltage is 3,500 V in positive mode and 3,000 V in negative mode, the nebulizer is 20 psi, the sheath gas temperature is 400 °C, and the sheath gas flow rate is 10 L / min.
[0080] Step 3: Metabolomics data preprocessing and metabolite identification
[0081] Metabolomics data preprocessing steps include peak extraction, peak alignment, baseline correction, peak filtering, isotope peak removal, and missing value imputation. Specifically, firstly, peak extraction and alignment are performed on the raw mass spectrometry data. Then, baseline correction is applied to remove noise signals and retain the original signal peaks. The raw data is then converted into centrally discrete data, and isotope peaks in the mass spectrometry data are removed. Next, the retention times of peaks in a single sample are aligned with the peaks in the chromatograms of other samples. Isotope peaks and low-intensity adducts in the same analyte are removed, and characteristic peaks with missing values >50% are eliminated. Missing values are imputed using the MICE Forest chain equation to obtain the final mass spectrometry matrix data.
[0082] Metabolite identification is performed by analyzing raw data using software to obtain spectral information of the parent ion and secondary fragment ions, such as the mass-to-charge ratio (m / z) of the primary mass spectrometer, fragments of the secondary ions, and retention time. The metabolites are then qualitatively identified by matching their spectral information with that of primary and secondary metabolites in databases. Commonly used metabolite databases include HMDB (www.hmdb.ca), PubChem (https: / / pubchem.ncbi.nlm.nih.gov), MassBank (http: / / www.massbank.jp), and MassBank of North America (https: / / massbank.us). Metabolites identified based on these databases are then validated using retention times, MS1, and MS2 mass spectrometry data obtained from separation of standards under the same chromatographic column and mass spectrometry conditions. The criteria for metabolite identification are a retention time difference within 0.1 min and a molecular weight error of less than 10 ppm.
[0083] Step 4: Diagnostic methods for coronary heart disease based on combined clinical indicators and metabolomics characteristics
[0084] The 179 samples used in both the modeling and validation groups were actual samples collected from two different medical centers, and the samples in the modeling and validation groups were different. The specific distribution is as follows: the modeling group had 134 plasma samples, including 77 from healthy individuals and 57 from individuals with coronary heart disease; the validation group had 45 plasma samples, including 26 from healthy individuals and 19 from individuals with coronary heart disease (Table 2). All plasma samples were collected in the morning on an empty stomach, and all collected plasma samples were stored at -80°C.
[0085] Table 2. Sample Information
[0086]
[0087] 2. Blood biomarker screening methods
[0088] 1) Data standardization: Log-transformation was performed on the clinical indicator data and metabolomics data of the modeling group and the validation group respectively to scale the original data to a specific range, eliminating the order-of-magnitude differences between different samples and obtaining a standardized data matrix.
[0089] 2) t-test: For the standardized clinical indicator data and metabolomics data of the modeling group, a t-test was performed between the healthy group and the coronary heart disease group. PClinical indicator data and metabolomics data with a value <0.05 were included in subsequent correlation analysis.
[0090] 3) Correlation Analysis: Highly correlated variables may lead to overfitting in the model and cause significant fluctuations in practical applications. To reduce the risk of overfitting and improve model stability, features with correlations exceeding a preset threshold will be removed, thereby ensuring the model has good discriminative power and stability.
[0091] Specifically:
[0092] |R x,y |≥β, for x,y=1,2,…,n;
[0093] Where the 1st, 2nd, ..., nth features are K1, K2, ..., Kn respectively, and the correlation coefficient between two features is R, R x,y Representing feature K x K y The correlation coefficient between two features K. A threshold β is set based on the correlation coefficient; when two features K... x and K y The correlation coefficient R between them x,y If the absolute value of β is greater than or equal to β, then these two features are considered a highly correlated feature pair. Features that appear in multiple highly correlated feature pairs are preferentially removed.
[0094] Therefore, Pearson correlation analysis was performed on the clinical indicators and metabolomics features of the modeling group. Then, all indicator pairs and feature pairs with an absolute value of correlation coefficient R greater than or equal to 0.9 were extracted and removed starting from the clinical indicators and metabolomics features that participated the most, until the correlation of any pair of clinical indicators and metabolomics features was less than the threshold of 0.9.
[0095] 4) Univariate ROC analysis: After removing some clinical indicators and metabolomics features through t-tests and correlation analysis, univariate ROC analysis was performed between the healthy group and the coronary heart disease group of the modeling group for the remaining clinical indicators and metabolomics features to obtain the area under the curve (AUC) values of the clinical indicators and metabolomics features.
[0096] Based on an AUC > 0.73, 10 clinical indicators (direct bilirubin, aspartate aminotransferase, high-density lipoprotein, total protein, globulin, urine specific gravity, hematocrit, albumin, red blood cell count, and mean corpuscular hemoglobin concentration) and 8 metabolomics features (pyroglutamate, N-formyl-L-methionine, hypoxanthine, menobiose, L-valine-alanine, choline, allantoic acid, and xanthine), totaling 18 characteristic variables, were selected as blood biomarkers. These 18 blood biomarkers can effectively predict the early risk of coronary heart disease.
[0097] Then, from the initial 10 clinical indicators and 8 metabolomics features, 7 more clinical indicators (direct bilirubin, aspartate aminotransferase, high-density lipoprotein, total protein, globulin, urine specific gravity, and hematocrit) and 6 more metabolomics features (pyroglutamate, N-formyl-L-methionine, hypoxanthine, menobiose, L-valine-alanine, and choline) were further selected, totaling 13 characteristic variables more relevant to coronary heart disease, which are the blood biomarkers. These 13 blood biomarkers can also effectively predict the early risk of coronary heart disease.
[0098] 5) Lasso Regression Feature Screening: After combining the data from the selected 10 clinical indicators and 8 metabolomics features into a single dataset, the Lasso regression feature screening method was used to ultimately select 8 feature variables most relevant to coronary heart disease, which are the blood biomarkers used for the construction of early coronary heart disease risk prediction models. These include 3 clinical indicators (direct bilirubin, aspartate aminotransferase, and high-density lipoprotein) and 5 metabolomics features (pyroglutamate, N-formyl-L-methionine, hypoxanthine, menobiose, and L-valine-alanine) (as shown in Table 3).
[0099] Table 3. Eight blood biomarkers distinguishing between the healthy group and the coronary heart disease group
[0100]
[0101] 3. Construction of an early coronary heart disease risk prediction model
[0102] To evaluate the reliability of the eight selected blood biomarkers in distinguishing between healthy and coronary artery disease (CAD) groups, multivariate ROC curve analysis was performed on the eight blood biomarkers. Three-quarters of the samples from the modeling group were randomly selected as the training set, and the remaining one-quarter as the test set. A support vector machine (SVM) was randomly iterated 1000 times. An early CAD risk prediction model was constructed by statistically analyzing the average accuracy of the final model. Figure 1The results showed an AUC of 0.977 (sensitivity: 92.4%, specificity: 94.7%), indicating that the constructed model has high predictive efficacy. Generally, model performance is evaluated based on the area under the ROC curve (AUC), typically between 0.5 and 1. A UC value closer to 1 indicates better model performance and diagnostic effectiveness; an AUC value less than 0.5 indicates poor model accuracy. Sensitivity and specificity reflect the model's ability to identify positive and negative samples, respectively. Sensitivity, also known as the true positive rate (TPR), represents the probability that the model correctly diagnoses a positive case; specificity, also known as the true negative rate (TNR), represents the probability that the model correctly diagnoses a negative case in the control group. Higher sensitivity indicates a stronger ability to identify positive samples, while higher specificity indicates a stronger ability to distinguish negative samples.
[0103] 4. Validation of the early coronary heart disease risk prediction model
[0104] To validate the model's effectiveness, only the dataset containing the eight selected blood biomarkers was retained from the standardized dataset of the validation group. This dataset was then used as an unknown sample and fed into the early coronary artery disease risk prediction model constructed using the modeling group. Figure 2 The results showed that the AUC = 0.984 (sensitivity: 89.5%, specificity: 96.2%), indicating that the early coronary heart disease risk prediction model based on the combination of eight clinical indicators and metabolomics characteristics also has good predictive ability in the validation dataset.
[0105] Based on the 10 selected clinical indicators and 8 metabolomics features, 13 more relevant characteristics to coronary artery disease (CAD) were further screened, which were then identified as blood biomarkers. Using the SVM method, an early CAD risk prediction model was constructed in the modeling group sample for these 13 blood biomarkers. Then, the validation group sample, as unknown samples, was added to the prediction model to verify the effectiveness of different blood biomarkers in predicting early CAD risk. Figure 3 The results showed that the AUC was 0.989 (sensitivity: 85.4%, specificity: 99.8%), indicating that the early coronary heart disease risk prediction model based on a combination of 13 blood biomarkers also has good predictive ability.
[0106] Based on metabolomics and clinical indicators, 10 clinical indicators and 8 metabolomics features were initially screened using univariate AUC. For these 18 blood biomarkers, a SVM method was used to construct an early coronary artery disease risk prediction model in the modeling group sample. Then, the validation group sample, as unknown samples, was added to the prediction model to verify the effectiveness of different blood biomarkers in predicting the early coronary artery disease risk. Figure 4 The results showed an AUC of 0.996 (sensitivity: 89.5%, specificity: 96.2%), indicating that the early coronary heart disease risk prediction model based on a combination of 18 blood biomarkers also has good predictive ability. In summary, the blood biomarkers, screening methods, and applications for early coronary heart disease diagnosis provided by this invention not only utilize plasma metabolomics characteristics but also incorporate clinical blood indicators through routine blood tests and blood biochemistry. Compared to other methods, this more comprehensively reflects the body's physiological and pathological state, reducing the possibility of biased results. Therefore, using multidimensional blood biomarkers can more accurately predict the risk of early coronary heart disease. This invention, using multidimensional blood biomarkers, can also improve the specificity and sensitivity of screening and diagnosing early coronary heart disease. Furthermore, the method of diagnosing early coronary heart disease by combining clinical indicators and metabolomics characteristics has advantages such as being non-invasive, easy to obtain samples, and non-invasive.
[0107] It should be understood that the application of the present invention is not limited to the examples above. Those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.
Claims
1. A method for screening of blood biomarkers for early coronary heart disease diagnosis, characterized by, It comprises the following steps: Clinical index data of blood routine and blood biochemical detection of the coronary heart disease group and the healthy group are respectively acquired; Plasma of the coronary heart disease group and the healthy group is respectively extracted to obtain an extract; Water phase in the extract is detected by using a chromatograph-mass spectrometer to obtain detection data; The detection data is pretreated and then metabolite identification is performed to obtain metabolomics data; The clinical index data and the metabolomics data are sequentially subjected to standardization processing, t test and correlation analysis to obtain remaining clinical indexes and metabolomics characteristics; Univariate ROC analysis is performed on the remaining clinical indexes and metabolomics characteristics to screen blood biomarkers for early coronary heart disease diagnosis; The blood biomarkers comprise blood clinical indexes and plasma metabolomics characteristics, the blood clinical indexes comprise direct bilirubin, aspartate aminotransferase and high-density lipoprotein, and the plasma metabolomics characteristics comprise pyroglutamic acid, N-formyl-L-methionine, hypoxanthine, pinobetsin and L-valyl-alanine.
2. The screening method of blood biomarkers for early coronary artery disease diagnosis according to claim 1, characterized in that, The blood clinical indexes further comprise total protein, globulin, urine specific gravity and hematocrit, and the plasma metabolomics characteristics further comprise choline.
3. The method for screening of blood biomarkers for early coronary heart disease diagnosis according to claim 2, characterized in that, The blood clinical indexes further comprise albumin, red blood cell count and mean hemoglobin concentration, and the plasma metabolomics characteristics further comprise allantoic acid and xanthine.
4. The method for screening of blood biomarkers for early diagnosis of coronary heart disease according to claim 1, characterized in that, The blood biomarkers are composed of blood clinical indexes and plasma metabolomics characteristics, the blood clinical indexes are composed of direct bilirubin, aspartate aminotransferase and high-density lipoprotein, and the plasma metabolomics characteristics are composed of pyroglutamic acid, N-formyl-L-methionine, hypoxanthine, pinobetsin and L-valyl-alanine.
5. Application of a blood biomarker for early coronary heart disease diagnosis in the preparation of a product for diagnosing early coronary heart disease, wherein the blood biomarker comprises blood clinical indexes and plasma metabolomics characteristics, the blood clinical indexes comprise direct bilirubin, aspartate aminotransferase and high-density lipoprotein, and the plasma metabolomics characteristics comprise pyroglutamic acid, N-formyl-L-methionine, hypoxanthine, pinobetsin and L-valyl-alanine. The blood clinical indexes further comprise total protein, globulin, urine specific gravity and hematocrit, and the plasma metabolomics characteristics further comprise choline.
6. Use of the blood biomarkers for early coronary artery disease diagnosis according to claim 5 for the preparation of a product for the diagnosis of early coronary artery disease, characterized in that, The blood clinical indexes further comprise albumin, red blood cell count and mean hemoglobin concentration, and the plasma metabolomics characteristics further comprise allantoic acid and xanthine.
7. Use of the blood biomarkers for early coronary artery disease diagnosis according to claim 5 for the preparation of a product for the diagnosis of early coronary artery disease, characterized by, The blood biomarkers are composed of blood clinical indexes and plasma metabolomics characteristics, the blood clinical indexes are composed of direct bilirubin, aspartate aminotransferase and high-density lipoprotein, and the plasma metabolomics characteristics are composed of pyroglutamic acid, N-formyl-L-methionine, hypoxanthine, pinobetsin and L-valyl-alanine.
8. Use of the blood biomarkers for early coronary artery disease diagnosis according to claim 5 for the preparation of a product for the diagnosis of early coronary artery disease, characterized in that, The product comprises a detection reagent or a kit.
9. Use of the blood biomarkers for early coronary artery disease diagnosis according to claim 5 for the preparation of a product for the diagnosis of early coronary artery disease, characterized in that, 10. A product for diagnosing early coronary heart disease, characterized in that, Blood biomarkers for early coronary heart disease diagnosis, comprising blood clinical indicators including direct bilirubin, aspartate aminotransferase, and high-density lipoprotein, and plasma metabolome features including pyroglutamic acid, N-formyl-L-methionine, hypoxanthine, pinobanksin, and L-valyl-alanine.
11. A product for diagnosing early coronary heart disease according to claim 10, characterized in that, The blood clinical indicators further include total protein, globulin, specific gravity, and hematocrit, and the plasma metabolome features further include choline.
12. The product for diagnosing early coronary artery disease according to claim 10, wherein The blood clinical indicators further include albumin, red blood cell count, and mean corpuscular hemoglobin concentration, and the plasma metabolome features further include allantoic acid and xanthine.
13. The product for diagnosing early coronary artery disease according to claim 10, wherein The blood biomarkers consist of blood clinical indicators consisting of direct bilirubin, aspartate aminotransferase, and high-density lipoprotein, and plasma metabolome features consisting of pyroglutamic acid, N-formyl-L-methionine, hypoxanthine, pinobanksin, and L-valyl-alanine.
14. The product for diagnosing early coronary artery disease according to claim 10, wherein The products include detection reagents or kits.
Citation Information
Patent Citations
Colorectal cancer diagnosis biomarker and application thereof
CN117929749A
Metabolic marker for diagnosing coronary artery stenosis degree and application thereof
CN118409095A