Kit for early diagnosis of young acute coronary syndrome based on plasma transcriptomics and application thereof
By analyzing the expression levels of specific proteins in plasma and combining machine learning methods, a risk scoring model for acute coronary syndrome in young people was established, which solved the problem of insufficient existing diagnostic methods and achieved early, accurate diagnosis and personalized treatment of acute coronary syndrome in young people.
Patent Information
- Application Number
- CN202510680127.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-09-19
AI Technical Summary
Existing diagnostic methods for acute coronary syndrome in young people mainly rely on clinical research and experience in middle-aged and elderly people, which makes them difficult to apply to young people. In addition, there is a lack of large-scale, multi-omics research, resulting in insufficient diagnostic standards and treatment plans, and an inability to effectively address the high mortality and disability rates of acute coronary syndrome in young people.
By analyzing the expression levels of twelve proteins in plasma, including FAM178B, ARVCF, SLC6A9, BNC2, DKK2, CLSTN3, SEMA6B, DAB2IP, LUC7L, NIPAL4, RBFOX3, and ST18, and combining random forest, LASSO regression, and XGBOOST machine learning methods, a risk scoring model for acute coronary syndrome in young people was established, providing a non-invasive and rapid early diagnosis method.
It has achieved early and accurate diagnosis of acute coronary syndrome in young people, improved the sensitivity and specificity of diagnosis, reduced mortality and disability rates, and provided personalized treatment plans.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biomedicine, and in particular to a plasma transcriptomics-based early diagnosis kit for acute coronary syndrome in young people and a preparation and application method thereof. Background Art
[0002] Acute coronary syndrome (ACS) is a group of clinical syndromes caused by a sudden decrease in coronary blood flow, including unstable angina (UA), non-ST-segment elevation myocardial infarction (NSTEMI), and ST-segment elevation myocardial infarction (STEMI). The primary pathological mechanism is the rupture or erosion of coronary atherosclerotic plaques, followed by thrombosis, leading to partial or complete obstruction of the coronary arteries, resulting in myocardial ischemia and necrosis. ACS is highly prevalent worldwide, particularly in industrialized countries. Statistics show that ACS causes millions of deaths annually and is one of the leading causes of death from cardiovascular disease. The high mortality and disability rates of ACS are primarily due to its sudden onset, rapid progression, and the severe complications such as heart failure and arrhythmias that can ensue from extensive myocardial necrosis. Studies have shown that ACS costs the United States over $400 billion annually in direct medical costs (such as hospitalizations, surgeries, and medications) and indirect economic losses (such as lost labor). In my country, the average hospitalization cost per visit for coronary heart disease and ischemic heart disease in 2018 was 13,083.90 yuan, including 28,879.30 yuan for acute myocardial infarction. Adjusting for inflation, the average annual growth rate for AMI has been 6.09% since 2004. Acute coronary syndrome not only poses a serious threat to the health of individual patients but also imposes a heavy economic burden on society.
[0003] In recent years, the incidence of acute coronary syndrome (ACS) in young adults (aged 18-45 years) has been increasing, a phenomenon that has attracted widespread attention. ACS has traditionally been considered a common disease among middle-aged and elderly individuals. However, a growing number of studies suggest that young adults are also at increased risk of ACS. Traditional ACS risk factors, such as hypertension, diabetes, hyperlipidemia, obesity, environmental pollution, and smoking, are increasing in prevalence among young adults, potentially contributing to the rising incidence of ACS in young adults. However, while the early onset of risk factors is associated with premature ACS, the specific pathological mechanisms remain unclear. Multiple factors, including genetics, environmental factors, lifestyle, and psychological stress, may contribute to the development and progression of ACS in young adults.
[0004] Currently, the diagnosis and treatment of acute coronary syndrome in young people are primarily based on clinical research and experience with acute coronary syndrome in middle-aged and elderly people. However, significant differences may exist between the etiology, course, and clinical manifestations of acute coronary syndrome in middle-aged and elderly people and young people. Young patients with acute coronary syndrome often experience sudden myocardial infarction without obvious premonitory symptoms, often accompanied by severe coronary artery spasm and thrombosis. Therefore, targeted research and adjustments are urgently needed for the diagnostic criteria and treatment plans for acute coronary syndrome in young people. However, due to the relatively low incidence of acute coronary syndrome in young people and the lack of relevant research, existing clinical guidelines and treatment plans are mostly based on data from middle-aged and elderly people, making them difficult to fully apply to young patients.
[0005] Currently, large-scale, prospective, multi-omic studies of ACS in young adults are relatively scarce. Multi-omic approaches, including genomics, transcriptomics, proteomics, and metabolomics, can reveal disease pathogenesis and biomarkers from multiple levels and perspectives. Although some small-scale studies have attempted to use multi-omic approaches to explore ACS in young adults, limited sample size, study design, and technical means have yet to yield systematic research results. Furthermore, the high heterogeneity and complex etiology of ACS in young adults pose even greater challenges to research. Therefore, large-scale, multicenter, long-term follow-up, prospective, multi-omic studies are urgently needed to fully uncover the pathogenesis and potential therapeutic targets of ACS in young adults.
[0006] Acute coronary syndrome (ACS) is a disease caused by myocardial ischemia or necrosis due to stenosis or obstruction of the coronary arteries. It is a major type of cardiovascular disease. It is one of the leading causes of death and disability worldwide and poses a major threat to public health. The hazards of coronary heart disease include: angina pectoris and myocardial infarction: patients may experience angina pectoris after ACS, and the myocardium may repeatedly suffer from ischemia and dysfunction. Patients often feel chest pain and discomfort during activities, affecting their daily life and work ability. In severe cases, acute myocardial infarction (AMI) may occur. AMI is one of the most serious manifestations of coronary heart disease. It is usually caused by complete blockage of the coronary arteries, leading to myocardial necrosis, and can be fatal in severe cases. After myocardial ischemia, patients will experience a decrease in effectively contracting myocardial cells, a decline in quality of life, and heart failure, resulting in a huge public health burden. With the development of molecular biology and genomics, the application of transcriptomics in the early diagnosis of diseases has gradually attracted attention. By analyzing the RNA expression profile in the blood, non-invasive and accurate disease diagnosis and prognosis assessment can be achieved. The necessity and advantages of transcriptomics in the detection of coronary heart disease include: Early diagnosis: Transcriptomics can detect changes in gene expression before the onset of coronary heart disease symptoms, achieve early diagnosis of the disease, help early intervention, and reduce mortality and disability rates. Non-invasive detection: RNA information in whole blood or plasma can be obtained through simple blood sample collection, avoiding the risks and discomfort brought by invasive examinations, and improving patient acceptance and compliance. High throughput and comprehensiveness: Transcriptomics technology can detect the expression levels of thousands of genes at the same time, providing comprehensive molecular information, which helps to reveal the complex mechanisms of the disease.
[0007] The application of machine learning in clinical medicine is becoming increasingly widespread and important. By processing large amounts of complex biomedical data, machine learning methods can help improve disease diagnosis, prognosis, and treatment outcomes. This article introduces three commonly used machine learning methods: Random Forest, LASSO regression, and XGBOOST (Extreme Gradient Boosting), their respective strengths and weaknesses, and the advantages of combining these three methods for feature learning.
[0008] Random Forest: Random forest is an ensemble learning method that improves model accuracy and stability by building multiple decision trees and combining their predictions. Each decision tree is trained on a different subsample, and only a subset of features is considered at each split. Advantages: 1) Robustness: Random forests are insensitive to noise and outliers in the data. 2) Overfitting prevention: By combining multiple decision trees, random forests can effectively reduce overfitting. 3) High-dimensional data processing: Random forests can handle datasets with a large number of features and assess feature importance. Disadvantages: 1) High computational cost: Because a large number of decision trees must be built, both the training and prediction processes require high computing resources. 2) Poor interpretability: Although feature importance can be assessed, the combined results of individual decision trees are difficult to interpret.
[0009] LASSO (Least Absolute Shrinkage and Selection Operator) regression is a linear regression method that uses the L1 regularization term to select features and prevent overfitting. LASSO regression can reduce certain regression coefficients to zero, thereby achieving feature selection. Advantages: 1) Feature Selection: It can automatically select important features, reduce model complexity, and prevent overfitting. 2) Processing high-dimensional data): It is particularly suitable for processing datasets with more features than the number of samples. Disadvantages: 1) Selection bias: LASSO may ignore some important but highly correlated features. 2) Linear assumption: LASSO assumes that the relationship between features and target variables is linear, which limits its performance when dealing with nonlinear relationships.
[0010] XGBOOST is a boosting algorithm that builds a series of weak learners (typically decision trees), each improving upon the previous one to continuously reduce prediction error. XGBOOST improves model performance and computational efficiency through parallel processing and regularization techniques. Advantages: 1) Efficiency: XGBOOST utilizes parallel computing and optimization algorithms, resulting in fast training speed and high resource utilization. 2) High Accuracy: XGBOOST has performed well in many competitions and real-world applications, achieving high prediction accuracy. 3) High Flexibility: It can handle various data types (such as numerical and categorical data) and supports custom loss functions. Disadvantages: 1) Complex Parameter Tuning: XGBOOST has many hyperparameters, requiring extensive tuning to achieve optimal performance. 2) Poor Interpretability: Similar to random forests, while feature importance can be assessed, the overall model exhibits low interpretability.
[0011] In clinical medicine, a single machine learning method may not fully capture the complexity and diversity of data. Therefore, combining random forest, LASSO regression, and XGBOOST for feature learning can leverage their respective strengths to improve model accuracy and robustness. Advantages of combining these three machine learning methods for feature learning: 1) Diversity in feature selection: LASSO regression selects important features through L1 regularization, removing redundant and irrelevant features. Random forest assesses feature importance, providing a feature selection method based on nonlinear relationships. XGBOOST, combined with a boosting algorithm, can capture complex feature interactions and further optimize feature selection. 2) Improved model stability and robustness: Random forest and XGBOOST use ensemble learning methods to reduce the risk of overfitting in a single model and improve model robustness. LASSO regression uses regularization techniques to further reduce model complexity and overfitting. 3) Adaptability to diverse data characteristics: Combining these methods can handle high-dimensional data, nonlinear relationships, and different types of data, improving model adaptability and generalization. 4) Model interpretability and visualization: Feature importance evaluation using random forest and XGBOOST, combined with coefficient analysis using LASSO regression, can provide multi-angle feature interpretations to help clinicians better understand the model's prediction results.
[0012] The application of machine learning methods in clinical medicine provides powerful tools for early disease diagnosis, prognosis prediction, and treatment optimization. By combining random forest, LASSO regression, and XGBOOST, we can leverage their respective strengths to optimize feature selection and model building, improving the accuracy and robustness of predictive models and thus better serve clinical practice.
[0013] Existing methods for diagnosing acute coronary syndrome (ACS) have the following major drawbacks: 1. A physician's medical history review can only provide a reference and cannot accurately determine whether an ACS is present, especially in elderly women and diabetic patients, who often do not have typical angina symptoms. 2. The electrocardiogram (ECG) is the first-line diagnostic tool for ACS, but changes can completely or partially disappear with relief of angina, necessitating continuous ECG monitoring. 3. Measurements of myocardial injury markers (cTnT, cTnI, or CK-MB) vary in sensitivity and specificity, and each has its own diagnostic window, often requiring the combination of multiple indicators and continuous monitoring. 4. Coronary angiography is an invasive procedure. Furthermore, the angiography equipment is extremely inconvenient to move, and the cost of each measurement is high. Therefore, an improved product for the early diagnosis of ACS in young people is currently in demand. Summary of the Invention
[0014] The purpose of the present invention is to provide a non-invasive detection method and kit for early diagnosis of acute coronary syndrome in young people. Based on the expression level of protein transcription factors in plasma, by detecting the abnormal expression of these transcription factors in patients with acute coronary syndrome, a rapid, accurate and non-invasive detection method is provided for early diagnosis, which has high clinical application value. The present invention provides a computer device, comprising a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the following steps: S1) Data reception: Receive the expression levels of twelve proteins, including FAM178B, ARVCF, SLC6A9, BNC2, DKK2, CLSTN3, SEMA6B, DAB2IP, LUC7L, NIPAL4, RBFOX3, and ST18, in plasma samples from patients to be tested; S2) Data processing: Substituting the expression levels of the twelve proteins into the prognostic risk prediction model to calculate a risk score value for the patient to be tested as a high-risk young adult with coronary heart disease; the risk prediction model for the patient to be tested as a high-risk young adult with coronary heart disease is as follows: Risk score = Y = 4.98918 + (FAM178B * (-0.09285)) + (ARVCF * 0.00261) + (SLC6A9 * (-0.00406)) + (BNC2 * (-0.00243)) + (DKK2 * (-0.06780)) + (CLSTN3 * 0.00041) + (SEMA6B * (0.00385)) + (DAB2IP * (-0.00894)) + (DAB2IP * (-0.00894)) + (NIPAL4 * 0.02325) + (RBFOX3 * (-0.03305)) + (ST18 * (-0.07321)) Formula 1, S3) Data output: Analyze the possibility that the patient to be tested is a high-risk young coronary heart disease patient based on the risk score value.
[0015] Among them, FAM178B, ARVCF, SLC6A9, BNC2, DKK2, CLSTN3, SEMA6B, DAB2IP, LUC7L, NIPAL4, RBFOX3, and ST18 in the formula all represent the expression levels of proteins.
[0016] Furthermore, in step S3), the possibility that the patient to be tested with a risk score ≥0.669 is a high-risk young patient with coronary heart disease is greater than the possibility that the patient to be tested with a risk score <0.669 is a high-risk young patient with coronary heart disease.
[0017] The present invention provides a computer-readable storage medium having a computer program / instruction stored thereon, which implements the steps of the above method when the computer program / instruction is executed by a processor.
[0018] Furthermore, the computer-readable storage medium refers to a carrier for storing data, which may be a floppy disk, an optical disk, a DVD, a hard disk, a flash memory, a USB flash drive, a CF card, an SD card, an MMC card, an SM card, a memory stick (Memory Stick) or an xD card.
[0019] The present invention provides a computer program product, comprising a computer program, which, when executed by a processor, implements the following steps: S1) Data reception: Receive the expression levels of twelve proteins, including FAM178B, ARVCF, SLC6A9, BNC2, DKK2, CLSTN3, SEMA6B, DAB2IP, LUC7L, NIPAL4, RBFOX3, and ST18, in plasma samples from patients to be tested; S2) Data processing: Substituting the expression levels of the twelve proteins into the prognostic risk prediction model to calculate a risk score value for the patient to be tested as a high-risk young adult with coronary heart disease; the risk prediction model for the patient to be tested as a high-risk young adult with coronary heart disease is as follows: Risk score = Y = 4.98918 + (FAM178B * (-0.09285)) + (ARVCF * 0.00261) + (SLC6A9 * (-0.00406)) + (BNC2 * (-0.00243)) + (DKK2 * (-0.06780)) + (CLSTN3 * 0.00041) + (SEMA6B * (0.00385)) + (DAB2IP * (-0.00894)) + (DAB2IP * (-0.00894)) + (NIPAL4 * 0.02325) + (RBFOX3 * (-0.03305)) + (ST18 * (-0.07321)) Formula 1, S3) Data output: Analyze the possibility that the patient to be tested is a high-risk young coronary heart disease patient based on the risk score value.
[0020] Furthermore, in step S3), the possibility that the patient to be tested with a risk score ≥0.669 is a high-risk young patient with coronary heart disease is greater than the possibility that the patient to be tested with a risk score <0.669 is a high-risk young patient with coronary heart disease.
[0021] Any of the following applications of proteins and / or substances for detecting said proteins should also fall within the scope of protection of the present invention: D1) Use in the preparation of products for predicting or assisting in predicting the likelihood that a patient to be tested is a high-risk young adult with coronary heart disease; D2) Application in devices for predicting or assisting in predicting the likelihood that a patient to be tested is a high-risk young adult with coronary heart disease; The proteins are FAM178B, ARVCF, SLC6A9, BNC2, DKK2, CLSTN3, SEMA6B, DAB2IP, LUC7L, NIPAL4, RBFOX3, and ST18.
[0022] Furthermore, the substance is a reagent, a kit and / or an instrument.
[0023] The present invention is based on a comparison of the expression levels of proteins FAM178B, ARVCF, SLC6A9, BNC2, DKK2, CLSTN3, SEMA6B, DAB2IP, LUC7L, NIPAL4, RBFOX3, and ST18 in the peripheral venous blood plasma of young patients with acute coronary syndrome and young patients with stable angina pectoris and normal coronary arteries. It was found that the content of target plasma transcription factors in samples of patients with acute coronary syndrome was different from that in patients with young stable angina pectoris and normal coronary arteries, indicating that the target transcript panel detection can be used as a marker combination for young acute coronary syndrome. The present invention provides a kit and device for early diagnosis of acute coronary syndrome in young patients. By detecting the transcriptional expression levels of proteins FAM178B, ARVCF, SLC6A9, BNC2, DKK2, CLSTN3, SEMA6B, DAB2IP, LUC7L, NIPAL4, RBFOX3, and ST18 in the peripheral venous blood plasma of young patients, it can be quickly determined whether the patient has acute coronary syndrome. The detection is non-invasive, the diagnosis is rapid and accurate, the operation is simple, and the repeatability is good. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 The plasma transcriptome of young patients with coronary heart disease consecutively enrolled in our center screened for transcriptional markers related to acute coronary syndrome in young people.
[0025] Figure 2 The performance of three machine learning algorithms in the validation set (A: Random Forest; C: LASSO; E: XGBoost), and the initial screening of feature variables based on different algorithms (B: variables selected based on Random Forest; D: variables selected based on LASSO algorithm; F: variables selected based on XGBoost algorithm).
[0026] Figure 3The expression levels of the 12 transcripts FAM178B, ARVCF, SLC6A9, BNC2, DKK2, CLSTN3, SEMA6B, DAB2IP, LUC7L, NIPAL4, RBFOX3, and ST18 that finally entered the model in the training set (A) and validation set (B), as well as the significant differences between different groups. Cyan represents the non-ACS group, and yellow represents the ACS group. P < 0.05 indicates a significant difference.
[0027] Figure 4 Figure 3 ROC (A: training set ROC; B: validation set ROC) and prediction status (C: training set prediction probability visualization diagram, cyan marks represent ACS patients, more ACS patients have higher risk scores and are located on the upper right side of the diagram; red marks represent non-ACS patients, most of the non-ACS patients have lower risk scores and are located on the lower side of the diagram. D is the training set and validation set visualization prediction probability diagram).
[0028] Figure 5 The PR curves (A, B) and calibration curves (C, D) in the training set and validation set indicate that the risk prediction model of the above transcript combination has good accuracy and the predicted situation is highly consistent with the actual situation.
[0029] Figure 6 The DCA curves (A, B) and CIC (C, D) in the training set and validation set indicate that the risk prediction model of the above transcript combination has a good guiding role in clinical prediction and patient condition selection.
[0030] Figure 7 This is a nomogram of the risk prediction model, visually displaying the integral of different transcription factors in ACS risk assessment. This is more helpful for clinicians to manage and predict risk.
[0031] Figure 8 The forest plot of the 12 transcription factors used in the final model construction indicates the contribution of different transcripts to the risk. DETAILED DESCRIPTION
[0032] The present invention will be further described in detail below in conjunction with specific embodiments. The examples provided are only for illustrating the present invention and are not intended to limit the scope of the present invention. The examples provided below can serve as a guide for further improvements by those skilled in the art and are not intended to limit the present invention in any way.
[0033] Unless otherwise specified, the experimental methods in the following examples are conventional methods and were performed according to the techniques or conditions described in the literature in the field or according to the product instructions. The materials and reagents used in the following examples, unless otherwise specified, were all commercially available.
[0034] Unless otherwise specified, the quantitative tests in the following examples were performed three times, and the results were averaged. Summary of the invention: In order to reveal the pathogenesis and internal laws of acute coronary syndrome in young people, the present invention conducted a large-scale prospective multi-omics exploratory study. The research team conducted a systematic analysis of a large number of samples from young patients with acute coronary syndrome from multiple levels such as genomics, transcriptomics, proteomics and metabolomics. The study found that some specific gene mutations, increased levels of inflammatory factors and metabolic disorders may be closely related to the occurrence of acute coronary syndrome in young people. In particular, a series of potential diagnostic markers for acute coronary syndrome were identified. The expression of these markers in young patients with acute coronary syndrome was significantly abnormal and had a high diagnostic value. This research result provides new ideas and methods for the early diagnosis and personalized treatment of acute coronary syndrome in young people, and lays the foundation for further clinical research and application in the future.
[0036] The present invention uses three machine learning methods, random forest, LASSO regression, and XGBOOST, to take the union set to screen potential candidate transcription factors for predicting the onset of ACS in young people, and then performs secondary screening through classic multivariate regression. Finally, 12 transcription factors (FAM178B, ARVCF, SLC6A9, BNC2, DKK2, CLSTN3, SEMA6B, DAB2IP, LUC7L, NIPAL4, RBFOX3, ST18) are selected to establish a prediction model for predicting ACS in young people, achieving excellent predictive efficiency: training set AUC 0.9672, 95% CI (0.9435-0.9908), specificity 0.9482, sensitivity 0.8929, accuracy 0.9154, positive predictive value 0.96115, negative predictive value 0.8594, and Youden index 0.8411. The validation set AUC was 0.8462, 95%CI (0.7958-0.9589), specificity was 0.8461, sensitivity was 0.7568, accuracy was 0.7936, positive predictive value was 0.8750, negative predictive value was 0.7097, and Youden index was 0.6029.
[0037] 1) FAM178B (family with sequence similarity 178 member B, Gene ID: 51252, updated on 10-Dec-2024): FAM178B is a transcription factor that plays a role in various biological processes. Although limited research directly links FAM178B to cardiovascular disease, the protein encoded by this gene has been linked to bipolar disorder, body mass index (BMI), and cell adhesion. However, studies in gliomas have revealed a potential role for this gene in cell proliferation and apoptosis. Aberrant expression of FAM178B has been associated with the development and progression of certain cancers, suggesting that it may play a similar role in the pathogenesis of cardiovascular disease.
[0038] 2) ARVCF (ARVCF delta catenin family member, Gene ID: 421, updated on 4-Jan-2025): A member of the delta catenin family. In one study, mutations in FAM178B were associated with hypertrophic cardiomyopathy. These mutations could be repaired using the base editing technology ABEmax-NG, suggesting that FAM178B may play a role in the pathogenesis of cardiomyopathy. This family plays a key role in the formation of adherens junction complexes, which are believed to facilitate communication between the intracellular and extracellular environments. Studies have reported that ARVCF plays a role in the pathological mechanisms of diabetes-induced coronary artery calcification.
[0039] 3) SLC6A9 (Solute Carrier Family 6 Member 9, Gene ID: 6536, updated on 4-Jan-2025) is an amino acid transporter protein belonging to the solute carrier 6 (SLC6) family. SLC6A9 is primarily involved in amino acid transport, particularly glycine. Using membrane potential as a driving force, it mediates the transport of neutral amino acids from the extracellular environment into the intracellular space or into storage vesicles, primarily involved in neurotransmitter transport. While its role in cardiovascular disease remains unclear, studies in neurological disorders such as schizophrenia suggest that this gene may indirectly affect cardiovascular health by influencing nervous system function.
[0040] 4) Transcriptional induction of BNC2 (basonuclin zinc finger protein 2, Gene ID: 54796, updated on 5-Jan-2025) is a specific feature of myofibroblast activation in fibrotic tissue. Mechanistically, BNC2 expression and activity allow for the integration of profibrotic stimuli, including TGFβ and Hippo / YAP1 signaling, to induce matrisome genes, such as those encoding type I collagen, which play a key role in cell differentiation and proliferation. Although BNC2 is less well-studied in cardiovascular disease, its profibrotic role suggests that it may also play an important role in cardiovascular disease pathogenesis.
[0041] 5) DKK2 (dickkopf WNT signaling pathway inhibitor 2, Gene ID: 27123, updated on 4 January 2025) is an inhibitor of the WNT signaling pathway. Dkk2 knockdown significantly reduced genes associated with classical M1-polarized macrophages but increased markers of alternative M2-polarized macrophages. Furthermore, Dkk2 silencing significantly attenuated foam cell formation, as evidenced by increased expression of markers associated with cholesterol efflux and decreased expression of markers of cholesterol influx. Downregulated Dkk2 expression in macrophages may contribute to macrophage inactivation by targeting the Wnt / β-catenin pathway, thereby impacting cardiovascular disease progression.
[0042] 6) CLSTN3 (calsyntenin 3, Gene ID: 9746, updated on 4-Jan-2025): CLSTN3 has been found to be involved in lipid energy metabolism. Studies have shown that overexpression of CLSTN3 can improve lipid metabolism disorders, gluconeogenesis, and energy homeostasis, and reduce liver damage, inflammation, and oxidative stress. However, silencing CLSTN3 leads to the opposite effect, indicating that the CLSTN3 gene is closely related to lipid metabolism disorders in NAFLD. We believe that CLSTN3 participates in the regulation of cardiovascular disease primarily by regulating lipid metabolism.
[0043] 7) SEMA6B (semaphorin 6B, Gene ID: 10501, updated on 4-Jan-2025). In colorectal cancer, overexpression of SEMA6B is associated with poor prognosis and is linked to a suppressive tumor microenvironment. Studies have shown that tumor tissues with higher levels of SEMA6B expression are associated with gene clusters related to immune response and inflammatory activity, particularly with infiltration levels of CD4+ T cells, macrophages, regulatory T cells (Tregs), neutrophils, and dendritic cells. Furthermore, SEMA6B expression is associated with upregulation of immunosuppressive molecules and immune checkpoints, which may help tumors evade immune surveillance. Currently, research on the relationship between SEMA6B and cardiovascular disease is insufficient, and analysis suggests that it may play a role in cardiovascular disease by participating in the functional regulation of immune cells.
[0044] 8) DAB2IP (DAB2 interacting protein, Gene ID: 153090, updated on 4-Jan-2025) is a protein with multiple biological functions. In prostate cancer, DAB2IP has been shown to regulate autophagy and affect cell sensitivity to radiotherapy and chemotherapy. DAB2IP also inhibits Ras protein activity through its Ras-GTPase activating protein (GAP) domain, thereby affecting cell proliferation and survival. It also inhibits the PI3K-AKT signaling pathway by interacting with PI3K and AKT. It inhibits tumor cell proliferation and survival by regulating multiple signaling pathways. For example, in esophageal squamous cell carcinoma (ESCC), decreased DAB2IP expression is associated with increased resistance to chemotherapy and radiotherapy and is an independent predictor of disease-specific survival in ESCC patients. DAB2IP may regulate atherosclerosis and other cardiovascular diseases through autophagy and other regulatory signaling pathways.
[0045] 9) LUC7L (Putative RNA-binding protein Luc7-like 1, Gene ID: 55692, updated on 4-Jan-2025) is a protein found in humans. It belongs to the RNA-binding protein family and shares similarity with the Luc7p subunit of the yeast U1 snRNP splicing complex. LUC7L plays a key role in splicing regulation, particularly in 5' splice site selection. Its role in neurological diseases has been studied, but its role in cardiovascular disease requires further investigation. LUC7L's function in RNA splicing suggests that it may regulate cardiovascular function by affecting gene expression.
[0046] 10) NIPAL4 (NIPA like domain containing 4, Gene ID: 348938, updated on 4-Jan-2025) is closely associated with skin barrier homeostasis, and its mutations may lead to abnormal skin barrier function. For example, in one study, NIPAL4 mutations were found to disrupt the expression of sebum lipid oxygenase and transglutaminase 1, which are part of a common metabolic pathway required for maintaining skin barrier homeostasis. However, its function in cardiovascular disease is still unclear. NIPAL4's effects on endothelial barrier function may mediate the pathology of atherosclerosis.
[0047] 11) RBFOX3 (RNA binding fox-1 homolog 3, Gene ID: 146713, updated on 4-Jan-2025), also known as NeuN antigen, is an RNA-binding protein highly expressed in neural tissue and a member of the RNA-binding FOX protein family. It primarily regulates the alternative splicing of pre-mRNA and has a significant impact on neural tissue development and adult brain function. RBFOX3 plays a crucial role in neural development, participating in the maturation and differentiation of neurons, particularly in the assembly of the axonal initial segment. Although the specific role of RBFOX3 in cardiovascular disease remains unclear, its role in gene expression regulation suggests that it may exert an indirect effect by influencing gene expression in the cardiovascular system.
[0048] 12) ST18 (ST18 C2H2C-type zinc finger transcription factor, Gene ID: 9705, updated on 4-Jan-2025) can promote autoantibody-mediated epidermal cell-cell adhesion abnormalities and the secretion of proinflammatory mediators by keratinocytes. ST18 can destabilize cell-cell adhesion in a TNFalpha-dependent manner. ST18 plays an important role in gene expression regulation and cell signaling. Its specific mechanism of action in cardiovascular disease may require further research to clarify.
[0049] In the following embodiments, the method for extracting peripheral blood mononuclear cells (PBMCs) includes the following steps: 1. When young patients with acute coronary syndrome and healthy young people undergo coronary artery intervention, after implantation of an arterial sheath and before injection of heparin into the sheath, draw 10 ml of arterial blood into an EDTA purple anticoagulant tube; 2. After gently inverting the tube upside down ten times to mix, immediately place the whole blood upright at 4°C or in an ice box; 3. Return Ficoll and PBS from low temperature to room temperature; 4. Add 10 mL of Ficoll to a 50 mL centrifuge tube (marked A), and then preferably let it stand for a period of time; add 10 mL of PBS to another 50 mL centrifuge tube (marked B); 5. Dilution: Gently shake the anticoagulant blood collection tube, remove the top cover with sterile gauze and set aside, transfer 10 mL of blood to centrifuge tube B, and gently mix to obtain a PBS mixture (20 mL); 6. Tilt centrifuge tube A appropriately and slowly add 20 mL of PBS mixture along the wall (note that it should be slow, as it will penetrate the Ficoll if it is too fast). layer, which greatly affects the extraction effect); 7. Gradient centrifugation: Carefully place the centrifuge tube into the centrifuge (do not disturb the liquid layer), centrifuge at 400g for 20 minutes at room temperature, and set the acceleration / deceleration to 1 / 0 respectively; 8. After centrifugation, carefully remove it and place it on the operating table. It can be seen that it is divided into four layers ( Figure 1 9. Use a P1000 pipette to aspirate as much PBMC as possible and transfer to a 15 mL centrifuge tube, minimizing the amount of plasma and Ficoll. 10. Add PBS to bring the volume to 10 mL, cap, and gently shake five times. 11. Centrifuge at 300g for 10 minutes at room temperature, with the acceleration / deceleration settings set to 9 / 9. Remove the supernatant and gently flick the end of the centrifuge tube until the cell pellet is completely resuspended in the remaining PBS to obtain the peripheral blood cell resuspension.
[0050] In the following examples, the method for isolating and high-throughput sequencing RNA from peripheral blood mononuclear cells includes the following steps: 1. Adding TriZol to the peripheral blood cell resuspension at a volume ratio of approximately 3:1 to ensure complete lysis of the cells; RNA extraction and library preparation for sequencing: RNA from the total sample was isolated and purified using TRIzol (Invitrogen, CA, USA) according to the manufacturer's protocol. The total RNA quantity and purity were then quality controlled using a NanoDrop ND-1000 (NanoDrop, Wilmington, DE, USA). RNA integrity was then tested using a Bioanalyzer 2100 (Agilent, CA, USA) and verified by agarose gel electrophoresis. A concentration >50 ng / μL, RIN value >7.0, OD260 / 280 >1.8, and total RNA >2 μg were sufficient for downstream experiments. Ribosomal RNA (rRNA) was removed using the Epicentre Ribo-ZeroGold Kit (Illumina, San Diego, USA). The remaining RNA was fragmented using the NEBNext® Magnesium RNA Fragmentation Module (NEBNext® Magnesium RNA Fragmentation Module, Cat. No. E6150S, USA) at 94°C for 5-7 minutes. The fragmented RNA was then synthesized into cDNA using reverse transcriptase (Invitrogen SuperScript™ II Reverse Transcriptase, Cat. No. 1896649, CA, USA). Secondary strand synthesis was then performed using E. coli DNA polymerase I (NEB, Cat. No. m0209, USA) and RNase H (NEB, Cat. No. m0297, USA) to convert the DNA-RNA duplex into a DNA duplex. Simultaneously, dUTP Solution (Thermo Fisher, Cat. No. R0133, CA, USA) was incorporated into the duplex to blunt-end the duplex. An A base was then added to each end to allow ligation to a T-terminated adapter. The fragments were then size-screened and purified using magnetic beads. The second strand was digested with UDG enzyme (NEB, Cat. No. m0280, MA, US) and then subjected to PCR: initial denaturation at 95°C for 3 minutes, followed by eight cycles of denaturation at 98°C for 15 seconds each, annealing at 60°C for 15 seconds, extension at 72°C for 30 seconds, and a final extension at 72°C for 5 minutes, to generate a library with a fragment size of 300 bp ± 50 bp.Finally, we used the Illumina Novaseq™ 6000 (LC Bio Technology CO., Ltd. Hangzhou, China) to perform paired-end sequencing according to standard procedures, with the sequencing mode being PE150. 2. After extracting total RNA from the sample using a total RNA extraction kit, a strand-specific library was constructed by removing ribosomal RNA (rRNA depletion). After the library passed quality inspection, high-throughput sequencing was performed using the Illumina Novaseq™ 6000.
[0051] In the following examples, sequencing data was preprocessed using normalized raw data (raw counts) to measure gene expression levels in different samples. Feature vector selection for sequencing data analysis involved random forest analysis using the R "randomForest" package. LASSO analysis was performed using the "glmnet" package, the "xgboost" package, the "pROC" package, and ROC plotting using the "ggvenn" package. Data was converted to long format using the "reshape2" package. Scatter plots of transcription factor expression were drawn using "ggpubr." DCA plots were drawn using the "rms" package, and forest plots were drawn using the "autoReg" package.
[0052] Example 1 Determination of markers like Figure 1 As shown, Figure 1 The image shows the transcriptional markers related to acute coronary syndrome in young patients with coronary heart disease who were screened by plasma transcriptome detection in consecutive young patients enrolled in our center. The flow chart of the patent project is shown. 205 patients with chest pain were enrolled consecutively. The results of coronary angiography, troponin, electrocardiogram, etc. were used to determine whether the patients were ACS or non-ACS patients. 14,612 transcripts were screened by transcriptomics. The training set and validation set were split for feature engineering and model construction. Three machine learning algorithms (random forest, LASSO and XGBoost) were used to preliminarily screen potential markers. Then, the multivariate regression method was used to further select transcript features (such as Figure 2 As shown in the figure), finally 12 transcripts were retained for model construction, namely FAM178B, ARVCF, SLC6A9, BNC2, DKK2, CLSTN3, SEMA6B, DAB2IP, LUC7L, NIPAL4, RBFOX3, and ST18.
[0053] Example 2 Model Construction 1. Data Acquisition This example involves a cohort study of young adults with coronary heart disease. A total of 205 patients underwent transcriptomic analysis, of whom 121 (69.42%) were diagnosed with ACS and 84 (39.58%) were diagnosed with non-ACS. For this unbalanced cohort, the dataset was split using R using the "caret," "rsample," "tidyverse," and "modeldata" R packages. The random seed was set to "1234." The dataset was split into a training set and a validation set with a 70:30 ratio. The ACS and non-ACS populations in the training and validation sets were split to maintain a nearly equal ratio. The final training set consisted of 142 individuals (84 with ACS and 58 with non-ACS), and the validation set consisted of 63 individuals (37 with ACS and 26 with non-ACS). The data were well matched, and further statistical analysis and visualization were performed.
[0054] Peripheral blood was collected from the patients mentioned above, and peripheral plasma proteomic analysis was performed according to the Olink detection method to obtain the expression levels of relevant proteins. The OlinkTarget Cardiovascular Disease [CVD] II kit from Olink Proteomics (Uppsala, Sweden) was used to perform protein detection based on the Olink method. The results were processed according to the sequencing data preprocessing method. The results are shown in the figure below. Figure 3 As shown, Figure 3 The expression levels of the 12 transcripts FAM178B, ARVCF, SLC6A9, BNC2, DKK2, CLSTN3, SEMA6B, DAB2IP, LUC7L, NIPAL4, RBFOX3, and ST18 that finally entered the model in the training set (A) and validation set (B), as well as the significant differences between different groups. Cyan represents the non-ACS group, and yellow represents the ACS group. P < 0.05 indicates a significant difference.
[0055] 2 Construction of prediction model By combining random forest, LASSO regression, and XGBOOST, three machine learning methods were used to select potential candidate transcription factors for predicting the onset of ACS in young people. Secondary screening was performed using classic multivariate regression. A prediction model for predicting ACS in young people was established for the selected 12 transcription factors (FAM178B, ARVCF, SLC6A9, BNC2, DKK2, CLSTN3, SEMA6B, DAB2IP, LUC7L, NIPAL4, RBFOX3, ST18). The prediction model is as follows: Y=4.98918+(FAM178B*(-0.09285))+(ARVCF*0.00261)+(SLC6A9*(-0.00406))+(BNC2*(-0.00243))+(DKK2*(-0.06780))+(CLSTN3* 0.00041)+(SEMA6B*(0.00385))+(DAB2IP*(-0.00894))+(DAB2IP*(-0.00894))+(NIPAL4*0.02325)+(RBFOX3*(-0.03305))+(ST18*(-0.07321)) When the risk model result Y is greater than or equal to 0.669, it has a strong diagnostic performance for judging that the subject has acute coronary syndrome (specificity 94.8%, sensitivity 89.3%).
[0056] like Figure 4-8 As shown, Figure 4 Figure 3 ROC (A: training set ROC; B: validation set ROC) and prediction status (C: training set prediction probability visualization diagram, cyan marks represent ACS patients, more ACS patients have higher risk scores and are located on the upper right side of the diagram; red marks represent non-ACS patients, most of the non-ACS patients have lower risk scores and are located on the lower side of the diagram. D is the training set and validation set visualization prediction probability diagram). Figure 5 The PR curves (A, B) and calibration curves (C, D) in the training set and validation set indicate that the risk prediction model of the above transcript combination has good accuracy and the predicted situation is highly consistent with the actual situation. Figure 6 The DCA curves (A, B) and CIC (C, D) in the training set and validation set indicate that the risk prediction model of the above transcript combination has a good guiding role in clinical prediction and patient condition selection. Figure 7 This is the nomogram of the risk prediction model. Visually displaying the integral of different transcription factors for ACS risk assessment is more helpful for clinicians to manage and predict risks. Figure 8This forest plot shows the contribution of different transcripts to risk in the final model. The risk model achieved excellent predictive performance: training set AUC 0.9672, 95% CI (0.9435-0.9908), specificity 0.9482, sensitivity 0.8929, accuracy 0.9154, positive predictive value 0.96115, negative predictive value 0.8594, and Youden index 0.8411. The validation set AUC 0.8462, 95% CI (0.7958-0.9589), specificity 0.8461, sensitivity 0.7568, accuracy 0.7936, positive predictive value 0.8750, negative predictive value 0.7097, and Youden index 0.6029.
[0057] The present invention is based on a comparison of the expression levels of proteins FAM178B, ARVCF, SLC6A9, BNC2, DKK2, CLSTN3, SEMA6B, DAB2IP, LUC7L, NIPAL4, RBFOX3, and ST18 in the peripheral venous blood plasma of young patients with acute coronary syndrome and young patients with stable angina pectoris and normal coronary arteries. It was found that the content of target plasma transcription factors in samples of patients with acute coronary syndrome was different from that in patients with young stable angina pectoris and normal coronary arteries, indicating that the target transcript panel detection can be used as a marker combination for young acute coronary syndrome. The present invention provides an early diagnosis kit for acute coronary syndrome in young patients. By detecting the transcriptional expression levels of proteins FAM178B, ARVCF, SLC6A9, BNC2, DKK2, CLSTN3, SEMA6B, DAB2IP, LUC7L, NIPAL4, RBFOX3, and ST18 in the peripheral venous blood plasma of young patients, it can be quickly determined whether the patient has acute coronary syndrome. The detection is non-invasive, the diagnosis is rapid and accurate, the operation is simple, and the repeatability is good.
[0058] The present invention has been described in detail above. It will be apparent to those skilled in the art that the present invention may be practiced over a wide range of parameters, concentrations, and conditions without departing from the spirit and scope of the present invention and without unnecessary experimentation. Although specific embodiments have been given herein, it should be understood that further modifications may be made to the present invention. In summary, this application is intended to encompass any variations, uses, or improvements to the present invention, including those made by conventional techniques known in the art that depart from the scope of the present invention. Applications of the essential features may be made within the scope of the following claims.
Claims
1. A computer device comprising a memory, a processor, and a computer program stored in the memory, wherein: The processor executes the computer program to implement the following steps: S1) Data reception: Receive the expression levels of twelve proteins, including FAM178B, ARVCF, SLC6A9, BNC2, DKK2, CLSTN3, SEMA6B, DAB2IP, LUC7L, NIPAL4, RBFOX3, and ST18, in plasma samples from patients to be tested; S2) Data processing: Substituting the expression levels of the twelve proteins into the prognostic risk prediction model to calculate a risk score value for the patient to be tested as a high-risk young adult with coronary heart disease; the risk prediction model for the patient to be tested as a high-risk young adult with coronary heart disease is as follows: Risk score = Y = 4.98918 + (FAM178B * (-0.09285)) + (ARVCF * 0.00261) + (SLC6A9 * (-0.00406)) + (BNC2 * (-0.00243)) + (DKK2 * (-0.06780)) + (CLSTN3 * 0.00041) + (SEMA6B * (0.00385)) + (DAB2IP * (-0.00894)) + (DAB2IP * (-0.00894)) + (NIPAL4 * 0.02325) + (RBFOX3 * (-0.03305)) + (ST18 * (-0.07321)) Formula 1, S3) Data output: Analyze the possibility that the patient to be tested is a high-risk young coronary heart disease patient based on the risk score value.
2. The computer device according to claim 1, wherein: In step S3), the possibility that the patient to be tested with a risk score ≥0.669 is a high-risk young patient with coronary heart disease is greater than the possibility that the patient to be tested with a risk score <0.669 is a high-risk young patient with coronary heart disease.
3. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method according to claim 1 are implemented.
4. The computer-readable storage medium according to claim 31, wherein The computer-readable storage medium refers to a carrier for storing data, which may be a floppy disk, an optical disk, a DVD, a hard disk, a flash memory, a USB flash drive, a CF card, an SD card, an MMC card, an SM card, a Memory Stick or an xD card.
5. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the following steps are implemented: S1) Data reception: Receive the expression levels of twelve proteins, including FAM178B, ARVCF, SLC6A9, BNC2, DKK2, CLSTN3, SEMA6B, DAB2IP, LUC7L, NIPAL4, RBFOX3, and ST18, in plasma samples from patients to be tested; S2) Data processing: Substituting the expression levels of the twelve proteins into the prognostic risk prediction model to calculate a risk score value for the patient to be tested as a high-risk young adult with coronary heart disease; the risk prediction model for the patient to be tested as a high-risk young adult with coronary heart disease is as follows: Risk score = Y = 4.98918 + (FAM178B * (-0.09285)) + (ARVCF * 0.00261) + (SLC6A9 * (-0.00406)) + (BNC2 * (-0.00243)) + (DKK2 * (-0.06780)) + (CLSTN3 * 0.00041) + (SEMA6B * (0.00385)) + (DAB2IP * (-0.00894)) + (DAB2IP * (-0.00894)) + (NIPAL4 * 0.02325) + (RBFOX3 * (-0.03305)) + (ST18 * (-0.07321)) Formula 1, S3) Data output: Analyze the possibility that the patient to be tested is a high-risk young coronary heart disease patient based on the risk score value.
6. The computer program product according to claim 5, characterized in that In step S3), the possibility that the patient to be tested with a risk score ≥0.669 is a high-risk young patient with coronary heart disease is greater than the possibility that the patient to be tested with a risk score <0.669 is a high-risk young patient with coronary heart disease.
7. Any of the following uses of a protein and / or a substance for detecting the protein: D1) Use in the preparation of products for predicting or assisting in predicting the likelihood that a patient to be tested is a high-risk young adult with coronary heart disease; D2) Application in devices for predicting or assisting in predicting the likelihood that a patient to be tested is a high-risk young adult with coronary heart disease; The proteins are FAM178B, ARVCF, SLC6A9, BNC2, DKK2, CLSTN3, SEMA6B, DAB2IP, LUC7L, NIPAL4, RBFOX3, and ST18.
8. The use according to claim 7, characterized in that: The substance is a reagent, a kit and / or an instrument.