Prediction model and kit for identifying myocardial infarction recurrence risk

By screening out three microRNA biomarkers—MIR101-1, MIR1204, and MIR21—and combining them with a support vector machine model and qRT-PCR technology, the accuracy and dynamism of myocardial infarction recurrence risk assessment were solved, achieving efficient and minimally invasive risk warning and identification.

CN121565460APending Publication Date: 2026-02-24HUNAN QINGGENG BIOLOGICAL IND INNOVATION RES INST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511742542.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-25
Publication Date
2026-02-24

AI Technical Summary

Technical Problem

Existing methods for assessing the risk of myocardial infarction recurrence are insufficient in terms of accuracy, dynamism, and individualized early warning capabilities. Traditional clinical indicators and biochemical testing methods are difficult to identify individuals at high risk of recurrence in a timely and accurate manner, and there is a lack of dynamic monitoring tools that combine multiple biomarkers.

Method used

Using peripheral blood transcriptome data from the GEO database, three microRNAs, MIR101-1, MIR1204, and MIR21, were screened as biomarkers. A prediction model was established using a support vector machine model, and detection was performed using qRT-PCR technology to construct a highly sensitive and specific risk warning system.

Benefits of technology

It enables prospective identification of the risk of myocardial infarction recurrence, improves the sensitivity and accuracy of prediction, and is minimally invasive, low-cost, and scalable, making it suitable for large-scale population follow-up in primary healthcare institutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121565460A_ABST
    Figure CN121565460A_ABST
Patent Text Reader

Abstract

The invention discloses a prediction model and a kit for identifying the recurrence risk of myocardial infarction. Different from traditional lagging diagnosis means depending on electrocardiogram abnormity or troponin rise after AMI attack, the scheme of the invention can provide risk early warning before patients have typical recurrent symptoms by detecting abnormal expression of microRNA in blood, so as to realize prospective dynamic monitoring of people with high recurrent risk. Meanwhile, direct comparison between'first AMI non-relapse patients' and'relapse patients' is adopted, and basic cardiovascular disease background interference caused by comparison only by healthy people is avoided, so that the screened marker has higher relapse event specificity, and the prediction sensitivity and accuracy are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a predictive model and kit for identifying the risk of recurrent myocardial infarction. Background Technology

[0002] Acute myocardial infarction (AMI) is a serious cardiovascular emergency caused by the rupture or erosion of atherosclerotic plaques in the coronary arteries, leading to platelet aggregation, thrombus formation, and acute occlusion of the coronary arteries, resulting in persistent ischemic necrosis of the myocardial tissue. AMI is one of the most critical clinical manifestations of coronary heart disease and remains one of the leading causes of cardiovascular death worldwide. In recent years, with the development of comprehensive treatment methods such as emergency percutaneous coronary intervention (PCI), thrombolysis, stent implantation, and antithrombotic drugs, the in-hospital mortality rate of acute myocardial infarction (AMI) has significantly decreased, and more patients are able to survive after their first episode. However, surviving patients still face a high risk of recurrent adverse cardiovascular events during long-term follow-up, with recurrent myocardial infarction (Recurrent AMI) being the most typical and dangerous. Clinical follow-up studies have shown that AMI patients have the highest risk of recurrence within the first year after discharge, with approximately 10%–20% of patients experiencing recurrent myocardial infarction or other major cardiovascular events within one year. With extended follow-up, the long-term recurrence and related event incidence can reach approximately 30%. Compared to a first-time myocardial infarction, recurrent acute myocardial infarction (AMI) is often more complex: patients frequently have multiple coronary artery disease, residual stenosis, or new plaque rupture, leading to a high risk of thrombosis, decreased myocardial reserve, and significantly increased disability and mortality rates. Recurrence not only results in readmission and worsening cardiac function but also significantly increases the patient's medical burden and social costs, reducing their quality of life. Therefore, effective prevention and identification of individuals at high risk of recurrence are crucial aspects of secondary prevention and long-term management of cardiovascular disease. International authoritative guidelines for secondary prevention of cardiovascular disease (such as ACC / AHA and ESC) all emphasize that after discharge, AMI patients must undergo risk factor identification, regular follow-up, and individualized intervention to delay or prevent recurrent myocardial infarction as much as possible. This requirement places higher demands on the accuracy and feasibility of clinical risk stratification tools. Currently, clinical assessment of AMI recurrence risk mainly relies on the following methods: Clinical risk factors include age, underlying diseases (diabetes, hypertension, etc.), smoking history, family history, number of coronary artery lesions, and cardiac function status. Biochemical indicators: such as high-sensitivity C-reactive protein (hs-CRP), N-terminal pro-brain natriuretic peptide (NT-CRP) precursor. ProBNP, troponin (cTnI, cTnT), etc., are used to assess inflammation levels, heart failure status, and the degree of myocardial damage; Imaging examinations, such as coronary angiography, coronary CT, and echocardiography, can directly display the degree of coronary artery stenosis, vascular condition, and ventricular function. These traditional risk assessment methods have played an important role in clinical practice. Despite the continuous development of various detection methods for assessing the risk of recurrence in patients with acute myocardial infarction (AMI), existing technologies still have many shortcomings and limitations. These issues restrict the accuracy, dynamism, and personalized early warning capabilities of current detection protocols in clinical practice, specifically in the following aspects: (1) The accuracy of traditional clinical indicators and scoring methods is limited. Current risk stratification methods primarily rely on clinical risk factors such as patient age, pre-existing conditions (e.g., hypertension, diabetes), smoking history, and number of coronary artery lesions, as well as traditional scoring tools like GRACE and TIMI. While these methods can provide overall risk trends, they are largely based on static, group-based statistical results and cannot reflect dynamic pathological changes at the individual level in a timely and accurate manner. Consequently, some potentially high-risk individuals may be missed. (2) Existing biochemical and imaging detection methods have limited specificity and insufficient dynamism. Currently widely used serological indicators (such as high-sensitivity C-reactive protein hs-CRP, NT-proBNP, troponin, etc.) can indicate myocardial damage and heart failure, but they are easily affected by non-specific factors such as infection and inflammation, making it difficult to predict the long-term risk of relapse. Coronary angiography, coronary CT, and other imaging examinations can visually display vascular stenosis and in-stent conditions, but they rely on specialized equipment and operation, are costly, involve radiation exposure in some examinations, are difficult to use frequently in large populations, and mostly provide single-point static results, which are difficult to meet the needs of dynamic follow-up of recurrence risk. (3) Lack of multi-biomarker combination prediction tools specifically targeting the risk of recurrent myocardial infarction Although various cardiovascular molecular markers have been used in the diagnosis of acute myocardial infarction (AMI) or in healthy control studies in recent years, and multiple follow-up studies (such as the FAST-MI study) have confirmed that standardized cardiac rehabilitation can reduce long-term mortality to some extent, existing detection methods for the risk of recurrent myocardial infarction in the same patient after the first onset still mostly rely on single indicators or static results. There is a lack of tools that combine information from multiple markers, can be dynamically monitored, and are suitable for long-term follow-up. It is difficult to identify individuals at high risk of recurrence in a timely manner, which affects the secondary prevention and precise intervention of cardiovascular events.

[0003] Therefore, existing technologies have shortcomings and need to be improved. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a predictive model and kit for identifying the risk of recurrent myocardial infarction, so as to solve the problems mentioned in the background art.

[0005] The technical solution of the present invention is as follows: A predictive model for identifying the risk of recurrent myocardial infarction is provided, established using the following method, including: S1: The GSE48060 dataset, publicly available in the GEO database, was selected. This dataset contains peripheral blood transcriptome expression data (transient abundance table data of each RNA molecule across the entire genome, measured using peripheral blood) from first-time AMI patients and normal controls. First-time AMI patients were divided into relapse group and non-relapse group. Missing values ​​were removed, log2 transformation and normalization were performed on the original microarray signals in this dataset to ensure the consistency and comparability of expression matrices between samples.

[0006] S2: Two-group comparisons were used to perform independent samples t-tests on each microRNA probe to calculate its log2Fold Change (logFC) and original p-value; then, a threshold p-value < 0.05 was used to preliminarily screen out microRNA probes with significant suggestive value.

[0007] S3: Differential expression analysis was used to screen the data from step S2 to identify three microRNAs that were highly associated with myocardial infarction recurrence: MIR101-1, MIR1204, and MIR21. These three microRNAs showed significant expression differences between patients with recurrence and those without.

[0008] S4: Select the three microRNAs screened in step S2, namely MIR101-1, MIR1204, and MIR21, extract their expression values ​​in all patient samples, and perform standardization. Divide the standardized data into training and test sets, with the training set accounting for 60%-80% and the test set accounting for 40%-20%. Then, use a support vector machine (SVM) model, fitting the training set to the SVM model and evaluating the performance of the SVM model on the test set. The results of the SVM model show that when using these three microRNA variables, the SVM model exhibits good discriminative ability on the test set.

[0009] S5: After the model training and evaluation are completed, in order to enable model reuse and later deployment, the final support vector machine model is serialized and saved.

[0010] In step S5, the trained support vector machine model object is exported as an svm_mirna_model.pkl file using the joblib tool. In subsequent practical applications or diagnostic software development, the pre-trained model can be directly loaded using joblib.load() to predict newly input microRNA expression data, ensuring consistency and clinical usability.

[0011] The standardized data from step S4 were modeled using Logistic Regression and Random Forest models, and compared and validated with Support Vector Machine models under the same training and test set partitioning.

[0012] In step S4, the expression values ​​are standardized by scaling the features to make them have the same mean and variance, thereby improving the stability of the model during training and testing.

[0013] In step S2, to visually demonstrate the differential expression results, a volcano plot based on logFC and -log10 is generated. The differential expression results of microRNA probes can be categorized into three types: upregulation, downregulation, and not significant (Not Sig). For example, red indicates upregulation, blue indicates downregulation, and gray indicates not significant. The horizontal axis represents logFC, and the vertical axis represents -log10 (p-value). logFC represents the change in gene expression level, while -log10 (p-value) represents the confidence level of the differentially expressed gene. After identifying the differentially expressed genes, the relevant microRNA terms are found in the sequencing platform file (GPL570) used for this dataset, and the information is mapped to each differentially expressed gene. The location of the microRNA in the volcano plot is marked in green.

[0014] In step S3, to more intuitively demonstrate the relationship between the expression patterns of these microRNAs and whether a myocardial infarction recurrence occurs, a heatmap is drawn corresponding to their expression levels and patient grouping (recurrence / non-recurrence), using two different colors to indicate whether a recurrence has occurred or not. For example, green indicates no recurrence and purple indicates recurrence.

[0015] This invention also provides a kit for identifying the risk of recurrent myocardial infarction, characterized in that it comprises: detection reagents for MIR101-1, MIR1204, and MIR21. Its application method includes the following steps: A1: MicroRNA Sequence Acquisition; To achieve high sensitivity and specificity in peripheral blood microRNA detection, this study obtained the 3p / mature sequences of three target microRNAs based on human mature microRNA information registered in the miRBase database: MIR101-1, MIR1204, and MIR21. All sequences were derived from authoritative annotation libraries to ensure standardization and reproducibility across laboratories and platforms. The detection system was designed based on the TaqMan probe-based qRT-PCR principle, balancing specificity and detection sensitivity.

[0016] A2: Stem-loop RT primer and qPCR amplification primer design; cDNA was constructed using the mature stem-loop RT primer method, and qPCR amplification was performed using universal reverse primers and specific forward primers; each microRNA probe sequence carries a FAM fluorescent group and a BHQ1 quencher group to achieve highly specific fluorescence detection; the primer and probe sequences are shown in the table below:

[0017] A3: Sample Collection and Extraction; The sample required for testing is peripheral blood serum. After whole blood is drawn, serum or plasma is separated according to standard centrifugation methods. Total RNA is extracted using a commercially available RNA extraction kit and stored promptly to ensure RNA purity meets standards and is free from significant degradation. During the testing process, U6 snRNA is used as a standard internal reference RNA, and validated target microRNAs, namely MIR101-1, MIR1204, and MIR21, are selected as positive controls to ensure the effectiveness of the amplification reaction and accurate interpretation. A4: Reverse transcription reaction; the extracted total RNA needs to be reverse transcribed to generate cDNA of microRNA; each target microRNA uses a specially designed stem-loop RT primer, with the 3' end complementary to the end of the microRNA sequence and containing a fixed universal tail sequence to ensure consistent pairing in subsequent amplification.

[0018] A5: RT-qPCR amplification; After reverse transcription, the obtained cDNA is used for quantitative real-time PCR (qRT-PCR) amplification and detection; Each microRNA is detected using a specific forward primer, a universal reverse primer, and a specific TaqMan fluorescent probe (5'-FAM-sequence-BHQ1-3') that is complementary to the middle segment of the target sequence, in conjunction with the TaqMan fluorescent probe-specific reaction system.

[0019] A6: Data Transformation and Input; The Ct (Cycle Threshold) value obtained by qRT-PCR detection reflects the relative expression level of each microRNA; Since the constructed prediction model is based on the GEO microarray expression matrix, its unit is Log2 expression level or Z-score, the Ct value obtained by actual detection needs to be converted into an equivalent expression level using Equation 1: log2(expression) = (29.67 Ct) / 0.94 (Formula 1), where 29.67 and 0.94 are the intercept and slope obtained by regression fitting of the standard curve, respectively; the transformed log2(expression) value is input into the final support vector machine model obtained above for analysis.

[0020] The RT-qPCR amplification reaction system and reaction conditions are as follows:

[0021] The amplification program was fixed as follows: initial denaturation at 95℃ for 5 minutes, followed by 40-45 cycles. Each cycle included denaturation at 95℃ for 15 seconds and annealing / extension at 60℃ for 60 seconds. Fluorescence signals were recorded during the cycle, and the amplification curve of the target sequence was detected in real time. After amplification, the Ct values ​​of each target microRNA and internal control were output for subsequent data standardization and risk assessment.

[0022] The reverse transcription system contains reverse transcriptase, dNTPs, RNase inhibitors, and buffer. The primer binding, reverse transcription, and enzyme inactivation are completed according to the specified temperature program to generate a specific cDNA template. The specific system and parameters are as follows:

[0023] This invention, based on molecular-level abnormalities in the potential recurrence process after the first episode of acute myocardial infarction (AMI), screened for three core microRNAs: MIR101-1, MIR1204, and MIR21. These microRNAs showed significant expression differences between the two groups and were closely related to cardiovascular pathological processes. MIR101-1 has been shown to exert an inhibitory effect on the regulation of myocardial fibrosis and cardiovascular remodeling, affecting signaling pathways related to cardiomyocyte proliferation and apoptosis. While MIR1204 has been less reported in the cardiovascular field, it is closely related to the regulation of cellular stress and inflammatory responses. MIR21 is one of the most extensively studied cardiovascular-related microRNAs and has been shown to play an important regulatory role in myocardial injury repair, ventricular remodeling, and the chronic inflammatory microenvironment. These functional bases provide biological feasibility and theoretical support for their use as precursor molecular markers of AMI recurrence risk. Building upon this foundation, the Support Vector Machine (SVM) algorithm is further employed to jointly model and cross-validate the selected microRNA features, constructing an efficient and non-invasive predictive model that can be used to assess the risk of patient recurrence. Furthermore, by combining this model with the qRT-PCR standardization system, a risk warning can be issued before patients exhibit clinical symptoms of recurrence. Unlike traditional diagnostic methods that rely on delayed findings such as abnormal electrocardiograms or elevated troponin levels after an acute myocardial infarction (AMI), this invention detects abnormal expression of microRNAs in the blood, providing a risk warning before patients develop typical recurrence symptoms. This enables prospective and dynamic monitoring of high-risk individuals. Furthermore, this invention employs a direct comparison between patients with a first-time AMI who have not relapsed and those who have relapsed, avoiding interference from underlying cardiovascular disease backgrounds caused by using only healthy controls. This ensures that the selected biomarkers are more specific to recurrence events, significantly improving the sensitivity and accuracy of prediction.

[0024] Using the above-described approach, this invention provides a predictive model and kit for identifying the risk of recurrent myocardial infarction, which has the following technical advantages: 1. A novel combination of microRNA biomarkers: This invention proposes using MIR101-1, MIR1204, and MIR21 as combined biomarkers to identify the long-term recurrence risk in patients with first-time acute myocardial infarction (AMI). These microRNAs exhibit stable abnormal expression before the onset of myocardial infarction recurrence, possessing high sensitivity and good biological stability, and are an important functional complement to traditional single biomarkers such as troponin. 2. It can be applied to the prospective identification of recurrence risk; traditional AMI diagnosis and follow-up often rely on physiological or biochemical changes that occur after the event, such as elevated troponin or ST-segment elevation on electrocardiogram, which has a significant lag. This invention, based on the abnormal expression characteristics of microRNA, can provide risk warnings before patients develop typical clinical manifestations, which helps to achieve early monitoring and proactive intervention in high-risk groups for recurrence. 3. The detection method is minimally invasive and easy to promote; this invention uses peripheral blood microRNA detection, which does not require tissue biopsy or large imaging equipment, and has the advantages of being minimally invasive, low-cost, and easy to promote at the grassroots level. For MIR101-1, MIR1204, and MIR21, specific stem-loop RT primers, forward primers, universal reverse primers, and TaqMan fluorescent probes have been designed to form a standardized qRT-PCR detection system, ensuring stable conversion from peripheral blood samples to Ct values ​​and supporting cross-platform and cross-center applications. 4. Innovative clinical modeling strategy: Unlike the traditional design approach of simply using "healthy people vs. first-time AMI patients", this invention selects "first-time AMI patients without recurrence vs. patients with recurrence" for comparative analysis. This can effectively remove the interference of the underlying cardiovascular disease background and accurately identify specific molecular signals that are highly related to the recurrence event, making the model more in line with real clinical needs and more practical and reliable. 5. Machine Learning-Based Model Construction and Iterative Optimization: This invention employs Support Vector Machine (SVM) to model the selected microRNA features, establishing a relapse risk classifier. On the test set, the accuracy is approximately 80%, with an AUC exceeding 0.80, demonstrating excellent classification performance. This model can be seamlessly integrated with qRT-PCR detection results, directly inputting them into the model via Ct→log2(expression) transformation to achieve intelligent prediction. Attached Figure Description

[0025] Figure 1 Volcano diagram of differential expression of microRNA probes.

[0026] Figure 2 Heatmap of differentially expressed microRNAs (relapse group vs. non-relapse group).

[0027] Figure 3 ROC curves for an SVM model based on three microRNAs.

[0028] Figure 4 The relapse risk prediction performance of three microRNAs in different models. Detailed Implementation

[0029] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.

[0030] 1. Data Collection and Preprocessing The GSE48060 dataset, publicly available in the GEO database, was selected. The GSE48060 dataset, referenced in Suresh, R., et al., Transcriptome from circulating cells suggests dysregulated pathways associated with long-term recurrent events following first-time myocardial infarction. J Mol Cell Cardiol, 2014. 74: p. 13-21, contains peripheral blood transcriptome expression data from first-time AMI patients (divided into relapse and non-relapse groups) and normal controls. Missing values ​​were removed from the original microarray signals, and log2 transformation and normalization were performed. Specific processing methods can be found in Hanyu, Y., Z. Xiaoyong, and W. Lei, Review of Data Normalization Methods. Computer Engineering and Applications, 2023. 59(3): p. 13-22, ensuring consistency and comparability of expression matrices among samples. 2. Difference Screening Independent samples t-tests were performed on each microRNA probe using a two-group comparison method to calculate its log2 Fold Change (logFC) and original p-value (P.Value). The calculation method can be found in Rosati, D., et al., Differential gene expression analysis pipelines and bioinformatic tools for the identification of specific biomarkers: A review. Comput Struct Biotechnol J, 2024. 23: p. 1154-1168. Probes with suggestive significance were initially screened using a threshold P.Value < 0.05. To visually represent the differential results, a volcano plot based on logFC and -log10 (p-value) was plotted. Probes were categorized into three types: upregulated, downregulated, and not significant (Not Sig); red indicates upregulation, blue indicates downregulation, and gray indicates not significant; the horizontal axis represents logFC, and the vertical axis represents -log10 (p-value). logFC represents the change in gene expression level, while -log10 (p-value) indicates the confidence level of the differentially expressed gene. After identifying the differentially expressed genes, the microRNA-related terms were found based on the sequencing platform file (GPL570) used for this dataset, and information was then mapped to each differentially expressed gene. Please refer to [link to documentation]. Figure 1 The three microRNAs found are marked in green on the volcano map.

[0031] 3. microRNA selection and feature construction 3.1 Evaluation of candidate microRNAs In this dataset, three microRNAs highly associated with myocardial infarction recurrence were identified through differential expression analysis. Their logFC and P values ​​are shown in the table below.

[0032] Significantly different microRNAs (without vs. with recurrent events)

[0033] To more intuitively illustrate the relationship between the expression patterns of these microRNAs and the occurrence of myocardial infarction recurrence, a heatmap of their expression levels versus patient grouping (recurrence / non-recurrence) has been created. Please refer to [link to relevant documentation]. Figure 2Green indicates no recurrence, while purple indicates a recurrence.

[0034] The results showed that these three microRNAs exhibited significant expression differences between relapsed and non-relapsed patients, further supporting their feasibility as potential risk markers.

[0035] 3.2 Preliminary Modeling and Evaluation We first selected the three microRNAs (MIR101-1, MIR1204, and MIR21) with the most significant p-values ​​in the differential analysis, extracted their expression values ​​in all patient samples, and standardized them. Then, we used a support vector machine (SVM) model, fitting it to 70% of the training set and evaluating the model performance on the remaining 30% of the test set.

[0036] Please see Figure 3 Accuracy represents the proportion of correctly predicted samples, and AUC reflects the model's ability to distinguish between positive (AMI relapse) and negative (AMI non-relapse) samples. The closer the AUC is to 1, the better the performance. See Jin, H. and C.X. Ling, Using AUC and accuracy in evaluating learning algorithms. IEEE Transactions on Knowledge and Data Engineering, 2005. 17(3): p. 299-310. The results at this stage demonstrate that the combination of microRNA biomarkers has significant information content in AMI relapse discrimination. The model results show that when using these three microRNA variables, SVM exhibits good discrimination ability on the test set: • Accuracy is approximately 0.80 • The area under the curve (AUC) is 0.81. To further validate the performance of different algorithms, this technical solution also uses Logistic Regression and Random Forest methods to model the same dataset and compares them under the same training and test set partitioning. The results are shown in Figure 4. The AUC of the SVM model is significantly better than that of Logistic Regression (AUC = 0.69) and Random Forest (AUC = 0.50), indicating that the Support Vector Machine method has superior discriminative performance and generalization potential in predicting the relapse risk of AMI based on multiple microRNAs.

[0037] 3.3 Model Export and Iterative Optimization After model training and evaluation, in order to enable model reuse and later deployment, this study serialized and saved the final Support Vector Machine (SVM) model. Specifically, the joblib tool was used to export the trained SVM model object as an svm_mirna_model.pkl file.

[0038] In subsequent clinical applications or diagnostic software development, the pre-trained model can be directly loaded using joblib.load() to predict newly input microRNA expression data, ensuring consistency and clinical usability. The specific code is as follows: import joblib # Load the model svm_model = joblib.load('svm_mirna_model.pkl') # Execute prediction after inputting new data prediction = svm_model.predict(x_new_scaled).

[0039] Furthermore, this SVM model possesses excellent iterative optimization capabilities, providing a solid foundation for future continuous application and updates. As the clinical sample size accumulates, researchers can incorporate expression data from newly added AMI relapse and non-relapse patients into the existing training set. By retraining the model or introducing incremental learning strategies, the model's generalization ability and predictive accuracy can be continuously improved. This iterative optimization mechanism not only adapts to potential biological differences arising from different populations, regions, and testing platforms, but also promptly captures novel microRNA risk biomarkers discovered with the advancement of medical technology, ensuring the model maintains leading discriminative capabilities and clinical applicability.

[0040] 4. RT-qPCR Primer and Probe Design 4.1 MicroRNA Sequence Acquisition To achieve high sensitivity and specificity in the detection of peripheral blood microRNAs, this study obtained the 3p / mature sequences (MIR101-1, MIR1204, and MIR21) of three target microRNAs from the miRBase database (see Kozomara, A., M. Birgaoanu, and S. Griffiths-Jones, miRBase: from microRNA sequences to function. Nucleic Acids Res, 2019. 47(D1): p. D155-D162). All sequences were derived from authoritative annotation libraries to ensure standardization and reproducibility across laboratories and platforms. The detection system was designed based on the TaqMan probe-based qPCR principle, balancing specificity and detection sensitivity. 4.2 Design of stem-loop RT primers and qPCR amplification primers This project utilizes the established stem-loop RT primer method to construct cDNA, and then combines universal reverse primers with specific forward primers for qPCR amplification. Each microRNA probe sequence carries a FAM fluorescent group and a BHQ1 quencher group, achieving highly specific fluorescence detection. The primer and probe sequences are as follows:

[0041] 4.3 Experimental Design 4.3.1 Sample Collection and Extraction The required sample for the test is peripheral blood serum. After whole blood is drawn, serum or plasma is separated according to standard centrifugation method, and total RNA is extracted using a commercially available RNA extraction kit. It is stored in a timely manner to ensure that the RNA purity meets the standard and there is no obvious degradation. During the test, internal control RNA needs to be prepared (U6 snRNA is used as the standard internal control (refer to Chen, C., et al., Real-time quantification of microRNAs by stem-loop RT-PCR. Nucleic Acids Res,2005. 33(20): p. e179.)), and the validated target microRNA (three candidate microRNAs screened in this protocol) is selected as a positive control to ensure the effectiveness of the amplification reaction and the accuracy of interpretation.

[0042] 4.3.2 Reverse transcription reaction The extracted total RNA needs to be reverse transcribed to generate cDNA from microRNAs. Each target microRNA uses a specially designed stem-loop RT primer, with its 3' end complementary to the end of the microRNA sequence and containing a fixed universal tail sequence to ensure consistent pairing during subsequent amplification. The reverse transcription system contains reverse transcriptase, dNTPs, RNase inhibitors, and buffer. The primer binding, reverse transcription, and enzyme inactivation are completed according to the specified temperature program to generate a specific cDNA template. The specific system and parameters are as follows:

[0043] 4.3.3 RT-qPCR Amplification After reverse transcription, the obtained cDNA was used for quantitative real-time PCR (qRT-PCR) amplification and detection. Each microRNA was detected using a matching specific forward primer, a universal reverse primer, and a specific TaqMan fluorescent probe (5'-FAM-sequence-BHQ1-3') complementary to the target sequence mid-segment, in accordance with the TaqMan fluorescent probe-specific reaction system. The reaction system and conditions are as follows:

[0044] The amplification program was fixed as follows: initial denaturation at 95℃ for 5 minutes, followed by 40–45 cycles, each cycle consisting of denaturation at 95℃ for 15 seconds and annealing / extension at 60℃ for 60 seconds, during which fluorescence signals were recorded, and the amplification curve of the target sequence was detected in real time. After amplification, the Ct values ​​of each target microRNA and internal control were output for subsequent data standardization and risk assessment.

[0045] 5. Data Transformation and Data Input The Ct (Cycle Threshold) values ​​obtained by qRT-PCR detection reflect the relative expression level of each microRNA. Since the constructed prediction model is based on the GEO microarray expression matrix (unit: Log2 expression level or Z-score), the actual Ct values ​​obtained need to be converted to equivalent expression levels using Equation 1: log2(expression) = (29.67) Ct) / 0.94 (Formula 1); Wherein, 29.67 and 0.94 are the intercept and slope obtained by regression fitting of the standard curve, respectively. See Li, Y., et al., Systematic identification and validation of the reference genes from 60 RNA-Seq libraries in the scallop Mizuhopecten yessoensis. BMCGenomics, 2019. 20(1): p. 288. The transformed log2 (expression) value will be input into the prediction model (i.e., support vector machine model) used to identify the risk of recurrent myocardial infarction for analysis.

[0046] This invention has the following advantages: 1. Clinical grouped control design for AMI recurrence risk: This invention adopts the strategy of "first AMI patients without recurrence versus patients with recurrence" to accurately identify molecular expression changes directly related to recurrence events. This is different from the traditional research design of "healthy controls" or "CVD patients controls". It avoids interference signals from underlying diseases, is more in line with real clinical follow-up scenarios, and significantly improves the clinical specificity and practical value of biomarker screening. 2. High-performance relapse risk prediction model based on a small number of microRNAs: This invention is based on differential expression analysis combined with support vector machine (SVM) algorithm. It can efficiently identify the relapse risk of AMI using only 3 core microRNAs (MIR101-1, MIR1204, MIR21). The model achieves AUC≈0.81 and accuracy of about 0.80 on the test set. It has the advantages of high modeling efficiency, low input dimensionality and low computational cost, and is suitable for large-sample dynamic assessment system in clinical scenarios.

[0047] 3. Dedicated Primer-Probe Combinations and Standardized Detection Systems for Three MicroRNAs: This invention designs dedicated stem-loop RT primers, specific forward primers, universal reverse primers, and TaqMan fluorescent probes (labeled with FAM / BHQ1) for three microRNAs: MIR101-1, MIR1204, and MIR21. This constructs a standardized qRT-PCR detection system with high specificity, high sensitivity, and platform compatibility. This combination can be packaged as a core detection module into molecular diagnostic kits, providing technical support and a translational basis for low-invasive detection and intelligent prediction of AMI relapse risk.

[0048] Terminology Explanation 1. microRNA (miRNA): A class of non-coding RNA molecules about 22 nucleotides long, which are widely involved in gene expression regulation and can serve as potential biomarkers in a variety of diseases. 2. AMI: Acute Myocardial Infarction, is an acute event caused by coronary artery obstruction leading to myocardial ischemia and necrosis. 3. CVD: Cardiovascular Disease, including coronary heart disease, hypertension, heart failure and other clinical diseases. 4. AUC: Area Under the Curve, an important indicator for measuring the overall classification performance of a model. The closer it is to 1, the stronger its predictive ability. 5. Support Vector Machine (SVM): A commonly used supervised learning classification algorithm, suitable for scenarios with small samples and high-dimensional features, and has good nonlinear discrimination and generalization capabilities. 6. Standardization: One of the data preprocessing steps, which improves model stability by scaling features to have the same mean and variance. 7. Differential expression analysis: Identify molecules, such as genes and microRNAs, that show significant differences in expression levels between the disease group and the control group using statistical methods. 8. RT-qPCR (Real-Time Quantitative Polymerase Chain Reaction): This is a molecular detection technique that uses fluorescence signals to monitor the product amplification process to achieve quantitative RNA analysis. It is typically used to accurately quantify microRNA or mRNA expression levels. The experimental procedure includes two steps: RNA extraction and reverse transcription (RT) followed by real-time PCR amplification and detection. See Chen, C., et al., Real-time quantification of microRNAs by stem-loop RT-PCR. Nucleic Acids Res, 2005. 33(20): p. e179. 9. Reference gene: A relatively stable RNA molecule expressed in each sample, used as a calibration standard for qPCR to eliminate systematic errors.

[0049] In summary, this invention provides a predictive model and kit for identifying the risk of recurrent myocardial infarction, which has the following technical advantages: 1. Based on differential expression analysis, this invention screened out three key microRNAs: MIR101-1, MIR1204, and MIR21. A relapse risk classification model was established using support vector machine (SVM), achieving an accuracy of approximately 0.80 and an AUC of 0.81 on the test set. This significantly improved the sensitivity and specificity of the prediction, outperforming traditional single biomarker or static scoring methods. 2. The detection system is based on routine peripheral venous blood collection and combined with a mature qRT-PCR technology platform. It has advantages such as no need for tissue biopsy, simple operation process and low detection cost, and is suitable for primary medical institutions and large-scale population follow-up scenarios. 3. Specific stem-loop RT primers, forward primers, universal reverse primers, and TaqMan fluorescent probes have been designed for three target microRNAs, constructing a complete qRT-PCR detection system with high sensitivity, specificity, and batch-to-batch consistency, suitable for standardized mass production and kit development. 4. The test results can be directly integrated into the saved machine learning model through the Ct → log2(expression) data transformation process, constructing an integrated prediction closed loop from sample collection, molecular detection, data transformation to intelligent interpretation, which significantly improves diagnostic efficiency and clinical applicability. 5. This invention has achieved persistent model storage, has plug-and-play capability, supports iterative training and generalization optimization with new samples, is suitable for long-term application scenarios with multiple centers, multiple populations, and multiple time points, and can be extended to other cardiovascular disease types or multi-omics joint analysis applications, with broad market prospects.

[0050] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A predictive model for identifying the risk of recurrent myocardial infarction, characterized in that, The following methods are used to establish it, including: S1: The GSE48060 dataset, publicly available in the GEO database, was selected. This dataset contains peripheral blood transcriptome expression data of first-time AMI patients and normal controls. First-time AMI patients were divided into relapse group and non-relapse group. Missing values ​​were removed, log2 transformation and normalization were performed on the original chip signals in this dataset to ensure the consistency and comparability of expression matrices between samples. S2: Two sets of comparisons were used to perform independent samples t-tests on each microRNA probe to calculate its log2 FoldChange and original p-value. The abbreviation for log2 FoldChange is logFC, and the abbreviation for original p-value is p-value. Then, the threshold p-value < 0.05 was used to initially screen out microRNA probes with significant suggestive value. S3: Differential expression analysis was used to screen out three microRNAs that were highly associated with myocardial infarction recurrence from the data in step S2, namely MIR101-1, MIR1204, and MIR21. These three microRNAs showed significant expression differences between patients with recurrence and those without recurrence. S4: Select the three microRNAs screened in step S2, namely MIR101-1, MIR1204, and MIR21, extract their expression values ​​in all patient samples, and perform standardization. Divide the standardized data into training set and test set, with the training set accounting for 60%-80% and the test set accounting for 40%-20%. Then, use a support vector machine model, fit the training set to the support vector machine model, and evaluate the performance of the support vector machine model using the test set. S5: After the model training and evaluation are completed, in order to enable model reuse and later deployment, the final support vector machine model is serialized and saved.

2. The predictive model for identifying the risk of recurrent myocardial infarction according to claim 1, characterized in that, In step S5, the trained support vector machine model object is exported as an svm_mirna_model.pkl file using the joblib tool.

3. The predictive model for identifying the risk of recurrent myocardial infarction according to claim 2, characterized in that, In subsequent practical applications or diagnostic software development, the pre-trained model can be directly loaded using joblib.load() to predict newly input microRNA expression data, ensuring consistency and clinical usability.

4. The predictive model for identifying the risk of recurrent myocardial infarction according to claim 1, characterized in that, Logistic regression and random forest models are used to model the standardized data in step S4, and the models are compared and verified with support vector machine models under the same training and test set partitioning. In step S4, the expression values ​​are standardized by scaling the features to make them have the same mean and variance.

5. A predictive model for identifying the risk of recurrent myocardial infarction according to claim 1, characterized in that, In step S2, to visually demonstrate the differential results, a volcano plot based on logFC and -log10 is drawn. The differential results of microRNA probes are divided into three categories: upregulation, downregulation, and insignificance. The horizontal axis is logFC, and the vertical axis is -log10 (p-value). logFC represents the change in gene expression level, while -log10 (p-value) represents the confidence level of the differentially expressed gene. After finding the differentially expressed gene, the term related to microRNA is found according to the sequencing platform file used for this set of data. The sequencing platform file is GPL570 and corresponds to the information of each differentially expressed gene. The position of microRNA in the volcano plot is marked.

6. The predictive model for identifying the risk of recurrent myocardial infarction according to claim 1, characterized in that, In step S3, to more intuitively demonstrate the relationship between the expression patterns of these microRNAs and whether myocardial infarction recurrence occurs, a heatmap of their expression levels and corresponding patient groups is drawn. Patients are divided into two groups: relapse and non-relapse, and two different colors are used to mark the non-relapse and relapse groups.

7. A kit for identifying the risk of recurrent myocardial infarction, characterized in that, include: Detection reagents for MIR101-1, MIR1204, and MIR21.

8. A kit for identifying the risk of recurrent myocardial infarction according to claim 7, characterized in that, Its application method includes the following steps: A1: Establish a detection system for MIR101-1, MIR1204, and MIR21; the detection system is designed based on the TaqMan probe-based qRT-PCR principle, taking into account both specificity and detection sensitivity; A2: Stem-loop RT primer and qPCR amplification primer design; cDNA was constructed using the mature stem-loop RT primer method, and qPCR amplification was performed using universal reverse primers and specific forward primers; each microRNA probe sequence carries a FAM fluorescent group and a BHQ1 quencher group to achieve highly specific fluorescence detection; the primer and probe sequences are shown in the table below: A3: Sample Collection and Extraction; The sample required for testing is peripheral blood serum. After whole blood is drawn, serum or plasma is separated according to standard centrifugation methods. Total RNA is extracted using a commercially available RNA extraction kit and stored promptly to ensure RNA purity meets standards and is free from significant degradation. During the testing process, U6 snRNA is used as a standard internal reference RNA, and validated target microRNAs, namely MIR101-1, MIR1204, and MIR21, are selected as positive controls to ensure the effectiveness of the amplification reaction and accurate interpretation. A4: Reverse transcription reaction; the extracted total RNA needs to be reverse transcribed to generate cDNA of microRNA; each target microRNA uses a specially designed stem-loop RT primer, with the 3' end complementary to the end of the microRNA sequence and containing a fixed universal tail sequence to ensure consistent pairing in subsequent amplification; A5: RT-qPCR amplification; After reverse transcription, the obtained cDNA is used for quantitative real-time PCR amplification and detection; Each microRNA is detected using a specific forward primer, a universal reverse primer, and a specific TaqMan fluorescent probe complementary to the target sequence mid-segment, in conjunction with a TaqMan fluorescent probe-specific reaction system for amplification; A6: Data Transformation and Input; The Ct values ​​obtained from qRT-PCR detection reflect the relative expression level of each microRNA; Since the constructed prediction model is based on the GEO microarray expression matrix, its unit is Log2 expression level or Z-score, the Ct values ​​obtained from the actual detection need to be converted into equivalent expression levels using Equation 1: log2(expression)=(29.67 Ct) / 0.94 (formula 1), Wherein, 29.67 and 0.94 are the intercept and slope obtained by regression fitting of the standard curve, respectively; the transformed log2(expression) value is input into the final support vector machine model obtained by any one of claims 1-6 for analysis.

9. A kit for identifying the risk of recurrent myocardial infarction according to claim 8, characterized in that, The RT-qPCR amplification reaction system and reaction conditions are as follows: The amplification program was fixed as follows: initial denaturation at 95℃ for 5 minutes, followed by 40-45 cycles. Each cycle included denaturation at 95℃ for 15 seconds and annealing / extension at 60℃ for 60 seconds. Fluorescence signals were recorded during the cycle, and the amplification curve of the target sequence was detected in real time. After amplification, the Ct values ​​of each target microRNA and internal control were output for subsequent data standardization and risk assessment.

10. A kit for identifying the risk of recurrent myocardial infarction according to claim 7, characterized in that, The reverse transcription system contains reverse transcriptase, dNTPs, RNase inhibitors, and buffer. The primer binding, reverse transcription, and enzyme inactivation are completed according to the specified temperature program to generate a specific cDNA template. The specific system and parameters are as follows: 。