Use of an miRNA detection reagent in the preparation of a diagnostic reagent for predicting the risk of ischemic stroke
By using a prediction model constructed with 14 miRNAs and the XGBoost algorithm, hsa-miR-142-5p, hsa-miR-107, and hsa-miR-134-5p were selected as key miRNAs, solving the problem of predicting the risk of ischemic stroke and achieving efficient and economical disease screening.
Patent Information
- Application Number
- CN202511001597.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-07-21
AI Technical Summary
There is a lack of effective methods in the current technology to predict the risk of ischemic stroke, and the application of miRNA in this field has not been fully explored.
Fourteen specific miRNAs were used as biomarkers, and a prediction model was constructed using the XGBoost algorithm. hsa-miR-142-5p, hsa-miR-107, and hsa-miR-134-5p were selected as key miRNAs to construct a simplified prediction model to identify high-risk individuals for ischemic stroke.
It has achieved the ability to efficiently identify high-risk groups for ischemic stroke. The simplified model has a predictive effect comparable to the full model, reducing detection costs and meeting the cost-effectiveness of disease screening.
Smart Images

Figure CN120505414B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of biological medicine, and particularly relates to application of an miRNA detection reagent in preparation of a diagnostic reagent for predicting the risk of ischemic stroke. BACKGROUND
[0002] Ischemic stroke accounts for about 70% of all stroke cases, is an acute cerebrovascular disease characterized by brain tissue damage and neurological dysfunction, and is mainly caused by cerebral infarction caused by atherothrombosis or other large vessel emboli. Due to the serious consequences of ischemic stroke, its prevention is highly valued.
[0003] Small RNA (miRNA) is a single-stranded non-coding RNA, 18-25 nucleotides long, which can negatively regulate the translation of protein-coding genes by preventing the assembly of mRNA-ribosome complexes or accelerating the degradation of mRNA at the initial stage of translation.
[0004] It is still unknown whether miRNA has a predictive effect on the onset of ischemic stroke. Developing a diagnostic reagent capable of predicting the risk of ischemic stroke is a technical problem to be solved. SUMMARY
[0005] This section aims to summarize some aspects of the embodiments of the application and briefly introduce some preferred embodiments.
[0006] As one aspect of the application, the application provides application of an miRNA detection reagent in preparation of a diagnostic reagent for predicting the risk of ischemic stroke, wherein the miRNA comprises 14 miRNAs, and the 14 miRNAs are hsa-miR-107, hsa-miR-134-3p, hsa-miR-134-5p, hsa-miR-140-3p, hsa-miR-142-5p, hsa-miR-320a-3p, hsa-miR-320b, hsa-miR-320d, hsa-miR-361-5p, hsa-miR-423-5p, hsa-miR-483-5p, hsa-miR-484, hsa-miR-486-5p and hsa-miR-718.
[0007] The primer sequences of the 14 miRNAs are: hsa-miR-107 gcagcattgtacagggctatca, hsa-miR-134-3p gggccacctagtcaccaaaa, hsa-miR-134-5p atagcgctgtgactggttga, hsa-miR-140-3p ccacagggtagaaccacgg, hsa-miR-142-5p cgcccataaagtagaaagcactact, hsa-miR-320a-3p aaaagctgggttgagagggcga, hsa-miR-320b aaaagctgggttgagagggca, hsa-miR-320d ggaaaagctgggttgagagga, hsa-miR-361-5p gttatcagaatctccaggggtac, hsa-miR-423-5p gggcagagagcgagacttt, hsa-miR-483-5p aagacgggaggaaagaaggga, hsa-miR-484 tcaggctcagtcccctcccg, hsa-miR-486-5p ctgtactgagctgccccg and hsa-miR-718 ttacttccgccccgccggg.
[0008] As a preferred scheme of the application: using the XGBoost algorithm to construct the miRNA prediction model of ischemic stroke.
[0009] As a preferred scheme of the application: including randomly dividing the research subjects into a training set and a validation set according to a ratio of 7:3, first constructing the XGBoost prediction model of ischemic stroke in the training set, and testing the prediction ability of the model in the validation set.
[0010] As a preferred scheme of the application: the model parameters are: the target function is selected as binary:logistic, the random seed is set as 2025, the learning rate is set as 0.1, the maximum depth of the tree is set as 3, the minimum leaf weight is set as 1, the split threshold is set as 0.5, the row sampling rate is set as 0.5, and the column sampling rate is set as 1.
[0011] As a preferred scheme of the application: the miRNA consists of 3 miRNAs, and the 3 miRNAs are: hsa-miR-142-5p, hsa-miR-107 and hsa-miR-134-5p.
[0012] As a preferred scheme of the application: the detection sample of the miRNA detection reagent comprises a plasma sample.
[0013] The application has the beneficial effects that: the application screens 14 miRNAs which are all expressed in an elevated manner in the GEO data set and the previously collected patient plasma samples, and the 14 miRNAs are used as biomarkers for predicting the onset of ischemic stroke. A prediction model containing all the 14 miRNAs and a simplified prediction model containing only three relatively important miRNAs, i.e., hsa-miR-142-5p, hsa-miR-107 and hsa-miR-134-5p, are respectively constructed to identify high-risk groups of ischemic stroke, and finally it is found that the prediction model constructed by the three miRNAs can achieve the same ability of predicting the onset of ischemic stroke as the prediction model constructed by the 14 miRNAs, and is more in line with the cost-effectiveness of disease screening. BRIEF DESCRIPTION OF DRAWINGS
[0014] In order to more clearly illustrate the technical solutions of the embodiments of the application, the drawings needed in the embodiment description will be briefly introduced below, in which:
[0015] Figure 1 FIG. 1 is a diagram of the expression levels of miRNAs in the GEO data set.
[0016] Figure 2 FIG. 2 is the expression levels of 14 miRNAs in the case group and the control group.
[0017] Figure 3 FIG. 3 is the importance ranking of miRNAs for ischemic stroke.
[0018] Figure 4 FIG. 4 is the ROC curves and the areas under the curves (AUCs) of the 14-miRNA prediction model in the training set and the validation set (A) and the ROC curves and the AUCs of the 3-miRNA prediction model in the training set and the validation set (B). Figure 4 Figure 4
[0019] Figure 5 FIG. 5 is the ROC curve of a single miRNA for predicting ischemic stroke. DETAILED DESCRIPTION
[0020] In order to make the above-mentioned purposes, features and advantages of the application more apparent and easy to understand, the specific embodiments of the application will be described in detail below.
[0021] The GEO database is a high-throughput microarray, next-generation sequencing and other high-throughput sequencing data. In order to screen the differentially expressed miRNAs in ischemic stroke patients, the present application retrieves miRNA sequencing data sets containing plasma samples of ischemic stroke patients and healthy controls from the GEO database: GSE43618, GSE86291 and GSE95204. Among them, GSE43618 contains 8 control samples and 16 patient samples, GSE86291 contains 4 control samples and 7 patient samples, and GSE95204 contains 3 control samples and 3 patient samples. The above three GEO data sets are batch-processed using the R software package sva, and integrated into a data set containing 26 patients and 15 controls. We use the R software package limma to calculate the fold change and false discovery rate (FDR) adjusted P value between patients and controls. The miRNAs with fold change (FC) ≥ 1.5 and adjusted P value < 0.05 are considered as differentially expressed miRNAs.
[0022] The present application has carried out a prospective cohort study since July 2015, and the cohort has included 47516 research subjects, and a whole blood sample of each research subject is extracted and collected into 3 EDTA anticoagulation tubes, about 10 mL of peripheral blood is collected in each tube, one of the EDTA anticoagulation tubes is centrifuged at 3,000xg for 10 minutes at 4°C to separate the plasma and transfer it to another new anticoagulation tube. All samples are stored in a -80°C ultra-low temperature refrigerator.
[0023] Since the mature miRNA is only about 20 nucleotides long, it is not long enough to base-pair with the forward and reverse primers, therefore, in the process of miRNA reverse transcription, the present research uses PolyA polymerase to add a polyadenylate sequence to the 3' end of the miRNA sequence, increasing its length from the original 20 nucleotides to more than 80 nucleotides, then the reverse transcription primer binds to the PolyA sequence, and the cDNA synthesis is completed by reverse transcriptase, finally the real-time quantitative PCR technique is used to detect the miRNA expression level in the plasma sample, the specific steps are as follows:
[0024] 1) Take 2 μL of extracted RNA, and use the miRNA cDNA first strand synthesis kit to prepare the miRNA reverse transcription mixture in a PCR tube.
[0025] 2) Place the PCR tube in the PCR instrument for reverse transcription.
[0026] 3) Obtain the sequence of target miRNA from miRbase database, and design forward specific primers according to the sequence. Since the same polyadenylate sequence is added to the 3' end of each miRNA sequence by using PolyA polymerase, the reverse primer of each miRNA is Oligo dT sequence, i.e. containing 18 deoxythymine (dT), which is complementary to the polyadenylate sequence. The miRNA sequence and primer sequence are shown in Table 1.
[0027] 4) Dilute the cDNA obtained from the reverse transcription reaction by 5 times, and use TB Green Premix Ex Taq II kit and specific primers of target miRNA to prepare the real-time quantitative PCR reaction mixture.
[0028] 5) Add the reaction system to the 384-well plate, and set three replicate wells for each miRNA of each sample to ensure the stability of the results. Place the 384-well plate into the Roche LightCycler 480 real-time quantitative PCR instrument for expansion and detection.
[0029] Use Roche LightCycler software to calculate the Cq value of each well in the 384-well plate, i.e. the number of cycles required for the fluorescence intensity in each reaction system to reach the threshold value. The Cq value is inversely proportional to the expression level of the target miRNA in the sample, i.e. the higher the Cq value, the lower the expression level of the miRNA. In the present application, the Cq value of the miRNA > 40 is considered to be too low to be detected in the sample. The relative expression level of the miRNA is calculated by 2 -ΔCq, where ΔCq = Cq miRNA- Cq cel-miR-39 , Cq miRNA is the Cq value of the target miRNA, and Cq cel-miR-39 is the Cq value of the external reference cel-miR-39 in the same sample.
[0030] The present application also determines the specificity of the miRNA primer by the distribution of the melting curve. If the melting curve in the three wells has the same peak, the primer has good specificity, otherwise, it indicates that the primer has poor specificity. In addition to calculating the Cq value of the target miRNA, the present application also calculates the Cq value of UniSp2, UniSp4 and UniSp5 to control the quality of miRNA extraction, reverse transcription and quantification. The principle is that when UniSp2, UniSp4 and UniSp5 are added, their concentrations are gradually reduced (UniSp2: 2 mol / L, UniSp4: 0.02 mol / L, UniSp5: 0.00002 mol / L), so under normal circumstances, their Cq values will show a gradually increasing trend. However, when the sample quality is poor, the RNA extraction efficiency or transcription efficiency is inconsistent, the Cq values of UniSp2, UniSp4 and UniSp5 will be abnormal.
[0031] Table 1 miRNA and primer sequence thereof
[0032]
[0033] Statistical analysis: In the present application, the categorical variables are expressed by frequency and percentage (n, %), the continuous variables with normal distribution are expressed by mean and standard deviation (mean ± SD), and the continuous variables with skewed distribution are expressed by median, P25 and P75. Chi-square test is used to compare the differences of categorical variables between different groups. Student's t-test and Wilcoxon rank sum test are used to compare the differences of normally distributed and skewed distributed variables between groups, respectively.
[0034] To screen the key miRNAs in the process of ischemic stroke, we used Random Forest and Boruta algorithm to estimate the importance of each miRNA relative to ischemic stroke. Random Forest mainly uses the degree of decline in Gini index to measure the importance of variables. In short, Gini index is an indicator to measure the discriminant ability of the model in the range of 0 (no discrimination) to 100 (best discrimination). When each node on the decision tree reaches the highest purity, that is, all individuals falling into the node belong to the same category, the Gini coefficient of the model is the smallest, the purity is the highest, and the discriminant ability of the model is the highest, which can better classify individuals. Boruta algorithm screening is an algorithm for feature selection, which can identify the most important features in a given dataset and is widely used in machine learning. This algorithm is based on the Random Forest method, which estimates the importance of each feature by randomly shuffling it and mixing it with the original data features. By comparing the performance of the original features and the randomly mixed features, the algorithm determines whether the variable is significant in predicting the outcome. In this study, Random Forest and Boruta algorithm were implemented through R software "randomForest" package and "Boruta" package, respectively.
[0035] In this study, we also used XGBoost (eXtreme Gradient Boosting) algorithm to build a miRNA prediction model for ischemic stroke. XGBoost is a machine learning algorithm that belongs to the optimization implementation of the gradient boosting framework. Its core idea is to iteratively train multiple weak learners (usually decision trees) and continuously optimize the results based on the prediction errors of the previous models. Finally, the prediction results of multiple weak models are weighted and integrated into a strong model. XGBoost algorithm corrects the prediction error layer by layer by integrating multiple decision trees, and combines a unique regularization mechanism (L1 / L2 regularization term) to capture key risk factors (such as genetic markers, biochemical indicators, and lifestyle habits) in complex medical data while effectively avoiding overfitting problems. XGBoost has built-in missing value processing and parallel computing capabilities, which can quickly process high-dimensional, multi-modal medical data.
[0036] In this study, all study subjects were randomly divided into training set and validation set according to the ratio of 7:3. First, we built an XGBoost prediction model for ischemic stroke in the training set, and then tested the prediction ability of the model in the validation set. In this process, we optimized the model parameters by adjusting them. The specific parameter configurations are as follows:
[0037] Objective function (objective) according to the task type selection binary:logistic;
[0038] The random seed is set to 2025 to ensure the repeatability of the results;
[0039] The learning rate is set to 0.1;
[0040] The maximum depth of the tree is set to 3;
[0041] The minimum child weight is set to 1;
[0042] The split threshold (gamma) is set to 0.5;
[0043] The row sampling rate (subsample) is set to 0.5;
[0044] The column sampling rate (colsample_bytree) is set to 1;
[0045] The Receiver Operating Characteristic Curve (ROC) is drawn to visualize the discriminative ability of the prediction model, and the Area Under Curve (AUC) is calculated to quantify the overall performance of the model. At the same time, the Accuracy, Sensitivity, Specificity, Positive predictive value (PPV), and Negative predictive value (NPV) are used to evaluate the prediction effect of the model, and the calculation formulas of the above indicators are as follows:
[0046]
[0047] Among them, TP (True Positive) is the number of correctly identified positive samples; FN (False Negative) is the number of positive samples misjudged as negative; FP (false positive) is the number of negative samples misjudged as positive; TN (true negative) is the number of correctly identified negative samples.
[0048] The miRNA expression difference results in the GEO dataset are shown in Figure 1 Among the 327 miRNAs included in the GEO dataset, the expression of 23 miRNAs was differentially expressed with a fold change ≥1.5 and an adjusted P value <0.05. The fold change and adjusted P value of the differentially expressed miRNAs are shown in Table 2.
[0049] Table 2 Fold change and P value of differentially expressed miRNAs in GEO dataset and RT-qPCR detection
[0050]
[0051] *Seven miRNAs were undetectable in the vast majority of samples (Cq>40), and the primers for two miRNAs had low specificity. RT-PCR, real-time quantitative polymerase chain reaction.
[0052] RT-qPCR results: After excluding 7 miRNAs with extremely low expression levels (Cq value > 40) in most samples and 2 miRNAs with low primer specificity, we used RT-qPCR to detect the expression levels of 14 candidate miRNAs in plasma samples from ischemic stroke patients and controls: hsa-miR-107, hsa-miR-134-3p, hsa-miR-134-5p, hsa-miR-140-3p, hsa-miR-142-5p, hsa-miR-320a-3p, hsa-miR-320b, hsa-miR-320d, hsa-miR-361-5p, hsa-miR-423-5p, hsa-miR-483-5p, hsa-miR-484, hsa-miR-486-5p, and hsa-miR-718. Figure 2 All miRNAs were expressed higher in the patient group than in the control group, exhibiting consistent directionality compared to the GEO dataset, and the differences between the two groups were statistically significant as determined by the Wilcoxon signed-rank test. The fold change and p-values of miRNAs detected by RT-qPCR are shown in Table 2.
[0053] miRNA screening: The importance of each miRNA for ischemic stroke, such as... Figure 3 As shown, the importance of hsa-miR-142-5p, hsa-miR-107, and hsa-miR-134-5p, evaluated using the random forest algorithm and the Boruta algorithm, is significantly higher than that of other miRNAs.
[0054] Prediction of ischemic stroke risk by miRNAs: The cutoff values for the prediction of ischemic stroke by a single miRNA are shown in Table 3 (the expression level of the external reference cel-miR-39 was set to 1), and the area under the ROC curve for each miRNA was between 0.589 and 0.751. Figure 5 The prediction accuracy was low. The area under the curve (AUC) of the XGBoost prediction model based on 14 miRNAs, however, reached over 0.9. Figure 4AUC of the prediction model based on hsa-miR-142-5p, hsa-miR-107 and hsa-miR-134-5p also reached above 0.9 (Fig. 1B). In addition, the accuracy, sensitivity, specificity, positive predictive value and negative predictive value of the above two models reached a high level (Table 4). Figure 4
[0055] Table 3 Critical value of single miRNA predicting ischemic stroke
[0056]
[0057] Table 4 Performance of prediction model of risk of ischemic stroke
[0058]
[0059] The present application selected 604 research subjects diagnosed as ischemic stroke and 604 normal controls in the process of cohort follow-up. At the time of entering the cohort, the medical staff extracted the peripheral blood of each research subject and separated the plasma, that is, the plasma sample of each research subject was collected before the onset, which provided the basis for us to screen the disease markers in the plasma. Finally, we found that compared with the control group, 14 miRNAs were expressed in the plasma of ischemic stroke patients before the onset, which was consistent with the direction in the GEO data set. In addition, the present application used random forest algorithm and Boruta algorithm to screen hsa-miR-142-5p, hsa-miR-107 and hsa-miR-134-5p, which were more important for ischemic stroke. We constructed a prediction model containing 14 miRNAs, and the prediction ability of the model was higher than that of the prediction model constructed by a single miRNA. In addition, the prediction model based on hsa-miR-142-5p, hsa-miR-107 and hsa-miR-134-5p can achieve the same prediction efficiency as the 14-miRNA model. Therefore, fewer miRNAs can be used to predict the risk of ischemic stroke, so as to achieve approximate prediction efficiency at lower detection cost.
[0060] In summary, the present application screened 14 miRNAs that were all expressed at higher levels in GEO data sets and previously collected patient plasma samples, which served as biomarkers for predicting the onset of ischemic stroke. We constructed a prediction model containing all 14 miRNAs and a simplified prediction model containing only three relatively important miRNAs, hsa-miR-142-5p, hsa-miR-107, and hsa-miR-134-5p, to identify high-risk groups of ischemic stroke, and ultimately found that the prediction model constructed by the simplified three miRNAs could achieve the same ability to predict the onset of ischemic stroke as the prediction model constructed by the 14 miRNAs, and was more cost-effective for disease screening.
[0061] It should be noted that the above examples are only used to illustrate the technical solutions of the present application and are not limiting. Although the present application has been described in detail with reference to the preferred embodiments, it should be understood by those skilled in the art that the technical solutions of the present application can be modified or replaced equivalently without departing from the spirit and scope of the technical solutions of the present application, and they should be covered in the scope of the claims of the present application.
Claims
1. Use of an miRNA detection reagent in the manufacture of a diagnostic reagent for predicting the risk of onset of ischemic stroke, characterized in that: The miRNAs include 14 miRNAs, respectively: hsa-miR-107, hsa-miR-134-3p, hsa-miR-134-5p, hsa-miR-140-3p, hsa-miR-142-5p, hsa-miR-320a-3p, hsa-miR-320b, hsa-miR-320d, hsa-miR-361-5p, hsa-miR-423-5p, hsa-miR-483-5p, hsa-miR-484, hsa-miR-486-5p and hsa-miR-718.
2. Use according to claim 1, characterized in that: The detection sample of the miRNA detection reagent includes a plasma sample.
Citation Information
Patent Citations
Biomarker for prognosis or recurrence early warning evaluation of acute ischemic stroke and application of biomarker
CN114015759A
Novel biomarkers for use in stroke diagnosis and prognosis
WO2021251817A1