Biomarkers and methods for early prediction of intracranial atherosclerotic stenosis

CN122609720APending Publication Date: 2026-08-21HAINAN BOYA BIOTECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610729060.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-26
Publication Date
2026-08-21

AI Technical Summary

Technical Problem

[0003]然而,现有的诊断方法以数字减影血管造影、经颅多普勒超声、CT血管造影等影像学方法为主,这些方法通常具有操作复杂、成本高昂或辐射等局限性,这些缺点往往使人们不愿意在未表现出临床症状时主动接受这些检查

Benefits of technology

(1)本发明的生物标志物可对是否患有颅内动脉粥样硬化性狭窄提前做出有效预测。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122609720A_ABST
    Figure CN122609720A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of biological medicine, and discloses a biomarker and a method for early prediction of intracranial atherosclerotic stenosis, wherein the biomarker comprises SPRED2, TRAK1, NPHP3-ACAD11, ENPP5, LINC00689, PLXDC2, TMEM120B, COL20A1 and OLIG2. The biomarker can effectively predict intracranial atherosclerotic stenosis, and the curative effect can be predicted according to the biomarker; compared with some existing liquid biopsy markers, the biomarker has the advantages of safety, non-invasiveness, easy sample acquisition, high accuracy and convenient operation, and can provide accurate judgment for early prediction of intracranial atherosclerotic stenosis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biomedical technology, specifically to a biomarker and method for early prediction of intracranial atherosclerotic stenosis. Background Technology

[0002] Intracranial atherosclerotic stenosis (ICAS) is generally defined as a narrowing of the intracranial aorta diameter by 50% or more due to atherosclerosis. It is a significant risk factor for stroke and is closely related to the risk of recurrence of cerebrovascular diseases. Currently, the overall prevalence of ICAS in the population is approximately 4–13%. Although ICAS patients are mostly elderly, the disease has been found in all age groups and is showing a trend towards affecting younger people. In Asia, 30–50% of ischemic stroke patients also have ICAS. Therefore, ICAS has become an important predictor of stroke incidence and recurrence.

[0003] However, existing diagnostic methods primarily rely on imaging techniques such as digital subtraction angiography, transcranial Doppler ultrasound, and CT angiography. These methods typically have limitations such as complex operation, high cost, or radiation exposure, which often discourage people from undergoing these examinations when they do not exhibit clinical symptoms. Currently, research on using biomarkers to predict ICAS is very limited, and the accuracy of predictive ICAS is generally low. Therefore, it is essential to discover an accurate method for the early prediction of ICAS. Summary of the Invention

[0004] To address the aforementioned technical shortcomings, the purpose of this invention is to provide a biomarker and method for early prediction of intracranial atherosclerotic stenosis. This biomarker can effectively predict intracranial atherosclerotic stenosis, and the treatment efficacy can be predicted based on the biomarker. It has the advantages of being safe and non-invasive, having readily available samples, high accuracy, and convenient operation, and provides an accurate assessment of intracranial atherosclerotic stenosis.

[0005] To solve the above-mentioned technical problems, the present invention adopts the following technical solution: In a first aspect, the present invention provides a biomarker for early prediction of intracranial atherosclerotic stenosis, the biomarker including SPRED2, TRAK1, NPHP3-ACAD11, ENPP5, LINC00689, PLXDC2, TMEM120B, COL20A1 and OLIG2.

[0006] In a second aspect, the present invention provides the use of the biomarker, including: for constructing a model for early prediction of intracranial atherosclerotic stenosis, and for preparing a product for early prediction of intracranial atherosclerotic stenosis, the product being based on the model for predicting intracranial atherosclerotic stenosis to predict whether or not one has intracranial atherosclerotic stenosis.

[0007] Thirdly, the present invention provides a model for early prediction of intracranial atherosclerotic stenosis, wherein the input variable of the model is the content of the biomarker.

[0008] Furthermore, the method for determining the content of the biomarker is 5hmC high-throughput detection.

[0009] Fourthly, the present invention provides a method for constructing a model for early prediction of intracranial atherosclerotic stenosis, comprising the following steps: Step 1. Samples from multiple patients were tested to obtain 5hmC sequencing data of DNA; Step 2. Perform the sequencing data obtained in Step 1 through first filtering, screening, and second filtering to obtain biomarkers for predicting intracranial atherosclerotic stenosis. Step 3. Based on the six performance evaluation metrics of AUROC, Accuracy, Precision, Recall, AUPRC, and F1Score, the optimal machine learning algorithm combination was selected: Lasso + NaiveBayes. The biomarkers used were: SPRED2, TRAK1, NPHP3-ACAD11, ENPP5, LINC00689, PLXDC2, TMEM120B, COL20A1, and OLIG2. Step 4: Construct a Lasso model based on the screening results of Step 3, and extract the filtered features from the Lasso model; then, construct the final NaiveBayes model using the filtered features.

[0010] Furthermore, the first filtering includes: using bedtools to count the number of reads for each 5hmC peak in all samples, and removing peak information of 5hmC peak regions that appear only in 10 or fewer samples and have a read count of less than 50; The screening process included: using the JonckheereTerpstraTest function in the R package DescTools to perform the Jonckheere-Terpstra test, calculating the 5hmC values ​​with p-values ​​<0.01 for sequencing data from samples from different disease progression groups, and identifying 1413 upregulated and 926 downregulated differentially regulated biomarkers based on the 5hmC gradient; enrichment analysis was performed using the enrichGO function in the R package clusterProfiler to demonstrate that the identified biomarkers were associated with atherosclerosis, thus justifying the rationality of the biomarker screening process; The second filtering includes: for the biomarkers obtained through the screening, a more stringent threshold p-value <1e-5 is used to filter and obtain 14 / 17 biomarkers with significant gradient up / down regulation; Lasso regression and random forest machine learning algorithms are used to calculate the contribution of the feature biomarkers to the model for secondary filtering, resulting in 9 / 12 biomarkers respectively; subsequently, glmBoost, LDA, NaiveBayes and SVM machine learning algorithms are used to construct a model for predicting intracranial atherosclerotic stenosis.

[0011] Further, step 2 includes normalizing the biomarker and inputting it into the prediction model, and comparing the score with a threshold.

[0012] Furthermore, the 5hmC modification level was obtained from the plasma vesicle DNA of the subject being tested.

[0013] Furthermore, the threshold was obtained based on multiple control samples with known conditions of intracranial atherosclerotic stenosis.

[0014] Fifthly, the present invention provides a method for early prediction of intracranial atherosclerotic stenosis, comprising the following steps: 1) Determine the relative abundance of biomarkers in the samples of the subjects being tested; 2) Predict the risk of intracranial atherosclerotic stenosis based on the relative abundance of the biomarkers.

[0015] The beneficial effects of this invention are as follows: (1) The biomarkers of the present invention can make an effective prediction in advance whether or not someone has intracranial atherosclerotic stenosis.

[0016] (2) The prediction model of the present invention has the advantages of high specificity and high sensitivity. By applying the biomarkers and / or models described in the present invention, it is possible to make a safe, non-invasive and highly accurate prediction of whether a patient has intracranial atherosclerotic stenosis.

[0017] (3) According to the biomarkers and models provided by the present invention, the sample for predicting intracranial atherosclerotic stenosis is peripheral blood, which is easy to obtain; the prediction method relies on high-throughput sequencing, which has high detection efficiency; the prediction model is reliable and the prediction results are highly accurate. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0019] Figure 1 This is a prediction performance graph of the prediction model provided by the present invention; wherein (A) is the area under the curve (AUC) of the model predicting intracranial atherosclerotic stenosis; and (B) is the confusion matrix of the model predicting intracranial atherosclerotic stenosis. Detailed Implementation

[0020] The technical solutions of the present invention will be clearly and completely described below with reference to the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0021] This invention provides a biomarker for early prediction of intracranial atherosclerotic stenosis, the biomarker including SPRED2, TRAK1, NPHP3-ACAD11, ENPP5, LINC00689, PLXDC2, TMEM120B, COL20A1 and OLIG2.

[0022] SPRED2 is a negative regulator of the Ras / MAPK signaling pathway, which can inhibit cell proliferation and invasion and is downregulated in various tumors.

[0023] TRAK1 is involved in mitochondrial and intracellular vesicle transport; mutations can lead to epileptic encephalopathy and neurodevelopmental abnormalities.

[0024] NPHP3-ACAD11 is a gene fusion transcript that is associated with ciliary function and fatty acid metabolism, and its abnormal expression is often seen in tumors.

[0025] ENPP5 is an extracellular nucleotide hydrolase that participates in purine signaling and immune regulation, and is associated with the cardiovascular system and tumor microenvironment.

[0026] LINC00689 is a long non-coding RNA that primarily plays a role in tumor suppression, inhibiting tumor proliferation, migration, and invasion.

[0027] PLXDC2 is a transmembrane receptor protein that participates in cell migration and angiogenesis, and regulates tumor immunity and invasion capabilities.

[0028] TMEM120B is a transmembrane protein that participates in cell membrane homeostasis and cell signaling regulation, and is associated with some tumors and cardiovascular diseases.

[0029] COL20A1 is a member of the collagen family, constitutes the extracellular matrix, participates in tissue repair, and is associated with bone and joint diseases.

[0030] OLIG2 is a neurodevelopmental transcription factor that regulates oligodendrocyte differentiation and is an important biomarker for gliomas.

[0031] The present invention provides uses of the biomarker, the uses including: constructing a model for early prediction of intracranial atherosclerotic stenosis, and preparing a product for early prediction of intracranial atherosclerotic stenosis, the product being based on the model for predicting intracranial atherosclerotic stenosis to predict whether or not one has intracranial atherosclerotic stenosis.

[0032] The term "model for constructing early prediction of intracranial atherosclerotic stenosis" refers to the use of data from the aforementioned biomarkers (a combination of SPRED2, TRAK1, NPHP3-ACAD11, ENPP5, LINC00689, PLXDC2, TMEM120B, COL20A1, and OLIG2) to establish a mathematical or computational structure capable of early prediction of intracranial atherosclerotic stenosis through specific algorithms and computational methods.

[0033] This invention provides a model for early prediction of intracranial atherosclerotic stenosis, wherein the input variable of the model is the content of the biomarker.

[0034] Furthermore, the method for determining the content of the biomarker is 5hmC high-throughput detection. Preferably, it is 5hmC-Seal.

[0035] High-throughput 5hmC detection is a technique for large-scale, efficient quantitative analysis of 5-hydroxymethylcytosine (5hmC) modifications in biological samples. As an important epigenetic marker, the distribution and abundance of 5hmC in the genome are closely related to the occurrence and development of various diseases. High-throughput detection allows for the simultaneous analysis of 5hmC levels in a large number of samples or genomic regions, thus obtaining comprehensive and detailed epigenetic information. 5hmC-Seal (Selective Chemical Labeling and Enrichment for 5-Hydroxymethylcytosine) is a highly specific and sensitive high-throughput 5hmC detection method. The core of this technology lies in its unique chemical labeling and enrichment strategy, which can accurately capture and quantify 5hmC sites in the genome. Specifically, 5hmC-Seal technology typically utilizes T4 phage β-glucosyltransferase (T4-βGT) to specifically add glucose groups to 5hmC, forming glucosyl-5hmC. Subsequently, DNA fragments containing 5hmC were efficiently enriched using a biotin-labeled anti-glucosyl-5hmC antibody to separate them from the complex genomic background. Finally, the enriched DNA fragments were subjected to high-throughput sequencing to achieve precise identification and quantification of the 5hmC site.

[0036] This invention provides a method for constructing a model for early prediction of intracranial atherosclerotic stenosis, comprising the following steps: Step 1. Samples from multiple patients were tested to obtain 5hmC sequencing data of DNA; Step 2. The sequencing data obtained in Step 1 are sequentially subjected to a first filter, a second filter, and a third filter to obtain biomarkers for predicting intracranial atherosclerotic stenosis. Step 2 includes normalizing the biomarkers and inputting them into the prediction model, comparing the scores with a threshold. The threshold is obtained based on multiple control samples with known conditions of intracranial atherosclerotic stenosis.

[0037] Step 3. Based on the six performance evaluation metrics of AUROC, Accuracy, Precision, Recall, AUPRC, and F1Score, the optimal machine learning algorithm combination was selected: Lasso + NaiveBayes. The biomarkers used were: SPRED2, TRAK1, NPHP3-ACAD11, ENPP5, LINC00689, PLXDC2, TMEM120B, COL20A1, and OLIG2. Step 4. Pass: cv.Lasso_Model = cv.glmnet(x = Train_set, y = Train_label[["group"]], family = "binomial", nfolds = 10) Lasso_Model = glmnet(x = Train_set, y = Train_label[["group"]], family = "binomial", lambda = cv.Lasso_Model $lambda.min) Construct a Lasso model and then: Features = names(coef(Lasso_Model)) Extract the features filtered by the Lasso model; then, using the filtered features, apply the NaiveBayes algorithm: data = cbind(Train_set, Train_label["group"]) data[["group"]] = as.factor(data[["group"]]) NaiveBayes_Model = naiveBayes(group~., data = data) NaiveBayes_Model$subFeature = colnames(Train_set) Construct the final NaiveBayes model.

[0038] Furthermore, the first filtering includes: using bedtools to count the number of reads for each 5hmC peak in all samples, and removing peak information of 5hmC peak regions that appear only in 10 or fewer samples and have a read count of less than 50; The screening process included: performing the Jonckheere-Terpstra test using the JonckheereTerpstraTest function in the R package DescTools to calculate the 5hmC levels for sequencing data from samples from different disease progression groups, screening for 5hmC with p-values ​​<0.01, and obtaining 1413 upregulated and 926 downregulated differential biomarkers with 5hmC gradients; performing enrichment analysis using the enrichGO function in the R package clusterProfiler to prove that the found biomarkers are associated with atherosclerosis, thereby demonstrating the rationality of the biomarker screening process; the 5hmC modification level was obtained from the plasma vesicle DNA of the subjects being tested.

[0039] The second filtering includes: for the biomarkers obtained through the screening, a more stringent threshold p-value <1e-5 is used to filter and obtain 14 / 17 biomarkers with significant gradient up / down regulation; Lasso regression and random forest machine learning algorithms are used to calculate the contribution of the feature biomarkers to the model for secondary filtering, resulting in 9 / 12 biomarkers respectively; subsequently, glmBoost, LDA, NaiveBayes and SVM machine learning algorithms are used to construct a model for predicting intracranial atherosclerotic stenosis.

[0040] This invention provides a method for early prediction of intracranial atherosclerotic stenosis, comprising the following steps: 1) Determine the relative abundance of biomarkers in the samples of the subjects being tested; 2) Predict the risk of intracranial atherosclerotic stenosis based on the relative abundance of the biomarkers. Example

[0041] This invention provides a model for early prediction of intracranial atherosclerotic stenosis. The relevant parameters of nine 5hmC characteristic biomarkers in the model are shown in Table 1. Table 1

[0042] In the embodiments of the present invention, the 5hmC-Seal high-throughput sequencing method is used to sequence the samples. The 5hmC-Seal high-throughput sequencing method used in the following embodiments is explained as follows: 5hmC-Seal is a high-throughput sequencing method based on 5hmC. This method uses an improved chemical glycosylation marker combined with next-generation high-throughput sequencing technology to obtain the distribution information of 5hmC on genomic DNA.

[0043] Due to the high sensitivity of chemical labeling, the input DNA can be as low as 1-10 ng. The DNA can be fragmented genomic DNA or fragmented evDNAs. According to the requirements of next-generation sequencing, the DNA fragments are padded at both ends, and then an A tail is ligated to the 3' end. Sequencing Y-shaped adapters are ligated to both ends of each DNA fragment using AT specific ligation. These adapters contain index information that distinguishes the sample and sequences such as amplification primers. Next, the 5hmC labeling step was performed. First, UDP-6-N3-Glc was added, and under specific conditions, all 5hmC on the DNA reacted to become N3-5ghmC. Then, DBCO-PEG4-Biotin was added, and all N3-5ghmC were linked to biotin. Finally, through efficient and specific binding of biotin-magnetic beads, all DNA fragments containing 5hmC sites were screened. After PCR amplification and purification, the 5hmC-based DNA library was constructed. The size and distribution of DNA bands in each sample were analyzed using Fragment Analyzer for quality control. After precise quantification of the library by qPCR, high-throughput sequencing was performed using an Illumina Nextseq 500 sequencer to obtain the base sequences of all DNA fragments in the library.

[0044] Application Example 1 115 volunteers were enrolled, including 55 healthy individuals and 60 patients with intracranial atherosclerotic stenosis. Inclusion criteria were: 1. Stenosis of more than 50% in intracranial arteries confirmed by neuroimaging; 2. Age between 18 and 85 years. Exclusion criteria were: 1. Those who did not agree to participate in the experiment; 2. Those with severe cardiac, pulmonary, hepatic, renal, or coagulation dysfunction; 3. Patients with epilepsy; 4. Pregnant patients. Peripheral blood samples (3-4 mL) were collected from the 115 volunteers, and evDNAs were extracted from the plasma for 5hmC-Seal high-throughput sequencing. The sequencing throughput for each sample was 1.5 Gb, and the sequencing band size was 39 bp.

[0045] The 115 samples were randomly divided into a training cohort and a validation cohort at a ratio of 7:3. Peak information of plasma evDNAs from different patients in the training cohort (70 cases) was compared and filtered, with each peak appearing in at least 10 samples. The Jonckheere-Terpstra test was used to analyze sequencing data from samples from different disease progression groups, retaining 5hmC peak regions with more than 50 reads and screening for 5hmC with pvalue < 1e-5, thus obtaining differentially expressed biomarkers for 5hmC upregulation and downregulation. Then, Lasso regression and random forest machine learning algorithms were used for secondary filtering of feature biomarkers. Subsequently, glmBoost, LDA, NaiveBayes and SVM machine learning algorithms were used to build a model for predicting intracranial atherosclerotic stenosis. Then, the optimal combination of machine learning algorithms was selected using six performance evaluation metrics: AUROC, Accuracy, Precision, Recall, AUPRC and F1Score: Lasso + NaiveBayes. The biomarkers used were: SPRED2, TRAK1, NPHP3-ACAD11, ENPP5, LINC00689, PLXDC2, TMEM120B, COL20A1 and OLIG2. Finally, the optimal model constructed using Lasso + NaiveBayes was analyzed using receiver operating characteristic (ROC) analysis. The area under the curve (AUC), specificity, and sensitivity were calculated using pROC to evaluate the model's performance. The results of calculating the AUC using pROC are shown below. Figure 1 As shown, the constructed prediction model was used in Example 2 to predict the prevalence of intracranial atherosclerotic stenosis in the validation group.

[0046] Application Example 2 Peripheral blood sequencing results were obtained from volunteers in the validation group (51 cases, including 14 healthy individuals and 37 patients with intracranial atherosclerotic stenosis) in Application Example 1. The sequencing data were processed and analyzed to screen for 5hmC-enriched sites. Then, based on the screened 5hmC-enriched sites, the therapeutic effect was predicted using the prediction model obtained in Example 1. The prediction results are as follows: Figure 1 As shown, the results indicated that 14 out of 14 healthy individuals were predicted to be healthy, with a specificity of 1.00; and 36 out of 37 patients with ICAS were predicted to have ICAS, with a sensitivity of 0.973.

[0047] Therefore, it can be concluded that the prediction model constructed based on the above nine 5hmC characteristic biomarkers provided by the present invention can effectively predict in advance whether or not someone has intracranial atherosclerotic stenosis.

[0048] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A biomarker for early prediction of intracranial atherosclerotic stenosis, characterized in that, The biomarkers include SPRED2, TRAK1, NPHP3-ACAD11, ENPP5, LINC00689, PLXDC2, TMEM120B, COL20A1, and OLIG2.

2. The use of the biomarker as described in claim 1, characterized in that, The uses include: constructing a model for early prediction of intracranial atherosclerotic stenosis, and preparing a product for early prediction of intracranial atherosclerotic stenosis, the product being based on the model for predicting intracranial atherosclerotic stenosis to predict whether or not one has intracranial atherosclerotic stenosis.

3. A model for early prediction of intracranial atherosclerotic stenosis, characterized in that, The input variable of the model is the content of the biomarker.

4. The model as described in claim 3, characterized in that, The method for determining the content of the biomarker is 5 hmC high-throughput detection.

5. A method for constructing a model for early prediction of intracranial atherosclerotic stenosis, characterized in that, Includes the following steps: Step 1. Samples from multiple patients were tested to obtain 5hmC sequencing data of DNA; Step 2. Perform the sequencing data obtained in Step 1 through first filtering, screening, and second filtering to obtain biomarkers for predicting intracranial atherosclerotic stenosis. Step 3. Based on the six performance evaluation metrics of AUROC, Accuracy, Precision, Recall, AUPRC, and F1Score, the optimal machine learning algorithm combination was selected: Lasso + NaiveBayes. The biomarkers used were: SPRED2, TRAK1, NPHP3-ACAD11, ENPP5, LINC00689, PLXDC2, TMEM120B, COL20A1, and OLIG2. Step 4: Construct a Lasso model based on the screening results of Step 3, and extract the features filtered by the Lasso model; Subsequently, the final NaiveBayes model is constructed using the selected features.

6. The construction method as described in claim 5, characterized in that, The first filtering includes: using bedtools to count the number of reads for each 5hmC peak in all samples, and removing peak information of 5hmC peak regions that appear only in 10 or fewer samples and have a read count of less than 50; The screening process included: using the JonckheereTerpstraTest function in the R package DescTools to perform the Jonckheere-Terpstra test, calculating the sequencing data from samples from different disease progression groups, and screening for 5hmC with p-values ​​< 0.01, obtaining 1413 upregulated and 926 downregulated differential biomarkers with 5hmC gradients; enrichment analysis was performed using the enrichGO function in the R package clusterProfiler to prove that the found biomarkers are associated with atherosclerosis, thus demonstrating the rationality of the biomarker screening process; The second filtering includes: for the biomarkers obtained through the screening, a more stringent threshold p-value < 1e-5 is used to screen for 14 / 17 biomarkers with significant gradient up / down regulation; Lasso regression and random forest machine learning algorithms are used to calculate the contribution of the feature biomarkers to the model for secondary filtering, resulting in 9 / 12 biomarkers respectively; subsequently, glmBoost, LDA, NaiveBayes and SVM machine learning algorithms are used to construct a model for predicting intracranial atherosclerotic stenosis.

7. The construction method as described in claim 5, characterized in that, Step 2 includes normalizing the biomarker and inputting it into the prediction model, then comparing the score with a threshold.

8. The construction method as described in claim 6, characterized in that, The 5hmC modification level was obtained from the plasma vesicle DNA of the subject being tested.

9. The method of claim 7, wherein the threshold is obtained based on multiple control samples with known conditions of intracranial atherosclerotic stenosis.

10. A method for early prediction of intracranial atherosclerotic stenosis, characterized in that, Includes the following steps: 1) Determine the relative abundance of biomarkers in the samples of the subjects being tested; 2) Predict the risk of intracranial atherosclerotic stenosis based on the relative abundance of the biomarkers.