Marker for predicting recurrence of stroke and use thereof
By using methylation markers in the region from position 60727176 to 60733176 on chromosome 13 of the human reference genome hg38, the problem of inaccurate prediction of stroke recurrence in existing technologies has been solved, enabling precise assessment of stroke recurrence risk and personalized treatment guidance, while reducing the harm of invasive testing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BIOCHAIN BEIJING SCI & TECH
- Filing Date
- 2023-04-10
- Publication Date
- 2026-05-12
AI Technical Summary
The lack of effective molecular markers in current technologies for predicting stroke recurrence makes it impossible to accurately assess patients' risk of recurrence in clinical practice, increasing disability and mortality rates.
A methylation biomarker based on the region from position 60727176 to 60733176 of chromosome 13 on the human reference genome hg38 is provided for detecting the risk of stroke recurrence. Target sequences include SEQ ID NO:1, SEQ ID NO:2, etc. This biomarker screens asymptomatic individuals in a non-invasive manner and is used for prediction by combining a generalized linear model and a Cox regression model.
It can sensitively and specifically predict stroke recurrence, provide accurate recurrence risk assessment, guide personalized treatment and health management, reduce the harm of invasive testing, and achieve real-time monitoring.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] This application belongs to the field of molecular biology and relates to gene detection, specifically to a biomarker for predicting stroke recurrence and its uses. Background Technology
[0002] DNA methylation is one of the main mechanisms of epigenetic modification, playing a crucial role in maintaining normal cell function, genetic imprinting, embryonic development, and human tumorigenesis. With technological advancements, methods for studying methylation now cover almost every level, from individual genes to the entire genome. In mammalian cells, DNA methylation is characterized by the addition of a methyl group (-CH3) to the 5th carbon atom of the cytosine ring (5-methylcytosine; 5mC) under the action of DNA methyltransferases (DNMTs). This covalent addition of the methyl group typically occurs within the cytosine of a CpG dinucleotide, which is concentrated in a region called a CpG island. CpG sites appear in clusters at a expected frequency. Aberrant methylation is associated with most diseases, including cancer, neurodegenerative diseases, cardiovascular diseases, and autoimmune diseases. Analyzing DNA methylation patterns is crucial for understanding the underlying molecular mechanisms of these diseases. Furthermore, DNA methylation patterns can serve as a basis for clinical management, such as diagnosis, prognosis, and treatment response.
[0003] Stroke is a major cause of death and disability, and the number of cases is increasing year by year, with the age of onset gradually becoming younger. Clinical data shows that progressive stroke (PS) accounts for 26.4%-34% of acute ischemic stroke patients, and there are no effective treatments, leading to higher disability and mortality rates, placing a heavy burden on families and clinical work. For stroke patients, timely prediction of recurrence can effectively reduce mortality or disability. Currently, the clinical diagnosis of stroke recurrence mainly relies on clinical data assessment, and there are no clear molecular markers. Summary of the Invention
[0004] Based on the problems existing in the detection of stroke recurrence, the purpose of this application is to provide a biomarker, composition, kit and its use for predicting stroke recurrence.
[0005] The specific technical solution of this application is as follows:
[0006] 1. A methylation biomarker for predicting stroke recurrence, comprising a genomic region from position 60727176 to 60733176 on chromosome 13 of the human reference genome hg38.
[0007] 2. The methylation marker according to claim 1, wherein the target sequence of the methylation marker is as shown in SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, or SEQ ID NO:6.
[0008] 3. The methylation marker according to item 1 or 2, wherein the stroke is selected from one or more of large artery atherosclerotic stroke, cardioembolic stroke, small artery occlusive stroke, ischemic stroke and cerebral infarction.
[0009] 4. The methylation marker according to any one of items 1-3, wherein the stroke recurrence is a stroke recurrence within three months.
[0010] 5. Use of any of the methylation markers described in items 1-4 in the preparation of a kit for predicting stroke recurrence.
[0011] 6. Use of any one of the methylation markers in the preparation of a medicament for preventing recurrence of stroke, preferably, the medicament being a targeted drug designed based on the methylation marker.
[0012] 7. A composition for predicting stroke recurrence, said composition comprising:
[0013] Nucleic acid for detecting the methylation status of any one of the methylation markers in items 1-4,
[0014] The methylation state of the methylation marker is characterized by the methylation of the target sequence of the methylation marker.
[0015] 8. A kit comprising reagents for detecting the methylation markers described in any one of items 1-4.
[0016] 9. The kit according to claim 8, wherein the samples for which the kit is used to detect include: cell lines, histological sections, tissue biopsy / paraffin-embedded tissue, body fluids, feces, colonic effluent, urine, plasma, serum, whole blood, isolated blood cells, cells isolated from blood, or combinations thereof.
[0017] 10. A nucleic acid used to target any one of the methylation markers in items 1-4.
[0018] This application has the following beneficial effects:
[0019] The methylation markers proposed in this application can accurately predict the risk of stroke recurrence. Furthermore, through systematic analysis of an individual's prior basic information, epidemiological information, and test results, they can provide precise indications of the likely time point of recurrence, thereby providing guidance for precision treatment and health management, and have significant clinical application value.
[0020] Other features and advantages of this application will be described in detail in the following specific description and claims.
[0021] The effects of the invention
[0022] Based on a large number of stroke samples, this application screened and obtained methylation biomarkers for predicting stroke recurrence. These biomarkers can sensitively and specifically predict stroke recurrence, thereby assessing the risk of stroke recurrence within three months in patients experiencing their first stroke, and enabling better early prevention and treatment.
[0023] The composition of this application is used in a non-invasive manner for screening asymptomatic individuals, reducing the harm caused by invasive testing. The composition has higher sensitivity and accuracy, enabling real-time monitoring. Attached Figure Description
[0024] Figure 1 This is a diagram illustrating the 50% discount grouping;
[0025] Figure 2 AUC curves for 5 test sets of the markers in Example 3;
[0026] Figure 3 The AUC curves for the five test sets of the markers in Example 4. Detailed Implementation
[0027] The present application will now be described in detail. While specific embodiments of the present application are shown, it should be understood that the present application can be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present application and to fully convey the scope of the present application to those skilled in the art.
[0028] Unless otherwise stated, the implementation of this application will employ conventional molecular biology (including recombinant technology), microbiology, cell biology, biochemistry, and genetics techniques, all of which fall within the scope of conventional techniques in the art. Such techniques are described in detail in the literature, such as *Molecular Cloning: A Laboratory Manual*, 2nd edition (Sambrook et al., 1989); *Oligonucleotide Synthesis* (MJ Gait, 1984); *Animal Cell Culture* (RI Freshney, 1987); *Methods in Enzymology* (American Academic Publishing Co., Ltd.); *Current Protocols in Molecular Biology* (FMAusubel et al., 1987, and regularly updated); and *PCR: The Polymerase Chain Reaction* (Mullis et al., 1994). The primers, probes, blocking agents, and kits used in this application can be prepared using standard techniques known in the art.
[0029] Unless otherwise defined, the technical and scientific terms used in this application have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.
[0030] definition
[0031] In this application, "DNA methylation" refers to the addition of a methyl group to the 5-position of cytosine (C), which is usually (but not necessarily) in the case of CpG (cytosine followed by guanine) dinucleotides. As a stable modification state, it can be inherited by newly formed daughter DNA during DNA replication under the action of DNA methyltransferases, and is an important epigenetic mechanism. During DNA methylation, methylation of the gene promoter region can lead to transcriptional silencing of tumor suppressor genes, thus it is closely related to tumorigenesis. Abnormal methylation includes hypermethylation of tumor suppressor genes and DNA repair genes, hypomethylation of repetitive DNA sequences, and loss of imprinting of certain genes, all of which are associated with the development of various tumors.
[0032] As used herein, “increased methylation” or “significant methylation” refers to the presence of at least one methylated cytosine nucleotide in a DNA sequence, wherein the corresponding C in a normal control sample (e.g., a DNA sample extracted from a non-cancer cell or tissue sample or a DNA sample treated to methylate DNA residues) is unmethylated. In some embodiments, at least 2, 3, 4, 5, 6, 7, 8, 9, 10 or more Cs may be methylated, wherein the Cs at these positions in the control DNA sample are unmethylated.
[0033] In the implementation scheme, a variety of different methods can be used to detect DNA methylation alterations. Methods for detecting DNA methylation include, for example, methylation-sensitive restriction endonuclease (MSRE) assays using Southern or polymerase chain reaction (PCR) analysis, methylation-specific or methylation-sensitive PCR (MS-PCR), methylation-sensitive single nucleotide primer extension (Ms-SnuPE), high-resolution melting (HRM) analysis, bisulfite sequencing, pyrosequencing, methylation-specific single-strand conformation analysis (MS-SSCA), combined bisulfite restriction analysis (COBRA), methylation-specific denaturing gradient gel electrophoresis (MS-DGGE), methylation-specific melting curve analysis (MS-MCA), methylation-specific denaturing high-performance liquid chromatography (MS-DHPLC), and methylation-specific microarrays (MSO). These assays can be PCR analysis, quantitative analysis using fluorescent labels, or Southern blot analysis.
[0034] In this application, "methylation assay" refers to any assay that determines the methylation status of one or more CpG dinucleotide sequences within a DNA sequence.
[0035] In this application, "detection" refers to any process of observing a biomarker or biomarker alteration (e.g., a change in the methylation state of a biomarker or the expression level of a nucleic acid or protein sequence) in a sample, regardless of whether the biomarker or biomarker alteration is actually detected. In other words, the act of detecting a biomarker or biomarker alteration in a sample is "detection," even if the biomarker is determined to be absent or below a sensitivity level. Detection can be a quantitative, semi-quantitative, or non-quantitative observation and can be based on comparison with one or more control samples.
[0036] In this application, "amplification" refers to the process of obtaining multiple copies of a nucleic acid from a specific locus, such as genomic DNA or cDNA. Amplification can be achieved using any of a variety of known methods, including but not limited to polymerase chain reaction (PCR), transcription-based amplification, and strand displacement amplification (SDA).
[0037] In this application, "sensitivity" refers to the proportion of cancer detected in a certain cancer sample, and its calculation formula is: Sensitivity = (detected cancer / all cancers), while "specificity" refers to the proportion of normal samples detected in a certain normal sample, and its calculation formula is: Specificity = (detected negative / total negative).
[0038] Nucleic acid molecules can be detected using a variety of different methods. Nucleic acid detection methods include, for example, PCR and nucleic acid hybridization (e.g., Southern blotting, Northern blotting, or in situ hybridization). Specifically, oligonucleotides capable of amplifying target nucleic acids (e.g., oligonucleotide primers) can be used in PCR reactions. PCR methods typically include the following steps: obtaining a sample, isolating nucleic acids (e.g., DNA, RNA, or both) from the sample, and contacting the nucleic acids with one or more oligonucleotide primers that specifically hybridize with the template nucleic acid under conditions that allow amplification of the template nucleic acid to occur. In the presence of the template nucleic acid, an amplification product is generated. The conditions for nucleic acid amplification and detection of the amplification product are known to those skilled in the art. Various improvements to basic PCR techniques have been developed, including but not limited to anchored PCR, RACE PCR, RT-PCR, and ligase chain reaction (LCR). In the amplification reaction, the primer pair must anneal to the opposite strands of the template nucleic acid and should be kept at an appropriate distance from each other so that the polymerase can efficiently polymerize across regions and so that the amplification product can be easily detected, for example, by electrophoresis. For example, computer programs such as OLIGO (Molecular Biology Insights Inc., Cascade, Colo.) can be used to design oligonucleotide primers to facilitate the design of primers with similar melting temperatures. Typically, oligonucleotide primers are 9–30, 40, or 50 nucleotides in length (e.g., lengths of 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 nucleotides), but oligonucleotide primers can be longer or shorter, provided appropriate amplification conditions are used.
[0039] Detection of amplification products or hybridization complexes is typically achieved using detectable labels. The term "label," when referring to nucleic acids, is intended to include both direct labeling of nucleic acids by coupling (i.e., physically linking) a detectable substance to the nucleic acid, and indirect labeling of nucleic acids by reacting with another reagent that has directly labeled the detectable substance. Detectable substances include a variety of enzymes, prosthetic groups, fluorescent materials, cryoluminescent materials, bioluminescent materials, and radioactive materials. Examples of suitable enzymes include horseradish peroxidase, alkaline phosphatase, β-galactosidase, or acetylcholinesterase; examples of suitable prosthetic group complexes include avidin / streptin and avidin / biotin; examples of suitable fluorescent materials include umbelliferone, luciferin, luciferin isothiocyanate, rhodamine, dichlorotriazineamine luciferin, dansyl chloride, or phycoerythrin; examples of cryoluminescent materials include luminol; and examples of bioluminescent materials include luciferase, insect luciferin, and jellyfish protein. Examples of indirect labeling include end-labeling of nucleic acids with biotin, making the nucleic acid detectable using fluorescently labeled avidin streptavidin.
[0040] The "Generalized Linear Model (GLM)" in this application is an extension of the linear model, establishing a relationship between the expected value of the response variable and a linear combination of predictor variables through a connection function. Its key feature is that it does not forcibly alter the natural measure of the data, allowing the data to have a nonlinear and non-constant variance structure. It represents a development of the linear model in studying the non-normal distribution of response values and the concise and direct linear transformation of nonlinear models.
[0041] In this application, "5-fold cross-validation," or "five-fold cross-validation," is a commonly used testing method to assess algorithm accuracy. During validation, the dataset is divided into five parts, with four parts used alternately as training data and one part as test data. Each trial yields a corresponding accuracy (or error rate). The average of the accuracy (or error rate) from the five trials is used as an estimate of the algorithm's precision. Generally, multiple rounds of five-fold cross-validation (e.g., five trials) are performed, and the average is calculated again to further estimate the algorithm's accuracy.
[0042] In this application, the "LASSO algorithm," first proposed by Robert Tibshirani in 1996, is a compression estimator. It constructs a penalty function to obtain a more refined model, compressing some coefficients while setting others to zero. Therefore, it retains the advantage of subset shrinkage and is a biased estimator for handling data with complex collinearity.
[0043] In this application, the "Cox regression model" is also known as the "proportional hazards model (Cox model for short)," a semi-parametric regression model proposed by the British statistician D.R. Cox (1972). This model uses survival outcome and survival time as dependent variables, and can simultaneously analyze the impact of many factors on survival time. It can analyze data with truncated survival time and does not require estimation of the survival distribution type of the data.
[0044] In this application, the ROC curve can reflect the classification performance of the classifier to a certain extent. AUC is essentially the area under the ROC curve. AUC intuitively reflects the classification ability expressed by the ROC curve.
[0045] This application provides a methylation biomarker for predicting stroke recurrence, comprising a genomic region from position 60727176 to 60733176 of chromosome 13 on the human reference genome hg38.
[0046] In one specific implementation, the methylation marker is a genomic region at positions 60727176 to 60733176 on chromosome 13 of the human reference genome hg38.
[0047] The original nucleotides of the genomic region from positions 60727176 to 60733176 on chromosome 13 are as follows:
[0048] SEQ ID NO:1
[0049] GCTAGCTCAAGGTGCCTTGTACCTTGCAGCTGATGCAGGATATTTTTTTTTTTGACCGCTTCGAGGGATGGGTGACAGGG
[0050] ATTCCCCATTTACTCCTCCCACTGTGCTCAACCCCTTGTGGGAGGGAACACATAGGTGAGTGAGTACAGGAGCTAGCCA
[0051] GCTGCTTCAGCACTGGCAGAAGCAAACTCCATTTACTCAGGCCTGCTTCACCCTACCCCTTGTGGGAGGGAGCATGCAG
[0052] GTGAGCAAGTGCAGGAACCAGCTGCTTTGGTGTTGGCAGGAGCAAACTCCATGCAGGCCCTGTGGTAGCATACAGGTG
[0053] CAGGTGCCTGTGACCTCAAGGCCCCAGAGGGTGTGTTACTATGCTCTCTCAGCTCTGCTGTACGCAGATAGCAGTGTGT
[0054] TGTCAGCTCAGTGGGCCCTTTGCCTTGTTGTGTGGGGTGGCTGTCCTCCACTAATGAGGGTAAAGGGCCAGTGTGACAG
[0055] CCTTTTTGGGTACCCACACTTGGTGTATCCTGAATTCTTGTCCAGTGGCCAAGAAGAATGAAGTCACACAGACAAATTG
[0056] AAGGATGGTGAATGCAGACAATTTTATTGAGCACTGAAAGTGGCTCGCAGCAGACAGAGGAGCTGGAAAGAGTACGA
[0057] GAAGGGCAGGTCGCTTTCCTCTGAAGTCAAGTCACCTCTCTGCCTCTCTCCTTCGAAGTCAAGTTGCCTCTCTCTAATGT
[0058] CCAGCTACTTCTCTTAAGTCAAGACACCTCTTTTTGATGTCCAGCCACTTCTCCTCTCTACTGGCTGAGTCTAGGGTCTTT
[0059] GTAGGCACAGGATGGAGGGTGGGGCAGGCCTTAGGTAGTTTTGGAAAAGGCAGCATTCAATTAGTAAAAATACAATAT
[0060] TCAGAAAGAACCAATTGGGATAGAGTGGGCAAACAGGGATAGAAGTTCTCACTTTGGGCCATGGGTTTCAGGCTTTTCA
[0061] GCTTGAAGGTGGGGCTTTGGCAGTGACCTGTCCCTGTCAGCTGAGAGTTTCTCTGCCTCTTGTCTCTATCACAGCCACTA
[0062] GACTGTGGGCCAAAAGAAGCCATTGTTGGGAAGGGTATCCCATGCCATTGATGTATCACATTTCTCCTTTCTTGGTCATC
[0063] TCCTAAGAGCATTGCAACGCTCCTCTAAGTATTTCCAACTCTTTGCATTCAGAGTGGAACAGACTTTGACAGTAATGTTT
[0064] TCTGTAGAAAAGGCTCTGCTTGCCAGAATACTCAGTTCCTATTCTCATTCTCTTCCTCTCTTCCCTCATTATTACATTTTA
[0065] GCAATAGGTTGAACATGAAAAAATAGCAGAAATGAAGGGTTAATGATTAAAGTACTTTAAAACCTTTACTCATAGAAA
[0066] GAAGGAAATTACATGTTTGATCTGTCTCTTACTTCAATGATTAAATCTTTTTGCGAAATTGTAACAATGTACTCATAAGA
[0067] ATTATTCTCTCTCTTCAATGTGAAAGAACACTTTACAATTAAACCCATTGCATCCCATCAAGGTTTCCTCTTAAGCAACA
[0068] TAATTGTTCTGTCCTATTCGTGGTATATTTTAATATTCTCCTTGAGACTAAATTAATTAGGTTTTCATGTTGTCAAATGTT
[0069] TTTGAAATGACCTGTATATGACAACGTAGACATTGTGATTTCTAACTACAAGAGAGAAAGTATAATGAAAGAAAATATC
[0070] TTTTTGTTACTAAGTTTGTAGGAAGAAATTTGGGATTCTTAAGCAGGGAAGAATGCTATTATGTATTATTTATCTTCCTA
[0071] AGCATGATATTTAGAATTTCCTAAAACATACTTCTTGGAAATCTGTGTTTACCTTCTTTCAACCATGATCCATGGAAATT
[0072] GTTCTTTTTGTGAATGATAGTTTGTTATTAGTTTGAATTCTTAAAAAGTGTGTTTTTAATTTTTATGGATACATAATAGTC
[0073] ATACATATTTATGATAAAAAGGATGCTATAATGTGACATAGTTCCTGGGAAGTGAACTCTGTACAGCTTCAAGCATGTC
[0074] AGTTTCTGTCTGTTGTGATTTATCTATGTCTAGCAATGGAAGGATACATGTTGCATTGTATAAAATGTATTATGGTGACT
[0075] TTCTGGGCCAGGCTACAACCTAGGAAGAGTAGTAATTACTAACCACACAGAGGCACCCACTTCCTTAATTTTTGTCTTA
[0076] ATCATTTATTTTTACTTTAGTTTTTTTTTTTTTTTTTACTGGAAATAAAATTTGCAGCTTCAGGACTTAATATTTTAATCCG
[0077] GCACCATTTAGAAATTTTCAAACTGTTAGATCACCAGCCTTAGAAACCTTATGAATTTCCTCCACTTTAGTAGGAGCAAA
[0078] AGTGAGAAAGCCAGCTCAGTTATGGCCAATTTGGCAATTTCCAGACTGATCGTCTGTGTCCAGCCTGTAGTTTTTCCTTC
[0079] TATTTTCTTCTATATGATAATTTGCAATGAAAAAATTATAGTTTAGAGAATACTTATTTTCAACTTACTATATACCTGCTT
[0080] GCATAAAAATAGTTACTATAAGCATAAATCTAGAGGAATTCTGATGTCGTTATGTTTCTAGTGCCTTTGGGAATTTGGA
[0081] GAGGGCTTGGCTCCATTGTGCCTTGAAAAAACTAATTTTCTTTAAGGTATTTGTTAACAAATTGGTGTTTGTTCATTTCA
[0082] GTACAGCCTTTTGTAAACCCTAAATTTAGCTAGGAGGTGGGGCAGGTGGGAGATTCATAGTCACTACATAAAAGAGCTT
[0083] CAAGAAGATGTTATCAATTTAGAGGATGACTTTTGTGTTGGTTTTCTTCCTTTTATTTGAGTGAAATACTTTTGGAAAGA
[0084] AAGACTGAAAGAAGAAGAAGGAAAAAAAGGAAGGAGGGAAGAAGAAAAAAAGCAAGCAGGATTTCTATTAGTAGTT
[0085] TTCTCTACATAGACTAGGAAAATACAATAAAATAATGGCCAATTTATAAACCTTGAGCATTTATTGTTGATCATTACTGA
[0086] TGATGGATTCTACTATAGGTTCATCATGTATAGGTAGATATGTCCAAAGTCCATTGGCCTTATTGGAACAATGCAAATG
[0087] AAAATTAGGACCGGCATGAACATCTGTATATACGAACAATCATTTTATTGTAGAGCTTTAATAAACATTACTGATCTGT
[0088] AGGGTGTTACTTTTCTTCAACAGAAATTATTCATTTTTTATCTTTTTTTTTTCTTTTTTTGCTAACTGTTCTGCTTGGGAATTG
[0089] TAGTAAGACACTGTGTTAACTTTTAATGCAGTTACTTAGAACACTTCTCTCTCCCAAAACCTCCTGCAGGTTAGCCCTAA
[0090] GAAGCTTACCATCATTCTCATCACTGTAATCCACCAGAGAGAAGTAAACATGGCTTTGGCCCTTGGAAAAGTGATTTC
[0091] AACTCTAACTTTGTGCTAGGACATTAAAAACCACAGAAAGAAAGTCAGTGCAATTATCTTGTAGGCATACAGGCTGCC
[0092] ACCAAGTGTGGATAATGTCTGACCAAGCTATAAATCAGCAGAGCATGACAGGCCAATTCCCATTATCAAGACCCTCATG
[0093] GGGCATCATAGTTTCATTGTAGTGGAATTTTGTTATTTTCAGAAACTTCATAAAATTTCTAGTTTTCAGGACTATAGAT
[0094] GACATATGTGGCATCCAGTTTTCTGGAGGTCGTAAGAAGGTAGAAAAAGACCTACCGTGGAAGTAATAAAAGGTACAG
[0095] CAAGATTGTTGAATATAAGATCAACATATCAGAATCAATTTCTTTCCTACAACCTGGCAATAACCCATTAGAAAATAGA
[0096] ATAAAAGAAGGATTTCGTTTAGAATTGCAATAATATCTATATAGTGTCTAAGAATAAACCTAACAAAGAATGCAAGA
[0097] GCTTCAAGGAGAAATGTTTAATGAAAGGCTAACTTTAAATTAATACATTTACCATGATACATTTATATCCAGAATCAC
[0098] CCACAAAATATATTTCAATGTATTTAATATAAAAGTATAAAGAATGGACTAGAAGGGATATGTATCAAACTTACGCTTTG
[0099] GAAATTAGGTTTGGAACTGATATTGGAATGATGGTCAAAGGAGACTTCAGTTTTATTTTGTATTACAAAAAAACTTTTT
[0100] GTTTTATTTGCAATCTTCAATTTTTTAAAGAATTCACTCATATATTGTGTATTGAAAAAGTTTATACATGGAAATATTAG
[0101] AACACTAGAAGCAAATAGATTAGAATACCAATTCTGTATGCTGAGAATATGGGTGACTCTTATTTTCATTTTCTTTTTTG
[0102] TTTTAATATTTTCCAAAATATTATCTCCCATTGTTTCAGAAAAAAAGGTGACTTGCTATAAATTCAGTCAGAGAGATTAG
[0103] GGGCTAAAGCCAGAATGTTAGGAAGAGGAGAAAAGAGCCAGGTGCATCACGACATTGTAAAAAAAAAAAAGTTAT
[0104] AAAACAAAACAGAATTTGTACTCATCAAGCCAGGGTTGATGGTCTTCACTATGGAATTCAGGCAAATTGTTTGCCCAAT
[0105] GTGGTTAATGATGAGGTAGGGTCCCTCCATTTCCCACCGTTGTGCTTTTCCTGTTAACTGCTAGACCGAAAGGGGATCCC
[0106] TTCTAAAAGTTCTTTCTAACCTCCTAGAATATTACCTGACATTAAGTTTGTATGTTTACAATTGCTTTTCTCAGTTAAATA
[0107] ATTATGGATAAAAACTTTGGTTCTATTGATTCAATACCTAGCTGTGCTTTTTTTCCAAGTATAATTCTATTTCCTTGTACT
[0108] TAATGGAAAGCAGCATTATTATCCAGAGAATATCGAGACCATGATCTTCCCTACATAGAAGAGGCAAAAACAGAGGGA
[0109] GGGAGAGAAGAAAAGCACTGTAATTCACTTGAGAAATGGCAACAGGAGAAAATAACTGTGGCTGGGCAACCTTGAAA
[0110] AGTACAATATGTTAAAGGTAGTTAGCAGTAATACTATTTTTACTGGACAGTTCTATGTAGGCACAGAAGCGATTAGAAT
[0111] CGGGGGAAGGCAGAGATGAAGCCAACAGGCCCAGGGTTGCAAGGCCTCTTATCCAGTTACTCAGAATATGAATGGAGC
[0112] TGTTTATGTCTCTGCAGTGAGCAGGTTTCATTTCTCTGATGCAGTGTGTCTCCCGGGCTTGTTTCCTCACCTTCTTTCCCTT
[0113] TTCAGCATTGCCCTTCTCTGGTGTTCACCATGGCTACCTCTTCAGCTGCACCTAGGACCCTGGAGACTCATCTCAAACAT
[0114] TCTACAAGCAGCAACTCCCAGCTAGTTTTAACTGTTCAACCTAGAATGGAGACTGCATGTATTTAAAGAGAGGTATCAG
[0115] AGAGTTTGGCCATTTCATAATGACTCAAAAGCTTTAGGTATCACATTTTAGATTACTTCAGAGAAGGAGAAGCTGTGGC
[0116] TATCATAAGTGAGTGTTATACCTGCATGCCCCTCTTCTGTCTGGGTGCATTTTAAATATGAGATAGATAAATCCTTTCTG
[0117] GAAAGCACAGACATAAAAAAATCAATAATCAATTTACTTTTACCTATTGCTTCACAGATAATAAAGCAGAGGAATTTCA
[0118] CTTTTCTTCATAAGAATTTTTCTTTTCCCCAAATACAAATTTTCTCTCCAAAGCTTTTCTTTTAGTTCTATAGAAGAAGTA
[0119] GGCTTCAAGTAGATTCTAATATGTTGTTAATTCCCATGGTCCTCATTTTAATGAGTGGCACTGGTTCAAGGTAGAAACGT
[0120] CATAGCAGCTAGCCTGCAGCTAGGGGCCAGGGCCTAGAGTGGCCCCAGTGGATGGGAGCAGGTCGCTAGTACTGACAG
[0121] TGAACATTCAGAACAGTGAGTTAGGAATCTGTGAACCCAGAGGGCTTCCTATCTATCAGTAGTTATGAATCTCACATTT
[0122] ATGGATTTGGTGCTTGTACTTCAGAGAAAATGAAGGCTGTTTCATAAAAGAACATTGCTGGAAAAATTTCATACAGAAT
[0123] ATGCTTTTCATAATAGCAATGAAAAGTAGAAGTTGTGTGGAGAATAAAATCACCATATGCAGAAAAAAAGGTTTAACT
[0124] GAAAATTCAGCAGTGAATACAATAATAAGCAGAAGGCATTTTAT
[0125] The highly methylated sequence obtained by bisulfite treatment of the original sequence is as follows:
[0126] SEQ ID NO:2
[0127] GTTAGTTTAAGGTGTTTTTGTATTTTGTAGTTGATGTAGGATATTTTTTTTTTTTGATCGTTTCGAGGGATGGGTGATAG
[0128] GGATTTTTTATTTATTTTTTTTATTGTGTTTAATTTTTTGTGGGAGGGAATATATAGGTGAGTGAGTATAGGAGTTTAGTTA
[0129] GTTGTTTTAGTATTGGTAGAAGTAAATTTAATTTTATTTAGGTTTGTTTTATTTTATTTTTGTGGGAGGGATTAGGT
[0130] GAGTAAGTGTAGGAATTAGTTGTTTTGGTGTTGGTAGGAGTAAATTTTTATGTAGGGTTTTGTGGTAGTATATAGGTGTAGG
[0131] TGTTTGTGATTTTAAGGTTTTAGAGGGTGTGTTATTATGTTTTTTTAGTTTTGTTGTACGTAGATAGTAGTGTGTGGTTAG
[0132] TTTAGTGGGTTTTTTGTTTTGTTGTGTGTGGGTGGTTGTTTTTATTAATGAGGGTAAAGGGTTAGTGTGTATAGTTTTTTTG
[0133] GGTATTTATATTTGGTGTATTTTGAATTTTTGTTTAGTGGTTAAGAAGAATGAAGTTATATAGATAAATTGAAGGATGGT
[0134] GAATGTAGATAATTTTATTGAGTATTGAAAGTGGTTCGTAGTAGATAGAGGAGTTGGAAAGAGTACGAGAAGGTAGG
[0135] TCGTTTTTTTTTGAAGTTAAGTTTATTTTTTGTTTTTTTTTTCGAAGTTAAGTTGTTTTTTTTTAATGTTTAGTTATTTTTT
[0136] TAAGTTAAGATATTTTTTTTTGATGTTTAGTTATTTTTTTTTTTTATTGGTTGAGTTTAGGGTTTTTGTAGGTATAGGATGG
[0137] AGGTGGGGTAGGTTTTAGGTAGTTTTGGAAAAGGTAGTATTTAATTAGTAAAAATATAATATTTAGAAAGAATTAATT
[0138] GGGATAGAGTGGGTAAATAGGGATAGAAGTTTTTATTTTGGGTTATGGGTTTTAGGTTTTTTAGTTTGAAGGTGGGGTTT
[0139] TGGTAGTGATTTGTTTTTGTTAGTTGAGAGTTTTTTTGTTTTTGTTTTTATTATAGTTATTAGATTGTGGGTTAAAAGAA
[0140] GTTATTGTTGGGAAGGGTATTTTATGTTATTGATGTATTATATTTTTTTTTTTTTTGGTTATTTTTTAAGAGTATTGTAACGT
[0141] TTTTTTAAGTATTTTTAATTTTTTGTATTTAGAGTGGAATAGATTTTGATAGTAATGTTTTTTGTAGAAAAGGTTTTGTTT
[0142] GTTAGAATATTTAGTTTTTATTTTTTATTTTTTTTTTTTTTTTTTTTATTATTATTATATTTTAGTAATAGGTTGAATATGAAAAA
[0143] ATAGTAGAAATGAAGGGTTAATGATTAAAGTATTTTAAAATTTTTATTTATAGAAAGAAGGAAATTATATGTTTGATTT
[0144] GTTTTTTATTTTAATGATTAAATTTTTTTGCGAAATTGTAATAATGTATTTATAAAGAATTATTTTTTTTTTTTTTAATGTGAAA
[0145] GAATATTTTATAATTAAATTTATTGTATTTTATTATTAAGGTTTTTTTTTAAGTAATATAATTGTTTTGTTTTATTCGTGGTATA
[0146] TTTTAATATTTTTTTTGAGATTAAATTAATTAGGTTTTTATGTTGTTAAATGTTTTTGAAATGATTTGTATATGATAACGT
[0147] AGATATTGTGATTTTTAATTATAAGAGAGAAAGTATAATGAAAGAAAATATTTTTTTGTTATTAAGTTTGTAGGAAGAA
[0148] ATTTGGGATTTTTAAGTAGGGAAGAATGTTATTATGTATTATTTATTTTTTTAAGTATGATATTTAGAATTTTTTAAAATA
[0149] TATTTTTTGGAAATTTGTGTTTATTTTTTTTTAATTATGATTTATGGAAATTGTTTTTTTTGTGAAATGATAGTTTGTTATTA
[0150] GTTTGAATTTTTAAAAAGTGTGTTTTTAATTTTTATGGATATATAATAGTTATATATTTATGATAAAAAGGATGTTAT
[0151] AATGTGATATAGTTTTTGGGAAGTGAATTTTGTATAGTTTTAAGTATGTTAGTTTTTGTTTGTTGTGATTTATTTATGTTT
[0152] AGTAATGGAAGGATATATGTTGTATTGTATAAAATGTATTATGGTGATTTTTTGGGTTAGGTTATAATTTAGGAAGAGTA
[0153] GTAATTATTAATTATATAGAGGTATTTATTTTTTTAATTTTTTGTTTTAATTATTTATTTTTATTTTTAGTTTTTTTTTTTTTTTT
[0154] TATTGGAAATAAAATTTGTAGTTTTAGGATTTAATATTTTAATTCGGTATTATTTAGAAATTTTTAAATTGTTAGATTATT
[0155] AGTTTTAGAAATTTTATGAATTTTTTTATTTTTAGTAGGAGTAAAAGTGAGAAGTTAGTTTAGTTATGGTTAATTTGGT
[0156] AATTTTTAGATTGATCGTTTGTGTTTAGTTTGTAGTTTTTTTTTTATTTTTTTTATATGATAATTTGTAATGAAAAAATT
[0157] ATAGTTTAGAGAAATATTATTTTTAATTTTATTATATATTGTTTGTATAAAAATAGTTATTATAAGTATAAATTTAGAGG
[0158] AATTTTGATGTCGTTATGTTTTTAGGTTTTTGGGAATTTGGAGAGGGTTTGGTTTTATTGTGTTTTGAAAAAATTAATTT
[0159] TTTTTAAGGTATTGTTAATAAATTGGTGTTTGTTTATTTTAGTATAGTTTTTGTAAATTTAAATTTAGTTAGGAGGTG
[0160] GGGTAGGTGGGAGATTTATAGTTTTATATAAAAGAGTTTAAGAAGATGTTATTAATTTAGAGGATGATTTTTGGTTG
[0161] GTTTTTTTTTTTTTTTTGAGTGAAATATTTTTGGAAAGAAAGATTGAAAGAAAGAAAGGAAAAAAAGGAAGGAGGG
[0162] AAGAAGAAAAAAAGTAAGTAGGATTTTTTATTAGTAGTTTTTTTTATATAGATTAGGAAAATAATAATAAAATAATGGTTA
[0163] ATTTATAAATTTTGAGTATTTATTGTTGATTATTATTGATGATGGATTTTATTATAGGTTTATTATGTATAGGTAGATATG
[0164] TTTAAAGTTTATTGGTTTTATTGGAATAATGTAAATGAAAATTAGGATCGGTATGAATATTTGTATATACGAATAATTAT
[0165] TTTATTGTAGAGTTTTAATAAATATTATTGATTTTGTAGGGGTGTTATTTTTTTTTAATAGAAATATTATTTATTTTTTATTTTTTTTTT
[0166] TTTTTTTTTTTTGTTAATTGTTTTGTTTGGGAATTGTAGTAAGATATTGTGTTAATTTTTAATGTAGTTATTTAATATTT
[0167] TTTTTTTTAAAATTTTTTGTAGGTTAGTTTTAAGAAGTTTATTATTATTTTTTATTATTGTAATTTATTAGAGGAGAAGTA
[0168] AATTGGTTTTGGTTTTTGGAAGGTGATTTTAATTTTAATTTTGTGTTAGGATATTAAAAAATTATAGAAAGAAAGTTA
[0169] GTGTAATTATTTTGTAGGTATATAGGTTGTTATTAAGTGTGGATAATGTTTGATTAAGTTATAAATTAGTAGAGTATGAT
[0170] AGGTTAATTTTTATTATTAAGATTTTTATGGGGTATTATAGTTTTATTGTAGTGGAATTTTGTTATTTTTTGAAATTTTAT
[0171] AAAATTATTTAGTTTTTAGGATTATAGATGATATATGTGGTATTTAGTTTTTTGGAGGTCGTAAGAAGGTAGAAAAGATT
[0172] TATCGTGGAAGTAATAAAAGAGTATAGTAAGATTGTTGAATATAAGATTAATATATTAGAATTAATTTTTTTTTTATAAT
[0173] TTGGTAATAATTTATTAGAAAATAGAATAAAAGAAAGGATTTCGTTTAGAATTGTAATAATATTTATATAGTGTTTAAG
[0174] AATAAATTTAATAAAGAATGTAAGAGGTTTTAAGGAGAAATGTTTAATGAAAAGGTTAATTTTAAATTAATATTTATT
[0175] ATGATATATTTATATTTAGAATTATTTATAAAATATATTTTAATGTATTTAATATAAAAGTATAAAGATGGATTAGAAGG
[0176] GATATGTATTAAATTTACGTTTTGGAAATTAGGTTTGGAATTGATATTGGAATGATGGTTAAAGGAGATTTTAGTTTTAT
[0177] TTTTGTATTATAAAAATTTTTTGTTTTATTTGTAATTTTTAATTTTTTTAAAGAATTTATTTATATATTGTGTATTGAAAA
[0178] AGTTTATATATGGAAATATTAGAATATTAGAAGTAAATAGATTAGAATATTAATTTTGTATGTTGAGAATATGGGTGAT
[0179] TTTTATTTTTATTTTTTTTTTTGTTTTAATATTTTTTTAAATATTATTTTTTATTGTTTTAGAAAAAGGTGATTTGTTATAA
[0180] ATTTAGTTAGAGAAGATTAGGGGTTAAAGTTAGAATGTTAGGAAGAGGAGAAAAGAGTTAGGTGTATTACGATATTGT
[0181] AAAAAAAAAAAAAAGTTATAATAAAATAAAATAGAATTTGTATTTATTAAGTTAGGGTTGATGGTTTTTATTATGGAATTTA
[0182] GGTAAATTGTTTGTTTAATGTGGTTAATGATGAGTAGGGTTTTTTTATTTTTTATCGTTGTGTTTTTTTTGTTAATTGTTAG
[0183] ATCGAAAAAGGGATTTTTTTTAAAAGTTTTTTTTAATTTTTTAGAATATTATTTGATATTAAGTTTGTATGTTTATAATTG
[0184] TTTTTTTTAGTTAAATAATTATGGATAAAAATTTTGGTTTTATTGATTTAATATTTAGTTGTGTTTTTTTTTTAAGTATAAT
[0185] TTTATTTTTTTGTATTTAATGGAAAGTAGTATTATTATTTAGAGAATATCGAGATTATGATTTTTTTTATATAGAAGAGGT
[0186] AAAAATAGAGGGAGGGAGAGAAGAAAAGTATTGTAATTTATTTGAGAAATGGTAATAGGAGAAAATAATTGTGGTTGG
[0187] GTAATTTTGAAAAGTATAATATGTTAAAGGTAGTTAGTAGTAATATTATTTTTATTGGATAGTTTTATGTAGGTATAGAA
[0188] GCGATTAGAATCGGGGGAAGGTAGAGATGAAGTTAATAGGTTTAGGGTTGTAAGGTTTTTTATTTAGTTATTTAGAATA
[0189] TGAATGGAGTTGTTTATGTTTTTGTAGTGAGTAGGTTTTATTTTTTTGATGTAGTGTGTTTTTCGGGTTTGTTTTTTTATTT
[0190] TTTTTTTTTTTTTTAGTATTGTTTTTTTTTGGTGTTTATTATGGTTATTTTTTTAGTTGTATTTAGGATTTTGGAGATTTATTT
[0191] TAAATATTTTATAAGTAGTAATTTTTAGTTAGTTTTAATTGTTTAATTTAGAATGGAGATTGTATGTATTTAAAGAGAGG
[0192] TATTAGAGATTTGGTTATTTTATAATGATTTAAAAGTTTTAGGTATTATTATTTTAGATTATTTTAGAGAAGGAGAAGTT
[0193] GTGGTTATTATAAGTGAGTGTTATATTTGTATGTTTTTTTTTTGTTTGGGTGTATTTTAAATATGAGATAGATAAATTTTT
[0194] TTTGGAAAGTATAGATAAAAAAAATTAATAATTAATTTATTTTTATTTATTGTTTTATAGATAATAAAGTAGAGGAATT
[0195] TTATTTTTTTTTAAGAATTTTTTTTTTTTTAAAATAAATTTTTTTTTTAAAGTTTTTTTTTAGTTTTATAGAAGAAGT
[0196] AGGTTTTAAGTAGATTTTAATATGTTGTTAATTTTTATGGTTTTTATTTTAATGAGTGGTATTGGTTTAAGGTAGAAACGT
[0197] TATAGTAGTTAGTTTGTAGTTAGGGGTTAGGGTTTAGAGTGGTTTTAGTGGATGGGAGTAGGTCGTTAGTATTGATAGT
[0198] GAATATTTAGAATAGTGAGTTAGGAATTTGTGAATTTAGAGGGTTTTTTATTTATTAGTAGTTATGAATTTTATATTTAT
[0199] GGATTTGGTGTTTGTATTTTAGAGAAAATGAAGGTTGTTTTATAAAAGAATATTGTTGGAAAAATTTTATATAGAATATG
[0200] TTTTTTATAATAGTAATGAAAAGTAGAAGTTGTGTGGAGAATAAAATTATTATATGTAGAAAAAAAGGTTTAATTGAAA
[0201] ATTTAGTAGTGAATATAATAATAAGTAGAAGGTATTTTAT
[0202] The hypomethylated sequence obtained by bisulfite treatment of the original sequence is as follows:
[0203] SEQ ID NO:3
[0204]
[0205] The reverse complementary sequence of the original sequence is as follows:
[0206] SEQ ID NO:4
[0207] ATAAAATGCCTTCTGCTTATTATTGTATTCACTGCTGAATTTTCAGTTAAACCTTTTTTTCTGCATATGGTGATTTTA
[0208] TTCTCCACACAACTTCTACTTTTCATTGCTATTATGAAAAGCATATTCTGTATGAAATTTTTCCAGCAATGTTCTTTTATG
[0209] AAACAGCCTTCATTTTCTCTGAAGTACAAGCACCAAATCCATAAATGTGAGATTCATAACTACTGATAGATAGGAAGCC
[0210] CTCTGGGTTCACAGATTCCTAACTCACTGTTCTGAATGTTCACTGTCAGTACTAGCGACCTGCTCCCATCCACTGGGGCC
[0211] ACTCTAGGCCCTGGCCCCTAGCTGCAGGCTAGCTGCTATGACGTTTCTACCTTGAACCAGTGCCACTCATTAAAATGAG
[0212] GACCATGGGAATTAACAACATATTAGAATCTACTTGAAGCCTACTTCTTCTATAGAACTAAAAGAAAAGCTTTGGAGAG
[0213] AAAATTTGTATTTGGGGAAAAGAAAAATTCTTATGAAGAAAAGTGAAATTCCTCTGCTTTATTATCTGTGAAGCAATAG
[0214] GTAAAAGTAAATTGATTATTGATTTTTTTATGTCTGTGCTTTCCAGAAAGGATTTATCTATCTCATATTTAAAATGCACCC
[0215] AGACAGAAGAGGGGCATGCAGGTATAACACTCACTTATGATAGCCACAGCTTCTCCTTCTCTGAAGTAATCTAAAATGT
[0216] GATACCTAAAGCTTTTGAGTCATTATGAAATGGCCAAACTCTCTGATACCTCTCTTTAAATACATGCAGTCTCCATTCTA
[0217] GGTTGAACAGTTAAACTAGCTGGGAGTTGCTGCTTGTAGAATGTTTGAGATGAGTCTCCAGGGTCCTAGTGGCAGCTG
[0218] AAGAGGTAGCCATGGTGAACACCAGAGAGGGCAATGCTGAAAAGGGGAAAAGAGGTGAGGAAAACAAGCCCGGGAGA
[0219] CACACTGCATCAGAGAAATGAACCTGCTCACTGCAGAGACATAAACAGCTCCATTCATATTCTGAGTAACTGGATAAG
[0220] AGGCCTTGCAACCCTGGGCCTGTTGGCTTCATCTCTGCCTTCCCCCGATTCTAATCGCTTCTGTGCCTACATAGAACTGT
[0221] CCAGTAAAAATAGTATTACTGCTAACTACCTTTAACATATTGTACTTTTCAAGGTTGCCCAGCCACAGTTATTTCTCCT
[0222] GTTGCCATTTCCAAGTGAATTACAGTGCTTTTCTTCTCTCCCTCCCTCTGTTTTTGCCTCTTCTATGTAGGGAAGATCAT
[0223] GGTCTCGATATTCTCTGGATAATAATGCTGCTTTCCATTAAGTACAAGGAAATAGAATTATACTTGGAAAAAAAAGCACA
[0224] GCTAGGTATTGAATCAATAGAACCAAAGTTTTTATCCATAAATTATTTAACTGAGAAAAGCAATTGTAAACATACAAACT
[0225] TAATGTCAGGTAATATTCTAGGAGGTTAGAAAGAACTTTTAGAAGGGATCCCTTTTTCGGTCTAGCAGTTAACAGGAAA
[0226] AGCACAACGGTGGGAAAATGGAGGGACCCTACTCATCATTAACCACATTGGGCAAACAATTTGCCTGAATTCCATAGTG
[0227] AAGACCATCAACCCTGGCTTGATGAGTACAAATTCTGTTTTGTTTTATAACTTTTTTTTTTTTTTACAATGTCGTGATGCA
[0228] CCTGGCTCTTTTCTCCTCTTCCTAACATTCTGGCTTTAGCCCCTAATCTTCTCTGACTGAATTTATAGCAAGTCACCTTTTT
[0229] TTCTGAAACAATGGGAGATAATATTTGGAAAATATTAAAAAAAAAAAGAAATGAAAATAAGAGTCACCCATATTCTC
[0230] AGCATACAGAATTGGTATTCTAATCTATTTGCTTCTAGTGTTCTAATATTTCCATGTATAAACTTTTTCAATACACAATAT
[0231] ATGAGTGAATTCTTTTAAAAAATTGAAGATTGCAAATAAAAAAAAAAGTTTTTTGTAATACAAAGATAAAACTGAAGTC
[0232] TCCTTTGACCATCATTCCAATATCAGTTCCAAACCTAATTTCCAAAGCGTAAGTTTGATACATATCCCTTCTAGTCCATCT
[0233] TTATACTTTTATATTAAATACATTGAAATATATTTTGTGGGTGATTCTGGATATAAATGTATCATGGTAAATGTATTAAT
[0234] TTAAAGTTAGCCTTTTCATTAAACATTTCTCCTTGAAGCTCTTGCATTCTTTGTTAGGTTTATTCTTAGACACTATATAGA
[0235] TATTATTGCAATTCTAAACGAAATCCTTTCTTTTATTCTATTTTCTAATGGGTTATTGCCAGGTTGTAGGAAAGAAATTG
[0236] ATTCTGATATGTTGATCTTATATTCAACAATCTTGCTGTACTCTTTTATTACTTCCACGGTAGGTCTTTTCTACCTTCTTAC
[0237] GACCTCCAGAAAACTGGATGCCACATATGTCATCTATAGTCCTGAAAACTAGATAATTTTATGAAGTTTCTGAAAATAA
[0238] CAAAATTCCACTACAATGAAACTATGATGCCCCATGAGGGTCTTGATAATGGGAATTGGCCTGTCATGCTCTGCTGATT
[0239] TATAGCTTGGTCAGACATTATCCACACTTGGTGGCAGCCTGTATGCCTACAAGATAATTGCACTGACTTTCTTTCTGTGG
[0240] TTTTTTAATGTCCTAGCACAAAGTTAGAGTTGAAATCACTTTTCCAAGGGCCAAAGCCATGTTTACTTCTCCTCTGGTGG
[0241] ATTACAGTGATGAGAATGATGGTAAGCTTCTTAGGGCTAACCTGCAGGAGGTTTTGGGAGAGAGAAGTGTTCTAAGTA
[0242] ACTGCATTAAAAGTTAACACAGTGTCTTACTACAATTCCCAAGCAGAACAGTTAGCAAAAAAAGAAAAAAAAGATAAA
[0243] AAATGAATAATTTCTGTTGAAGAAAAGTAACACCCTACAGATCAGTAATGTTTATTAAAGCTCTACAATAAAATGATTG
[0244] TTCGTATATACAGATGTTCATGCCGGTCCTAATTTTCATTTGCATTGTTCCAATAAGGCCAATGGACTTTGGACATATCT
[0245] ACCTATACATGATGAACCTATAGTAGAATCCATCATCAGTAATGATCAACAATAAATGCTCAAGGTTTATAAATTGGCC
[0246] ATTATTTTATTGTATTTTCCTAGTCTATGTAGAGAAAACTACTAATAGAAATCCTGCTTGCTTTTTTTCTTCTTCCCTCCTT
[0247] CCTTTTTTTCCTTTCTTTTCTTTCAGTCTTTCTTTCCAAAAGTATTTCACTCAAATAAAAGGAAGAAAACCAACACAAAA
[0248] GTCATCCTCTAAATTGATAACATCTTCTTGAAGCTCTTTTATGTAGTGACTATGAATCTCCCACCTGCCCCACCTCCTAGC
[0249] TAAATTTAGGGTTTACAAAAGGCTGTACTGAAATGAACAAACACCAATTTGTTAACAAATACCTTAAAGAAAATTAGTT
[0250] TTTTCAAGGCACAATGGAGCCAAGCCCTCTCCAAATTCCCAAAGGCACTAGAAACATAACGACATCAGAATTCCTCTAG
[0251] ATTTATGCTTATAGTAACTATTTTTATGCAAGCAGGTATATAGTAAGTTGAAAATAAGTATTCTCTAAACTATAATTTTT
[0252] TCATTGCAAATTATCATATAGAAGAAAATAGAAGGAAAAACTACAGGCTGGACACAGACGATCAGTCTGGAAATTGCC
[0253] AAATTGGCCATAACTGAGCTGGCTTTCTCACTTTTGCTCCTACTAAAGTGGAGGAAATTCATAAGGTTTCTAAGGCTGGT
[0254] GATCTAACAGTTTGAAAATTTCTAAATGGTGCCGGATTAAAATATTAAGTCCTGAAGCTGCAAATTTTATTTCCAGTAA
[0255] AAAAAAAAAAAAAAACTAAAGTAAAAATAAATGATTAAGACAAAAATTAAGGAAGTGGGTGCCTCTGTGTGGTTAGT
[0256] AATTACTACTCTTCCTAGGTTGTAGCCTGGCCCAGAAAGTCACCATAATACATTTTATACAATGCAACATGTATCCTTCC
[0257] ATTGCTAGACATAGATAAATCACAACAGACAGAAACTGACATGCTTGAAGCTGTACAGAGTTCACTTCCCAGGAACTAT
[0258] GTCACATTATAGCATCCTTTTTATCATAAATATGTATGACTATTATGTATCCATAAAAATTAAAAACACACTTTTTAAGA
[0259] ATTCAAACTAATAACAAACTATCATTCACAAAAAGAACAATTTCCATGGATCATGGTTGAAAGAAGGTAAACACAGAT
[0260] TTCCAAGAAGTATGTTTTTAGGAAATTCTAAATATCATGCTTAGGAAGATAAATAATACATAATAGCATTCTTCCCTGCTT
[0261] AAGAATCCCAAATTTCTTCCTACAAACTTAGTAACAAAAAGATATTTTCTTTCATTACTTTCTCTCTTGTAGTTAGAA
[0262] ATCACAATGTCTACGTTGTCATATACAGGTCATTTCAAAACATTTGACAACATGAAAACCTAATTAATTTAGTCTCAA
[0263] GGAGAATATTAAATATACCACGAATAGGACAGAACAATTATGTTGCTTAAGAGGAACCTTGATGGGATGCAATGGG
[0264] TTTAATTGTAAAGTGTTCTTTCACATTGAAGAGAGAGAATAATTCTTATGAGTACATTGTTACAATTTCGCAAAAAGATT
[0265] TAATCATTGAAGTAAGAGACAGATCAAACATGTAATTTCCTTCTTTCTATGAGTAAAGGTTTTAAAGTACTTTAATCATT
[0266] AACCCTTCATTCTGCCTATTTTTTCATGTTCAACCTATTGCTAAAATGTAATAATGAGGGAAGAGAGGAAGAGAATGAG
[0267] AATAGGAACTGAGTATTCTGGCAAGCAGAGCCTTTTCTCAGAAAACATTACTGTCAAAGTCTGTTCCACTCTGAATGC
[0268] AAAGAGTTGGAAATACTTAGAGGAGCGTTGCAATGCTCTTAGGAGATGACCAAGAAAGGAGAAATGTGATACATCAA
[0269] GGCATGGGATACCCTTCCCAACAATGGCTTCTTTTGGCCCACAGTCTAGTGGCTGTGATAGAGACAAGAGGCAGAGAA
[0270] ACTCTCAGCTGACAGGGACAGGTCACTGCCAAAGCCCCACCTTCAAGCTGAAAAGCCTGAAACCCATGGCCCAAAGTG
[0271] AGAACTTCTATCCCTGTTTGCCCACTCTATCCCAATTGGTTCTTTCTGAATATTGTATTTTTACTAATTGAATGCTGCCTT
[0272] TTCCAAAACTACCTAAGGCCTGCCCCACCCTCCATCCTGTGCCTACAAAGACCCTAGACTCAGCCAGTAGAGAGGAGAA
[0273] GTGGCTGGACATCAAAAAGAGGTGTCTTGACTTAAGAGAAGTAGCTGGACATTAGAGAGAGGCAACTTGACTTCGAAG
[0274] GAGAGAGGCAGAGAGGTGACTTGACTTCAGAGGAAAGCGACCTGCCCTTCTCGTACTCTTTCCAGCTCCTCTGTCTGCT
[0275] GCGAGCCACTTTCAGTGCTCAATAAAATTGTCTGCATTCACCATCCTTCAATTTGTCTGTGTGACTTCATTCTTCTTGGCC
[0276] ACTGGACAAGAATTCAGGATACACCAAGTGTGGGTACCCAAAAAGGCTGTCACACTGGCCCTTTACCCTCATTAGTGGA
[0277] GGACAGCCACCCCACACAACAAGGCAAAGGGCCCACTGAGCTGACAACACACTGCTATCTGCGTACAGCAGAGCTGAG
[0278] AGAGCATAGTAACACACCCTCTGGGGCCTTGAGGTCACAGGCACCTGCACCTGTATGCTACCACAGGGCCTGCATGGA
[0279] GTTTGCTCCTGCCAACACCAAAGCAGCTGGTTCCTGCACTTGCTCACCTGCATGCTCCCTCCCACAAGGGGGTAGGGTGA
[0280] AGCAGGCCTGAGTAAATGGAGTTTGCTTCTGCCAGTGCTGAAGCAGCTGGCTAGCTCCTGTACTCACTCACCTATGTGTT
[0281] CCCTCCCACAAGGGGGTTGAGCACAGTGGGAGGAGTAAATGGGGAATCCCTGTCACCCATCCCTCGAAGCGGTCAAAAA
[0282] AAAAAATATCCTGCATCAGCTGCAAGGTACAAGGCACCTTGAGCTAGC
[0283] The hypermethylated sequence obtained by bisulfite treatment of the inverse complementary sequence is as follows:
[0284] SEQ ID NO:5
[0285]
[0286] The hypomethylated sequences obtained by bisulfite treatment of the reverse complementary sequences are as follows:
[0287] SEQ ID NO:6
[0288] ATAAAATGTTTTTTGTTTATTATTGTATTTATTGTTGAATTTTTAGTTAAATTTTTTTTTTTGTATATGGTGATTTTATT
[0289] TTTTATATAATTTTTATTTTTTATTGTTATTATGAAAAGTATATTTTGTATGAAATTTTTTTAGTAATGTTTTTTTATGAAA
[0290] TAGTTTTTATTTTTTTTGAAGTATAAGTATTAAATTTATAAATGTGAGATTTATAATTATTGATAGATAGGAAGTTTTTTG
[0291] GGTTTATAGATTTTTAATTTATTGTTTTGAATGTTTATTGTTAGTATTAGTGATTTGTTTTTATTTATTGGGGTTATTTTAG
[0292] GTTTTGGTTTTTAGTTGTAGGTTAGTTGTTATGATGTTTTTATTTTGAATTAGTGTTATTTATTAAAATGAGGATTATGGG
[0293] AATTAATAATATATTAGAATTTATTTGAAGTTTATTTTTTTTATAGAATTAAAAGAAAAGTTTTGGAGAGAAAATTTGTA
[0294] TTTGGGGAAAAGAAAAATTTTTATGAAGAAAAGTGAAATTTTTTTGTTTTATTATTTGTGAAGTAATAGGTAAAAGTAA
[0295] ATTGATTATTGATTTTTTTATGTTTGTGTTTTTTAGAAAGGATTTATTTATTTTATATTTAAAATGTATTTAGATAGAAGA
[0296] GGGTATGTAGGTATAATATTTATTTATGATAGTTATAGTTTTTTTTTTTTTGAAGTAATTTAAAATGTGATATTAAAGT
[0297] TTTTGAGTTATTATGAAATGGTTAAATTTTTTGATATTTTTTTTTAAATATATGTAGTTTTTATTTTTTAGGTTGAATAGTTAA
[0298] AATTAGTTGGGAGTTGTTGTTTGTAGAATGTTTGAGATGAGTTTTTAGGGTTTTAGGTGTAGTTGAAGAGGTAGTTATGG
[0299] TGAATATTAGAGAAGGGTAATGTTGAAAAGGGAAAGAAGGTGAGGAAATAAGTTTGGGAGATATATTGTATTAGAGAA
[0300] ATGAAATTTGTTTATTGTAGAGATATAAATAGTTTTATTTATATTTTGAGTAATTGGATAAGAGGTTTTGTAATTTTGGG
[0301] TTTGTTGGTTTTATTTTTGTTTTTTTTTGATTTTAATTGTTTTTGTGTTTATATAGAATTGTTTAGTAAAAATAGTATTATT
[0302] GTTAATTATTTTTAATATATTGTATTTTTTAAGGTTGTTTAGTTATAGTTATTTTTTTTTGTTGTTATTTTTTTAGTGAATT
[0303] ATAGTGTTTTTTTTTTTTTTTTTTTTTTTTGTTTTTGTTTTTTTTTTATGTAGGGAAGATTATGGTTTTGATATTTTTTGGATAATA
[0304] ATGTTGTTTTTTATTAAGTATAAGGAAATAGAATTATATTTGGAAAAAAAGTATAGTTAGGTATTGAATTAATAGAATT
[0305] AAAGTTTTATTTATAATTATTTAATTGAGAAAAGTAATTGTAAATATATAAATTTAATGTTAGGTAATATTTTAGGAGG
[0306] TTAGAAAGAATTTTTAGAAGGGATTTTTTTTTTGGTTTAGTAGTTAATAGGAAAAGTATAATGGTGGGAAATGGAGGGA
[0307] TTTTATTTATTATTAATTATATTGGGTAAATAATTTGTTTGAATTTTATAGTGAAGATTATTAATTTTGGTTTGATGAGTA
[0308] TAAATTTTGTTTTGTTTTATAATTTTTTTTTTTTTTTATAATGTTGTGATGTATTTGGTTTTTTTTTTTTTTTTTTAATATTTT
[0309] GGTTTTGTTTTTAATTTTTTTTGATTGAATTTATAGTAAGTTATTTTTTTTTTTTGAAATAATGGGAGATAATATTTGGAA
[0310] AATATTAAAATAAAAAAGAAAATGAAAATAAGAGTTATTTATATTTTTAGTATATAGAATTGGTATTTTAATTTATTTGT
[0311] TTTTAGTGTTTTAATATTTTTTATGTATAAATTTTTTTAATATATAATATATGAGTGAATTTTTTAAAAAATTGAAGATTGT
[0312] AAATAAAATAAAGGTTTTTTTGTAATATAAAGATAAAATTGAAGTTTTTTTTGATTATTATTTTAATATTAGTTTTTAAA
[0313] TTTAATTTTTAAAGTGTAAGTTTGATATATATTTTTTTTAGTTTATTTTTATATTTTTATATTAAATATATTGAAATATATT
[0314] TTGTGGGTGATTTTGGATATAAATGTATTATGGTAAATGTATTAATTTAAAGTTAGTTTTTTTATTAAATATTTTTTTTTG
[0315] AAGTTTTTGTATTTTTTGTTAGGTTTATTTTTTATATATTATATAGATATTATTGTAATTTTAAATGAAATTTTTTTTTTTAT
[0316] TTTATTTTTTAATGGGTTATTGTTAGGTTGTAGGAAAGAAATTGATTTTGATATGTTGATTTTATATTTAATAATTTTGTT
[0317] GTATTTTTTTATTATTTTTATGGTAGGTTTTTTTTATTTTTTTATGATTTTTAGAAAATTGGATGTTATATATGTTATTTAT
[0318] AGTTTTGAAAATTAGATAATTTTATGAAGTTTTTGAAAATAATAAAATTTTATTATAATGAAATTATGATGTTTTATGAG
[0319] GGTTTTGATAATGGGAATTGGTTTGTTATGTTTTGTTGATTTATAGTTTGGTTAGATATTATTTATATTTGGTGGTAGTTT
[0320] GTATGTTTATAAGATAATTGTATTGATTTTTTTTTTGTGGTTTTTTAATGTTTTAGTATAAAGTTAGAGTTGAAATTATTTT
[0321] TTTAAGGGTTAAAGTTATGTTTATTTTTTTTTTGGTGGATTATAGTGATGAGAATGATGGTAAGTTTTTTAGGGTTAATTT
[0322] GTAGGAGGTTTTGGGAGAGAGAAGTGTTTTAAGTAATTGTATTAAAAGTTAATATAGTGTTTTATTATAATTTTTAAGTA
[0323] GAATAGTTAGAAAAAAAAAAAAAAAGATAAAAAATGAATAATTTTTGTTGAAAGAAAGTAAATTTTTATAGATTAG
[0324] TAATGTTTATTAAAGTTTTATAATAAAATGATTGTTTGTATATATAGATGTTTATGTTGGTTTTAATTTTTTATTTGTATTGT
[0325] TTTAATAAGGTTAATGGATTTTGGATATATTATTTATATATGATGAATTTATAGTAGAATTTATTATTAGTAATGATTA
[0326] ATAATAAATGTTTAAGGTTTATAAATTGGTTATTATTTTATTGTATTTTTTTAGTTTATGTAGAGAAAATTATTAATAGAA
[0327] ATTTTGTTTGTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTAGTTTTTTTTTAAAAGTATTTTATTTAA
[0328] ATAAAAGGAAGAAAATTAAATAAAAGTTATTTTTTAAATTGATAATATTTTTTTGAAGTTTTTTTATGTAGTGATTATG
[0329] AATTTTTTATTTGTTTTATTTTTTTAGTTAAATTTAGGGTTTATAAAAAGGTTGTATTGAAATGAATAAATATTAATTTGTTA
[0330] ATAAATATTTTAAAAGAAAATTAGTTTTTTTAAGGTATAATGGAGTTAAGTTTTTTTAAATTTTTAAAAGGTATTAGAAAT
[0331] ATAATGATATTAATTTTTTTAGATTTATGTTTATAGTAATTATTTTTATGTAAGTAGGTATATAGTAAGTTGAAAATA
[0332] AGTATTTTTTAAATTATAATTTTTTTATTGTAAATTATTATATAGAAAAAAATAGAAGGAAAAAATTATAGGTTGGATATA
[0333] GATGATTAGTTTGGAAATTGTTAAATTGGTTATAATTGAGTTGGTTTTTTTATTTTTTGTTTTTATTAAAGTGGAGGAAATT
[0334]
[0335] Those skilled in the art will understand that the deletion of one or more nucleotides, the addition of one or more nucleotides, or the substitution of one or more nucleotides based on the nucleotide sequence shown in SEQ ID NO:1-6, but having 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity with the nucleotide sequence shown in SEQ ID NO:1-6, can also serve as a methylation marker for this application and should be understood to be covered within the scope of protection of this application.
[0336] "Cerebral stroke," also known as "stroke" or "cerebrovascular accident" (CVA), refers to an acute cerebrovascular disease. It is a group of diseases caused by the sudden rupture or blockage of blood vessels in the brain, preventing blood flow and resulting in brain tissue damage. These include ischemic and hemorrhagic strokes. The stroke covered in this application can encompass various types of stroke known in the art, such as large artery atherosclerotic stroke, cardioembolic stroke, small artery occlusive stroke, ischemic stroke, and cerebral infarction.
[0337] In one specific implementation, the stroke recurrence refers to a stroke recurrence within three months.
[0338] This application also provides the use of the above-mentioned methylation markers in the preparation of kits for predicting stroke recurrence.
[0339] This application also provides the use of the above-mentioned methylation markers in the preparation of medicaments for preventing recurrence of stroke.
[0340] In one specific implementation, the drug is a targeted drug designed based on the methylation marker.
[0341] This application also provides a composition for predicting stroke recurrence, the composition comprising: a nucleic acid for detecting the methylation state of the above-mentioned methylation marker, wherein the methylation state of the methylation marker is characterized by the methylation of the target sequence of the methylation marker.
[0342] This application also provides a kit containing reagents for detecting the above-mentioned methylation markers.
[0343] This application makes no restrictions on the reagents described, which can be selected as needed, such as primers, probes, etc.
[0344] The kit of this application can be used to detect samples selected from cell lines, histological sections, tissue biopsies / paraffin-embedded tissues, body fluids, feces, colonic effluent, urine, plasma, serum, whole blood, isolated blood cells, cells isolated from blood, or combinations thereof. Preferred samples are plasma and tissue.
[0345] This application also provides a nucleic acid for targeting the aforementioned methylation markers.
[0346] This application screened a large number of stroke patient samples to obtain the biomarkers for this application. The level of methylation of these biomarkers is associated with the risk of stroke recurrence in stroke patients. The results of the methylation biomarkers in this application showed that the AUC was above 0.6, indicating that the biomarkers screened in this application can accurately predict the risk of stroke recurrence in stroke patients, thereby providing guidance for precision treatment and health management.
[0347] The biomarkers screened in this application can not only predict the risk of relapse, but also provide accurate indications of the possible time point of relapse through systematic analysis of an individual's past basic information, epidemiological information, and test results, which has important clinical application value.
[0348] Example
[0349] The following description provides exemplary embodiments of this application, including various details to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this application. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0350] Unless otherwise specified, the experimental methods used in the following examples are conventional methods.
[0351] Unless otherwise specified, all materials and reagents used in the following examples are commercially available.
[0352] Example 1: Whole-genome methylation detection
[0353] 1. DNA extraction and quality control
[0354] 1.1 DNA Extraction
[0355] Sample lysis
[0356] Add 20 μl Proteinase K and 200 μl Leukocyte Lysis Buffer I to the sample storage tube using an electric continuous dispensing device, mix thoroughly, and incubate at 56°C for 30 min.
[0357] During sample lysis, the following reagents, as shown in Table 1, were added to the 96DW deep-well plate using a liquid workstation:
[0358] Table 1
[0359] Location Reagents & Dosage Sample Plate Buffer ML: 200μl Wash Plate I Buffer MWI: 750μl Wash Plate II Buffer MWI: 750μl Wash Plate III Buffer MWII: 750 μl Wash Plate IV Buffer MWIII: 750 μl Elution Plate Buffer EB: 150μl
[0360] Then, it is verified in plasma samples, and the experimental detection method is as follows:
[0361] Machine extraction
[0362] Add the lysis products to the Sample Plate in sequence; (Note: To prevent the lysis products from being added to the wrong location, the area after adding the sample must be immediately sealed with a silicone cap before adding the next sample.)
[0363] Turn on the AE2130 instrument, place the deep well plate into the instrument in sequence, and then run the program '2517-lysis';
[0364] After about 20 minutes, the instrument was paused. The 'Sample Plate' was removed, and 310 μl of a well-mixed mixture of magnetic beads and binding fluid was added to the deep well plate using an electric continuous dispensing device.
[0365] Reinsert the 'Sample Plate' into the AE2130 instrument and continue running the program;
[0366] After about 25 minutes, the program will finish running. Remove the deep well plate and magnetic sleeve from the instrument, transfer the elution product from the 'Sample Plate' to a 1.5 mL EP tube, freeze it, and take a picture for record.
[0367] 1.2 DNA quality inspection
[0368] 1.2.1 Concentration and purity determination
[0369] On a NanoDrop 2000, 2 μl of the extracted product was taken for concentration and purity determination.
[0370] 1.2.2 Qubit Concentration Determination
[0371] Transfer 2 μl of the extracted product to a PCR plate using a variable-gap pipette, dilute 10-fold, mix well, and then determine the concentration in a Qubit.
[0372] Completeness test
[0373] Take 6 μl of the diluted product from 1.2.2 and perform agarose gel electrophoresis to detect genomic DNA integrity.
[0374] Note: Downstream warehousing will commence immediately after quality inspection.
[0375] 2. Construction of the WGBS Library
[0376] 2.1 DNA fragmentation and screening
[0377] After thawing, the DNA samples (including positive control samples) are shaken and centrifuged, arranged in their original order, and can be temporarily stored at 4°C.
[0378] Sample homogenization: Take a low-adsorption PCR plate, add samples in a fixed order, and then add water to bring the volume to 130 μl;
[0379] Mechanical fragmentation procedure: Add 130 μl of the homogenized sample to the fragmentation tube, paying attention to the placement and order of the fragmentation tubes, and perform fragmentation according to the fragmentation procedure; after fragmentation, transfer 129 μl to a new purification plate, and add 1 μl of fragmented λDNA at a concentration of 0.05 ng / μl to each sample.
[0380] Interruption procedure: peak power 450, duty factor 30, cycle / burst 194, interrupt time 120s
[0381] Fragment screening: Balance the magnetic beads 30 minutes in advance and perform fragment screening on the automated workstation. Note that pipette tips and reagents should be prepared and the operating procedure verified 10 minutes in advance. Prepare the end-repair mix (see Table 2 for specific components) in advance and add it to a new low-adsorption PCR plate. The automated system will add the elution product to the mix and mix thoroughly (if paused, the elution product can be directly recovered and temporarily stored at -20℃).
[0382] Random sampling: The remaining products can be temporarily stored, and 1 to 3 samples from different locations will be randomly selected and diluted 10 times for concentration and peak quality inspection.
[0383] 2.2 Library Construction
[0384] 2.2.1 Terminal repair response + A-tail
[0385] The end-repair reaction solution can be prepared 10 minutes in advance according to the system shown in Table 2. After shaking and mixing, centrifuge at 400g for 5s, dispense into 8-tube PCR tubes, and then use a pipette to add the reaction solution to a new low-adsorption PCR plate on an ice box for later use.
[0386] Add 11 μl of the screened product to a low-adsorption PCR plate containing the end-repair mix, seal, shake, centrifuge, and run on the instrument.
[0387] Table 2
[0388]
[0389]
[0390] Reaction program: Hot cap 66℃, 25℃ 30 min, 56℃ 20 min, 4℃ forever
[0391] 2.2.2 Add connector
[0392] The reaction solution can be prepared 5 minutes in advance according to the system shown in Table 3. After vortexing and mixing, centrifuge at 400g for 5 seconds, and aliquot into 8-tube PCR tubes for later use on an ice box. Use a multi-channel pipette to add the reaction solution to the PCR plate containing the end-repair products, seal the plate, vortex, centrifuge, and run the reaction.
[0393] Table 3
[0394] reagents Single dosage (μl) enzyme 2 1 Buffer 2 1.5 Buffer U 0.5 Adapter (2μm) 2
[0395] Reaction procedure: Close the hot lid, 22°C for 40 minutes, 4°C forever.
[0396] 2.2.3 Purification of adapter products
[0397] Prepare columns and collection tubes in advance. Label the columns with the corresponding numbers. Alternatively, prepare clean 1.5mL EP tubes in advance and label them with the corresponding numbers.
[0398] Add 140 μl of DNA Binding Buffer to the sample, mix 5 times, and then transfer the entire amount to the corresponding column.
[0399] Centrifuge at 13000 rcf for 30 seconds (to ensure complete separation from the liquid);
[0400] Add 200 μl of DNA Wash Buffer to the column, centrifuge at 13000 rcf for 30 seconds (to ensure complete separation of the liquid);
[0401] Repeat the previous step;
[0402] After centrifugation, place the column on a previously prepared 1.5 mL EP tube and add 22 μl of EB Buffer;
[0403] Centrifuge at 13000 rcf for 30 seconds (to ensure complete separation from the liquid);
[0404] Take 20 μl of the liquid and proceed to the next step, Mix, for reaction.
[0405] 2.2.4 Inactivation reaction
[0406] The reaction solution can be prepared 10 minutes in advance according to the system shown in Table 4. After vortexing and mixing, centrifuge at 400g for 5s, and aliquot into 8-tube PCR tubes for later use on an ice box. Use a multipipe to add the reaction solution to a new low-adsorption PCR plate, then add the purified product, seal the plate, vortex, centrifuge, and run the reaction.
[0407] Table 4
[0408] reagents Single dosage (ul) Enzyme 4 0.2 Buffer 4.1 0.8 Buffer 4.2 1 Buffer U 2.5
[0409] Reaction procedure: 84℃ hot, 74℃ for 10 minutes, 4℃ forever
[0410] 2.3 Transformation
[0411] After inactivation, the sample was mixed with 130 μl of LC reagent using a 200 μl pipette, blow-mixed 10 times, sealed, and then subjected to the conversion reaction. After the reaction was completed, the sample could be stored at 4°C for 20 hours.
[0412] Reaction procedure: Hot cap 105℃. 98℃ for 8 minutes, 54℃ for 60 minutes, 4℃ forever.
[0413] After the reaction is complete, remove the sealing film and proceed with the following steps.
[0414] Prepare columns and collection tubes in advance. Write the corresponding numbers on the columns. Alternatively, prepare clean 1.5ml EP tubes in advance and write the corresponding numbers on them.
[0415] Add 600 μl of M-Binding Buffer to the column, then transfer all the sample to the column, cap it, and invert it 10 times.
[0416] Centrifuge at 13000 rcf for 30 seconds (to ensure complete separation from the liquid);
[0417] Add 100 μl of M-Wash Buffer to the column, centrifuge at 13000 rcf for 30 s (to ensure complete separation of liquid);
[0418] Add 200 μl L-Desulphonation Buffer, let stand at room temperature for 15 min, then centrifuge at 13000 rcf for 30 s (to ensure complete separation of the liquid);
[0419] Add 200 μl of M-Wash Buffer to the column, centrifuge at 13000 rcf for 30 s (to ensure complete separation of liquid);
[0420] Repeat the previous step;
[0421] After centrifugation, place the column on the previously prepared 1.5mL EP tube, add 10μl EB Buffer, 10000rcf, and centrifuge for 30s (to ensure complete separation of the liquid).
[0422] Take 9 μl of the liquid and proceed to the next step, Mix, for reaction.
[0423] 2.4 Library Amplification
[0424] The PCR reaction solution can be prepared 5 minutes in advance according to the system shown in Table 5, dispensed into 8-tube PCR tubes, and added to a new low-adsorption PCR plate using a pipette. Add index to each tube, record the results, and finally add the purified product. After sealing the membrane, shake, centrifuge, and run the PCR reaction.
[0425] Table 5
[0426] reagents Single dosage (ul) PCR enzyme 0.5 2X PCR Buffer 12.5 P5N primer (10 pmol / ul) 1.5 P7 primer (10 pmol / ul) 1.5 (single addition)
[0427] Reaction program: Hot cap 105℃. 94℃ 2min; (98℃ 10s, 52℃ 30s, 68℃ 15s) 12 cycles; 72℃ 1min; 4℃ forever
[0428] 2.5 Library Purification
[0429] After the reaction is complete, remove the sealing film and proceed with purification on the automated workstation. Note that you should prepare the pipette tips and reagents 10 minutes in advance, take photos for documentation, and verify the operating procedure.
[0430] During automated operation, prepare 1.5ml low-adsorption EP tubes for shipment, and complete label printing, affixing, and shipment information preparation. After purification is complete, use an adjustable-gap pipette to transfer the purified product into the labeled tube and record the results.
[0431] Subsequently, three samples were randomly selected for concentration and peak plot testing. Any abnormalities were reported to the person in charge of database construction.
[0432] 3. Document Quality Inspection
[0433] The Bioptic Qsep100 fully automated nucleic acid and protein analysis system or the Agilent 2100 bioanalyzer can be used to detect the fragment distribution of a library.
[0434] Effective concentration of qPCR quality control library.
[0435] 4. Lab session
[0436] The library sequencing platform was an Illumina Nova 6000 sequencer, and the sequencing mode was PE150.
[0437] 5. Data technical parameter requirements:
[0438] Sequencing data ≥90G / sample
[0439] BS conversion rate ≥99%,
[0440] Data redundancy <23% (number of duplicate reads in the mapping divided by the number of mapping reads).
[0441] Clean reads comprised >85% of the original data.
[0442] After deduplication: Genome coverage of 1× or higher >90%, Genome coverage of 5× or higher >85%
[0443] The coverage of 5×mCG (excluding sex chromosomes) is >70%.
[0444] Example 2: Biomarker Screening Process
[0445] This batch contained 930 samples (from Beijing Tiantan Hospital), including 662 patients with no recurrence of stroke and 268 patients with recurrent stroke. White blood cells (WBCs) were extracted for WGBS sequencing. The screening method for significantly relevant large fragments associated with recurrence was as follows:
[0446] Quality control: First, use quality control software such as FASTP to check the quality of 930 raw WGBS sequencing data, and filter, truncate or remove low-quality reads to obtain the corresponding clean data;
[0447] Alignment and deduplication: The quality-controlled clean data was aligned to the reference genome (hg38) using Bismark Bowtie2 alignment software; the BAM files from the initial alignment were deduplicated using deduplicate_bismark.
[0448] Extracting methylation site information: Using Bismark_methylation_extractor to extract the corresponding methylation site information, the final methylation CG file (including all individual CG site information files) is obtained;
[0449] The sliding window method calculates the methylation level of the target region: First, a sliding window is used on each chromosome of the reference genome (hg38) to calculate the overall methylation level of CpG sites within each window. For each window, the number of CG sites is counted. Since the methylation depth of cytosine at each CG site and the total depth of the site are known, the methylation level of the entire window can be calculated, which is the ratio of the sum of the methylation depths of cytosine at all CG sites to the sum of the total depths of all CG sites. Each window will obtain a corresponding methylation level through the above calculation method. The methylation depth of cytosine at each CG site is the number of reads showing methylated cytosine at that site in the sequencing results, i.e., the number of reads showing C (cytosine) at that site. The total depth of the site is the total number of all sequencing reads covering that site, i.e., the total number of reads showing C or T (thymine) at that site. The depth of methylated cytosine and the total depth of the site can be directly provided after analysis using sequencing software;
[0450] Biomarker filtering: The obtained biomarker files are filtered to select biomarker intervals with sequencing depth greater than or equal to 3, and biomarkers with a CpG count greater than 60 in 80% of the samples are retained.
[0451] Screening of relapse-related biomarkers: The samples were divided into five groups based on a fixed ratio of relapse and non-relapse within three months. Four groups were used as the training set, and one group was reserved as the test set, resulting in a total of five data groups. In each group, the biomarkers (approximately 5 million, calculated using the sliding window method mentioned above) from the test set were combined with clinical information (patient relapse time and relapse status) for Cox regression analysis. Specifically, for the Cox regression analysis, we used the log-likelihood test to compare the difference between relapse and non-relapse within three months (measured by p-value), screening out approximately 300,000 candidate biomarkers with p < 0.05. Then, we further filtered out biomarkers with small coefficients using lasso regression on the candidate DMV, obtaining approximately 100 biomarkers. Finally, biomarkers located in the promoter region were screened, and a generalized linear model was constructed using the test set to select the biomarker with the best AUC as the final methylation biomarker related to relapse, namely the biomarker of this application, located at positions 60727176 to 60733176 on chromosome 13 of the human reference genome hg38.
[0452] Example 3: Biomarker and Relapse-Related Classification AUC Display
[0453] In this embodiment, 930 samples were divided into 5-fold subgroups according to a fixed ratio of relapse to no relapse within three months. Four groups were selected as the training set to construct a generalized linear model, and the additional groups were used as the test set for validation (e.g., ...). Figure 1 (As shown). Ultimately, the AUC levels of the markers in this application on the five test sets were 0.578, 0.667, 0.636, 0.604, and 0.681, respectively, with an average AUC of 0.633 (as shown). Figure 2 (As shown).
[0454] Example 4: Biomarker subset and relapse-related AUC display
[0455] In this embodiment, methylation sites in the region from position 60727176 to 60733176 of chromosome 13 on the human reference genome hg38 were classified according to a certain distance (two CpG sites less than 100 bp), resulting in the following subsets (as shown in Table 6). These subsets were then segmented and modeled according to the method described in Example 3. The final AUC levels of the new classification combination markers in the five test sets were 0.572, 0.662, 0.566, 0.534, and 0.603, respectively, with an average AUC of 0.587 (as shown in Table 6). Figure 3 (As shown).
[0456] Table 6
[0457]
[0458]
[0459] The above description is merely a preferred embodiment of this application and is not intended to limit the application in any other way. Any person skilled in the art may make changes or modifications to the disclosed technical content to create equivalent embodiments. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of this application without departing from the scope of the technical solution of this application shall still fall within the protection scope of this application.
Claims
1. Use of a methylation biomarker in the preparation of a kit for predicting stroke recurrence, wherein the target sequence of the methylation biomarker is the sequence shown in SEQ ID NO:
1.
2. The use according to claim 1, wherein the stroke is selected from one or more of the following: large artery atherosclerotic stroke, cardioembolic stroke, small artery occlusive stroke, ischemic stroke, and cerebral infarction.
3. The use according to claim 1, wherein the stroke recurrence is a stroke recurrence within three months.
4. A composition for predicting stroke recurrence, said composition comprising: Nucleic acids used to detect the methylation status of methylation markers The methylation state of the methylation marker is characterized by the methylation of the target sequence of the methylation marker. The target sequence of the methylation marker is as shown in SEQ ID NO:
1.
5. The composition according to claim 4, wherein the stroke is selected from one or more of the following: large artery atherosclerotic stroke, cardioembolic stroke, small artery occlusive stroke, ischemic stroke, and cerebral infarction.
6. The composition of claim 4, wherein the stroke recurrence is a stroke recurrence within three months.
7. A kit comprising reagents for detecting methylation markers, wherein the target sequence of the methylation marker is a sequence as shown in SEQ ID NO:
1.
8. The kit according to claim 7, wherein, The kit is used to detect samples including: cell lines, histological sections, tissue biopsies, feces, colonic effusion, urine, plasma, serum, whole blood, isolated blood cells, or combinations thereof.