Screening method for cell-free DNA marker, DNA marker and use thereof
By screening and sequencing CpG sites in the human genome, potential free DNA methylation markers were screened out, which solved the problem of insufficient sensitivity and accuracy of blood marker detection in early AD diagnosis in the prior art, and achieved non-invasive and repetitive and efficient early diagnosis of AD.
Patent Information
- Application Number
- PCT/CN2024/124681
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-13
- Filing Date
- 2024-10-14
- Publication Date
- 2025-06-19
AI Technical Summary
The prior art lacks effective blood markers in the early diagnosis of Alzheimer's disease (AD), especially detection methods based on plasma protein markers have problems of insufficient sensitivity and accuracy, which is difficult to meet the needs of large-scale early detection.
By screening CpG sites in the human genome, whole genome sequencing combined with CpG segmentation method, the free DNA methylation markers with potential were screened out, and the samples were nucleotide sequenced by DNA methylation assay to determine the methylation level of the markers.
It realizes early diagnosis of AD without tissue samples, overcomes the problem of brain tissue acquisition, improves the sensitivity and accuracy of detection, and is non-invasive and repetitive through peripheral blood collection, suitable for large-scale promotion.
Smart Images

Figure CN2024124681_19062025_PF_FP_ABST
Abstract
Description
Screening method for free DNA marker, DNA marker and its application Technical Field
[0001] The present invention relates to the technical field of gene detection, and in particular to a method for screening free DNA markers, DNA markers and applications thereof. Background Art
[0002] Dementia is an important cause of brain health problems for the elderly in my country, and it directly endangers the health and well-being of the people. Alzheimer's disease (AD) is the main cause of Alzheimer's disease (more than 65%). As of the end of 2021, the number of elderly people aged 60 and above in my country reached 267 million, accounting for 18.9% of the total population. It is estimated that around 2035, the number of elderly people aged 60 and above will exceed 400 million, accounting for more than 30% of the total population, entering a stage of severe aging. There are currently about 47 million AD preclinical patients, 39 million MCI patients, and 15 million dementia patients in my country. Cognitive impairment cannot be reversed after the patient develops it, so "early detection, early diagnosis, and early treatment" are the key to the prevention and treatment of AD.
[0003] Extracellular amyloid (β-amyloid) plaques and intracellular neurofibrillary tau tangles are the two main pathological hallmarks of AD and serve as the basis for distinguishing AD dementia from non-AD dementia. Currently, the most established diagnostic methods for AD rely primarily on direct detection of Aβ plaques and tau tangles in the living brain using positron emission tomography (PET), or indirectly through measurement of AD-related pathological proteins such as Aβ42 and p-Tau in the cerebrospinal fluid (CSF). However, PET imaging is radioactive and expensive, and not all hospitals have imaging equipment. Furthermore, CSF extraction is tedious and painful. In recent years, blood-based early diagnosis techniques have gained traction, offering significant advantages due to their minimal invasiveness and the ability to be collected multiple times. Blood-based disease biomarkers hold significant clinical value, but many diseases currently lack effective diagnostic and monitoring markers, particularly for early diagnosis. The identification of markers includes aspects such as biological principle exploration, biochemical experiments, and bioinformatics computational methods. There is a great demand for highly sensitive and flexible disease marker identification methods. In the field of AD, the current focus of blood testing technology is almost entirely on plasma protein markers, and the detection of these plasma protein markers mostly requires the use of ultra-sensitive biomarker detection systems (such as SIMOA) or mass spectrometry platforms. The consumables and prices are relatively high, and the detection sensitivity and accuracy are not ideal. In summary, these defects have led to the inability of plasma protein biomarkers to meet the needs of large-scale early detection of AD in Chinese communities.
[0004] Cell-free DNA (CBD) in peripheral blood is a naturally occurring DNA fragment, mostly released into the peripheral blood after cell death. In healthy individuals, CBD primarily originates from blood cells and the liver. However, in patients with disease, diseased tissue releases CBD, triggering immune responses and causing changes in CBD. Therefore, CBD analysis can be used for disease diagnosis and detection. Furthermore, the metabolic half-life of CBD is approximately six hours, making it possible to monitor the human body in real time. Currently, CBD analysis focuses on areas such as cancer, pregnancy, and infectious diseases. DNA methylation is an important epigenetic modification of DNA. A common type of DNA methylation is 5-methylcytosine (5mC), which adds a methyl group to cytosine (C). In humans, 5mC occurs primarily on cytosine within CpG dinucleotides (cytosine and guanine linked by phosphates). It can influence genomic stability and regulate gene expression, exhibiting strong cell and tissue specificity. There are tissues and organs with different functions in the human body, such as the liver, lungs, kidneys, and brain. Their genomes are almost exactly the same, but there are significant differences in DNA methylation modifications. When a tissue in the human body is in a pathological state, causing its cells to die and release DNA into the peripheral blood, the pathological state can be detected and diagnosed by identifying its tissue-specific DNA methylation or sites. In addition, the immune response produced by the body under pathological conditions can also lead to abnormal changes in free DNA methylation. Free DNA methylation detection has been widely used in prenatal diagnosis and cancer diagnosis and has achieved great success. However, the identification of most disease markers of free DNA methylation requires collecting methylation data from disease-related tissues (or collecting relevant tissues and obtaining them through experiments by oneself), identifying disease markers by comparing them with peripheral blood or normal tissues, and then verifying them in peripheral blood. The different ways of obtaining methylation data have a significant impact on subsequent marker identification.
[0005] For example, DNA methylation chips (common ones include Illumina HumanMethylation450BeadChip, Illumina Infinium Methylation EPIC BeadChip, etc.) cover a fixed number of CpG sites with specific sequences. The capture depth of each CpG site in the experimental data is high, and the distance between adjacent CpG sites is relatively far, so each CpG site is generally analyzed separately. Methods based on specific recognition enzymes for enrichment, such as MEDIP-seq, have low resolution and are rarely used in disease diagnosis. The current mainstream is sequencing-based methods, such as WGBS, RRBS, EM-seq, TAPS and other technologies, which can cover almost all CpG sites. However, due to the limitation of sequencing depth, the quantitative deviation of the methylation level of a single CpG site is large. At the same time, considering the correlation between the methylation status of adjacent CpGs, multiple adjacent CpG sites will be used as a marker, and how to segment the CpG sites is very important. There are many mainstream approaches, such as using a single element or part of its sequence as a segment based on known regulatory elements such as CpG islands, CpG island shores, and gene promoters, but these methods still have shortcomings. In addition, due to the low concentration of free DNA (only about 7 ng of free DNA per milliliter of plasma), and the great damage of sulfites in traditional WGBS experiments to DNA, free DNA methylation experiments are very difficult; disease-related CpG sites only occupy a very small part of the genome, resulting in high detection costs. In summary, the development of early disease diagnosis technology based on free DNA methylation requires optimization in both experimental and computational methods.
[0006] Summary of the Invention
[0007] In view of this, the main purpose of the present invention is to provide a method for screening free DNA markers, DNA markers and applications thereof, in order to at least partially solve the above technical problems.
[0008] In order to achieve the above object, as a first aspect of the present invention, a method for screening free DNA markers is provided, comprising the following steps:
[0009] Select any human gene sequence and use its first CpG site as a candidate segment. For any candidate segment, examine its next adjacent CpG site. If the interval between the CpG site and the last CpG site in the candidate segment is less than or equal to a first preset threshold Di, merge the CpG site into the current candidate segment and continue to examine subsequent CpG sites until the interval condition is no longer met.
[0010] Inspect the current candidate segment, and if the number of CpG sites contained in it reaches the second preset threshold Num, mark the candidate segment as a formal segment, otherwise discard the candidate segment;
[0011] After the candidate segment is examined, the next adjacent CpG site is used as the starting point for a new candidate segment, and the above process is repeated until the preset end condition is reached;
[0012] Collect verified positive experimental samples and control group samples;
[0013] The positive experimental samples and the control group samples were sequenced using a DNA methylation assay:
[0014] For all samples, for all the formal segments determined above, the methylation quantitative values of all CpG sites contained therein are averaged as the methylation level of the formal segment;
[0015] For each formal segment, if the number or proportion of samples in the control group whose methylation levels are less than a preset threshold a is not less than x, and the number or proportion of samples in the positive experimental samples whose methylation levels are greater than another preset threshold b is greater than a preset threshold y, then the formal segment is regarded as a candidate marker; or
[0016] For each formal segment, if the number or proportion of samples in the control group whose methylation levels are greater than the preset threshold aa is not less than the preset threshold xx, and the number or proportion of samples in the positive experimental samples whose methylation levels are less than another preset threshold bb is greater than the preset threshold yy, then the formal segment is also regarded as a candidate marker.
[0017] Wherein, the positive experimental sample is a blood or body fluid sample whose cerebral cortex Aβ plaque pathology is confirmed to be positive by AβPET; preferably, it is a peripheral blood sample;
[0018] The control group samples are blood or body fluid samples whose cerebral cortex Aβ plaque pathology is negative as determined by AβPET;
[0019] Preferably, the blood or body fluid sample is obtained by a non-invasive method; further preferably, the sample is a peripheral blood sample obtained by a non-invasive method within 6 hours of ex vivo.
[0020] The DNA methylation determination method is selected from WGBS technology, RRBS technology, EM-seq sequencing method, TAPS technology, digital PCR, real-time fluorescence quantitative PCR and methylated DNA specific recognition enzyme binding technology.
[0021] wherein the first preset threshold Di is 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90 or 100 (bases);
[0022] The second preset threshold Num is 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 (CpG sites).
[0023] The method for calculating the average value is selected from arithmetic mean, median average, weighted average, etc.
[0024] Wherein, the preset threshold a is 10%, the preset threshold x is 90%, the preset threshold b is 20%, the preset threshold y is 25%, the preset threshold aa is 90%, the preset threshold xx is 90%, the preset threshold bb is 80%, and the preset threshold yy is 25%; or
[0025] The preset threshold a is 5%, the preset threshold x is N-2, the preset threshold b is 20%, the preset threshold y is 33%, the preset threshold aa is 95%, the preset threshold xx is N-1, the preset threshold bb is 80%, and the preset threshold yy is 33%; wherein N is the number of samples in the control group; or
[0026] The preset threshold a is 5%, the preset threshold x is N-1, the preset threshold b is 5%, the preset threshold y is 25%, the preset threshold aa is 95%, the preset threshold xx is N-1, the preset threshold bb is 95%, and the preset threshold yy is 25%; where N is the number of samples in the control group.
[0027] The positions of all or part of the CpG sites on all chromosomes in the human genome are screened; the human genome data are from GRCh38, GRCh37, GRCh36, T2TCHM13v2.0 / hs1 or Han1.
[0028] As a second aspect of the present invention, a candidate marker obtained by screening according to the above screening method is also provided;
[0029] Preferably, the candidate marker comprises the nucleotide sequence as described in SEQ ID No.1 to SEQ ID No.513.
[0030] As a third aspect of the present invention, there is also provided a primer, a primer amplification chain or a sample composition obtained by PCR amplification, replication, conversion and / or transformation based on the candidate markers as described above.
[0031] As a fourth aspect of the present invention, a methylation machine sequencing system is also provided, wherein the methylation machine sequencing system comprises a nucleic acid library of candidate markers as described above.
[0032] Based on the above technical solutions, it can be seen that the marker screening method, DNA marker and application thereof of the present invention have at least one of the following beneficial effects compared with the prior art:
[0033] 1. The present invention does not require tissue samples, overcoming the technical difficulty of obtaining brain tissue, especially brain tissue from early-stage AD patients, which is almost impossible to obtain;
[0034] 2. In the screening method of the present invention, whole genome sequencing is used, and the method of segmenting the genome in combination with CpG has a high degree of freedom. The effective gene segment can be as short as 10 bp (i.e., 10 bases) in length and contain only 3 CpG sites.
[0035] 3. The marker combination screened by the present invention has high noise tolerance for detection, and the control group has a clean signal;
[0036] 4. The detection method using the marker combination screened by the present invention has high sensitivity. It does not require a high accuracy rate for a single marker, but rather selects all potential markers.
[0037] 5. The detection method of the present invention only collects peripheral blood, is non-invasive to the human body, and can be collected repeatedly;
[0038] 6. Research in the field of cancer diagnosis has shown that disease-related cell-free DNA markers appear in the blood earlier than proteins, so cell-free DNA markers have a better application prospect in early diagnosis. In addition, most cell-free DNA is double-stranded DNA, which is highly stable and reproducible.
[0039] 7. The metabolic half-life of cell-free DNA is short, only 6 hours, so diseases can be monitored in real time. The cell-free DNA methylation detection technology is mature and can detect multiple sites at a time, making it suitable for large-scale promotion and use.
[0040] 8. The marker identification parameters of the present invention are not based on statistical tests, and statistically different sites are often not effective as markers. The present invention does not pursue a single marker to achieve good results, but rather identifies a certain amount of potential markers with strong anti-interference ability. The combined use of all or part of these markers (such as using average values, combining machine learning technology) can achieve better diagnostic results. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 is a dot-line graph comparing the free DNA methylation rate (methylation level) of a Seq. 21 marker in an early AD patient (asterisk) and a healthy elderly control (×) (each dot in the graph represents a CpG site);
[0042] Figures 2A and 2B are the arithmetic mean distributions of the methylation levels of all markers in healthy elderly people (left) and AD patients (right) based on Aβ PET imaging in two independent data sets, respectively (each point in the figure represents a patient);
[0043] FIG3 is a performance comparison diagram (ROC curve diagram) of the markers of the present invention and other AD markers (Aβ42 / Aβ40, p-Tau181, NfL, GFAP) in the classification of healthy controls and AD patients;
[0044] FIG4 is a graph comparing the performance of the free DNA marker of the present invention and other AD markers (Aβ42 / Aβ40, p-Tau181, NfL, GFAP) in predicting the intensity of Aβ accumulation signals in the patient's brain. DETAILED DESCRIPTION
[0045] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings.
[0046] Unless otherwise defined herein, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the invention belongs. Any methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present invention, however, preferred methods and materials are described.
[0047] All patents and publications mentioned herein, including all sequences disclosed within such patents and publications, are expressly incorporated by reference.
[0048] Numerical ranges include the numbers defining the range. Unless otherwise indicated, nucleic acids and amino acids are written left to right in 5' to 3' orientation and amino to carboxyl orientation, respectively.
[0049] The headings provided herein are not limitations of the various aspects or embodiments of the invention. Accordingly, the terms defined below are more fully defined in conjunction with the specification as a whole.
[0050] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the invention belongs. For clarity and ease of reference, certain terms are defined below.
[0051] As used herein, the term "sample" refers to a material or mixture of materials, typically, although not necessarily, in liquid form, which contains one or more analytes of interest.
[0052] As used herein, the term "nucleic acid sample" refers to a sample comprising nucleic acids. As used herein, nucleic acid samples can be complex because they comprise a variety of different molecules comprising sequences. Mammalian (e.g., mouse or human) genomic DNA is typical of complex samples. Complex samples can have more than 104, 105, 106, or 107 different nucleic acid molecules. DNA targets can be derived from any source, such as genomic DNA or artificial DNA constructs.
[0053] The term "nucleotide" is intended to include structures that contain not only the known purine and pyrimidine bases, but also other modified heterocyclic groups. These modifications include methylated purines or pyrimidines, acylated purines or pyrimidines, alkylated ribose or other heterocyclic compounds. In addition, the term "nucleotide" includes structures that contain haptens or fluorescent labels, and may contain not only conventional ribose and deoxyribose sugars but also other sugars. Modified nucleosides or nucleotides also include modifications of the sugar moiety, for example, where one or more hydroxyl groups are replaced by halogen atoms or aliphatic groups, or are functionalized as ethers, amines, etc.
[0054] The terms "nucleic acid" and "polynucleotide" are used interchangeably herein to describe polymers of any length composed of nucleotides, such as deoxyribonucleotides or ribonucleotides, e.g., greater than about 2 bases, greater than about 10 bases, greater than about 100 bases, greater than about 500 bases, greater than about 1000 bases, up to about 10,000 bases or more, and can be produced enzymatically or synthetically. Naturally occurring nucleotides include guanine, cytosine, adenine, and thymine (G, C, A, and T, respectively). DNA and RNA have deoxyribonucleotide and ribonucleotide sugar backbones, respectively.
[0055] As used herein, term " oligonucleotide " refers to about 2 to 200 nucleotides, up to a single-stranded nucleotide polymer of 500 nucleotide lengths. Oligonucleotide can be synthetic or can be made through an enzymatic process, and in some embodiments is 30 to 150 nucleotide lengths. Oligonucleotide can comprise ribonucleotide monomers (i.e., can be oligoribonucleotides) and / or deoxyribonucleotide monomers. Oligonucleotide can be, for example, 10 to 20, 21 to 30, 31 to 40, 41 to 50, 51 to 60, 61 to 70, 71 to 80, 80 to 100, 100 to 150 or 150 to 200 nucleotide lengths.
[0056] The term "CpG" as used herein refers to 5'-C-phosphate-G-3', ie, two nucleotides consisting of cytosine (C) and guanine (G) immediately adjacent to each other.
[0057] The term “human genome GRCh38” used herein refers to the 38th version of the human genome reference sequence released by the Genome Reference Consortium (GRC) in 2013.
[0058] As used herein, the term "primer" refers to a natural or synthetic oligonucleotide that can serve as a starting point for nucleic acid synthesis when forming a duplex with a polynucleotide template and extends from its 3' end along the template to form an extended duplex. The nucleotide sequence added during the extension process is determined by the template polynucleotide sequence. Typically, primers are extended by DNA polymerase. The primer length is typically adapted to its use in the synthesis of primer extension products, typically ranging from 8 to 100 nucleotides in length, such as 10 to 75, 15 to 60, 15 to 40, 18 to 30, 20 to 40, 21 to 50, 22 to 45, 25 to 40, and the like. Typical primers can range from 10 to 50 nucleotides in length, such as 15 to 45, 18 to 40, 20 to 30, 21 to 25, and the like, as well as any length within the range. In some embodiments, a primer is typically no more than about 10, 12, 15, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 55, 60, 65, or 70 nucleotides in length.
[0059] As used herein, the term "plurality" includes at least 2. In some cases, a plurality can be at least 10, at least 100, at least 100, at least 10,000, at least 100,000, at least 106, at least 107, at least 108, or at least 109 or more.
[0060] If two nucleic acids are "complementary," every base of one nucleic acid will base pair with the corresponding nucleotide of the other nucleic acid. Two nucleic acids need not be completely complementary to hybridize to each other.
[0061] As used herein, the term "sequencing" refers to a method by which the identity of at least 10 consecutive nucleotides of a polynucleotide is obtained (e.g., the identity of at least 20, at least 50, at least 100, or at least 200 or more consecutive nucleotides).
[0062] As used herein, the term "next generation sequencing" or "high-throughput sequencing" refers to so-called parallel synthesis sequencing or ligation sequencing platforms currently employed by Illumina, Life Technologies, and Roche, among others. Next generation sequencing methods may also include nanopore sequencing methods such as those commercialized by Oxford Nanopore Technologies, electrical detection-based methods such as the Ion Torrent method commercialized by Life Technologies, or single molecule fluorescence-based methods such as those commercialized by Pacific Biosciences.
[0063] As used herein, the term "oligonucleotide binding site" refers to the site in a target polynucleotide to which an oligonucleotide hybridizes. If an oligonucleotide "provides" a primer binding site, then a primer can hybridize to that oligonucleotide or its complementary strand.
[0064] The term "strand" as used herein refers to a nucleic acid consisting of nucleotides covalently linked together by covalent bonds, such as phosphodiester bonds. In cells, DNA is typically present in double-stranded form and therefore has two complementary nucleic acid chains, referred to herein as the "upper strand" and the "lower (bottom) strand". In some cases, the complementary strands of a chromosomal region may also be referred to as the "positive strand" and the "negative strand", the "first strand" and the "second strand", the "coding strand" and the "non-coding strand", the "Watson strand" and the "Crick strand" or the "sense strand" and the "antisense strand". The designation of the upper strand or the lower strand is arbitrary and does not imply any specific direction, function or structure.
[0065] As used herein, the term "tagging" refers to the addition of a sequence marker (which includes an identifier sequence) to a nucleic acid molecule. Sequence markers can be added to the 5' end, the 3' end, or both ends of the nucleic acid molecule. Sequence markers can be added to fragments and then ligated to the fragments using adapters, such as T4 DNA ligase or other ligases.
[0066] The term "marker," also referred to as "marker," as used herein refers to a parameter that can distinguish between two states.
[0067] The terms "sample identifier sequence" and "sample index" are nucleotide sequences attached to a target polynucleotide, wherein the sequence identifies the source of the target polynucleotide (i.e., the sample from which the target polynucleotide originates). When in use, each sample is labeled with a different sample identifier sequence (e.g., one sequence is attached to each sample, and different sequences are attached to different samples) and the labeled samples are combined. After the combined samples are sequenced, the sample identifier sequence can be used to identify the source of the sequence. The sample identifier sequence can be added to the 5' end of the polynucleotide or the 3' end of the polynucleotide. In some cases, a portion of the sample identifier sequence can be at the 5' end of the polynucleotide, and the remainder of the sample identifier sequence is at the 3' end of the polynucleotide. When the sample identification element has a sequence at each end, together, the 3' and 5' sample identification sequences identify the sample. In many examples, the sample identifier sequence is simply a subset of the bases attached to the target oligonucleotide.
[0068] The term "free DNA", also known as "circulating cell-free DNA" (cfDNA), used herein refers to DNA circulating in the patient's peripheral blood. The DNA molecules in the cell-free DNA may have a median size of less than 1 kb (e.g., in the range of 50 bp to 500 bp, 80 bp to 400 bp, or 100 to 1,000 bp), although fragments with a median size outside this range may also be present. Cell-free DNA may include circulating tumor DNA (ctDNA), i.e., tumor DNA that circulates freely in the blood of a cancer patient or circulating fetal DNA (if the subject is a pregnant female). cfDNA can be highly fragmented and in some cases can have a median fragment size of about 166 bp. cfDNA can be obtained by centrifuging whole blood to remove all cells and then isolating DNA from the remaining plasma or serum. Such methods are well known (see, for example, Lo et al., Am J Hum Genet 1998; 62: 768-75). cfDNA is mostly double-stranded, but can be made single-stranded by denaturation.
[0069] As used herein, the term "amplify" refers to the production of one or more copies of a target nucleic acid using the target nucleic acid as a template.
[0070] As used herein, the term "copy of a fragment" refers to a product of amplification, wherein the copy of the fragment can be the reverse complement of the fragment strand or have the same sequence as the fragment strand.
[0071] The present invention relates to the following methods for detecting methylation genes:
[0072] 1. Whole Genome Bisulfite Sequencing (WGBS)
[0073] WGBS technology uses sodium bisulfite to treat DNA, causing unmethylated C to become U, and then become T in the subsequent PCR and sequencing process, while methylated C is not affected, so that it can be identified which sites are methylated during detection. Sodium bisulfite treatment of DNA can easily cause damage, especially CpG islands contain a large number of unmethylated C, so coverage in this area is likely to be low. The library construction method of WGBS technology can be divided into two types. One is to first fragment the DNA and connect the adapters, and then treat the C->T with sulfite; the other is to first convert C->T, and then connect the adapters for expansion. The latter requires lower DNA input. It should be noted that WGBS technology cannot distinguish between 5-hmC and 5-mC.
[0074] Bismark, a commonly used WGBS alignment software, pre-processes the reference genome sequence with C->T and G->A conversions. During alignment, each read undergoes the same C->T and G->A conversions. This combination results in four different alignments for each read. The optimal alignment is selected from these to determine the strand direction and potential methylation sites.
[0075] 2. Reduced Representation Bisulfite Sequencing (RRBS) technology
[0076] RRBS technology works by enriching genomic DNA fragments rich in CCGG sites through restriction enzyme digestion. Bisulfite treatment and high-throughput sequencing are then used to perform single-base resolution methylation sequencing within CpG-rich regions of the genome. Compared to WGBS, RRBS is a cost-effective methylation sequencing solution that significantly reduces sequencing workload and has broad application value in large-scale clinical sample studies.
[0077] RRBS utilizes the ability of bisulfite to convert unmethylated cytosine (C) to thymine (T). By treating the genome with bisulfite and sequencing it, the methylation rate can be calculated based on the ratio of the number of reads that were not converted to C or T at a single C site to the total number of reads covered. This technology has important applications in comprehensively studying epigenetic mechanisms of embryonic development, aging, disease progression, and screening for disease-associated epigenetic markers.
[0078] The process of RRBS technology includes: (1) Testing of DNA samples, mainly including two methods: Qubit to accurately quantify DNA concentration; agarose gel electrophoresis to analyze the degree of DNA degradation and whether there is contamination. (2) Different extraction schemes are used according to different sample types to obtain high-quality genomic DNA; Qubit to detect the concentration of DNA samples, and agarose gel electrophoresis to detect the integrity of DNA samples; CpG fragments are enriched by enzyme digestion; adapter sequences and index sequences are introduced at both ends of DNA molecules; electrophoresis is used to recover and select 250-500bp fragments rich in CpG; λDNA is added to the recovered genomic DNA, and then the non-methylated base C is converted to U by bisulfite treatment; PCR amplification enrichment library and purification of PCR products to obtain the final library. (3) After the library is constructed, Qubit2.0 is used for preliminary quantification, and the library is diluted to 1ng / μl. Then, the insert size of the library is detected using Agilent 2100. If it meets the expectations, the effective concentration of the library is accurately quantified using the qPCR method to ensure the quality of the library. (4) Sequencing: After the library is qualified, different libraries are pooled according to the effective concentration and target data volume requirements and then sequenced on the Illumina Nova platform.
[0079] 3. Enzymatic conversion methylation detection - EM-seq sequencing
[0080] EM-seq sequencing combines enzymatic methylation library preparation with a targeted methylation capture system to analyze the methylation status of 3.98 million CpG sites across a 123-Mb region of the human genome. EM-seq sequencing is ideal for probing methylation levels in a variety of applications, including cancer metastasis, human development, and functional genomics. This method offers broad coverage, high sensitivity, and reduced sequencing costs, making it suitable for large cohort studies, particularly for methylation studies in cfDNA samples.
[0081] Enzymatic conversion requires a relatively low sample input amount; generally, 10-200 ng of DNA is sufficient. The DNA is sheared (e.g., by ultrasound; cfDNA is naturally degraded and does not require further shearing) and then subjected to a two-step enzymatic conversion process to distinguish unmethylated cytosine from 5mC and 5hmC. The enzymatic treatment process is gentler on DNA, minimizing damage. As a result, the resulting DNA is more intact, with more and longer inserts in the library, ultimately yielding longer sequences and higher alignment rates.
[0082] 4. TAPS (TET-assisted pyridine borane sequencing) technology
[0083] The core of TAPS technology is to convert methylated C bases to T bases without the use of bisulfite. Experiments have demonstrated that TET proteins can oxidize 5-methylcytosine (5mC) and 5-hydroxymethylcytosine (5hmC), ultimately converting them to 5-carboxycytosine (5caC). Furthermore, 5caC can be reduced to dihydrouracil (DHU) by pyridine borane and its derivatives (such as pyridine borane). DHU behaves identically to the natural U base in terms of chain amplification, allowing it to be converted to T after PCR amplification, thus achieving the goal of converting methylated C bases to T bases without the need for bisulfite treatment. In fact, not only 5caC can be reduced by pyridine borane, but 5-formylcytosine (5fC) can also be reduced to DHU, and both reduction efficiencies are very high. In addition, the reduction reactions of 5caC and 5fC with pyridine borane can be blocked by 1-ethyl-3-(3-dimethylaminopropyl)carbodiimide and hydroxylamine conjugation, respectively. This greatly increases the flexibility of TAPS technology. If combined with other methylation sequencing technologies on the market, TAPS technology can theoretically measure the several methylation conditions mentioned above.
[0084] Since bisulfite is not used, the advantages of TAPS technology are mainly reflected in the following aspects: (1) shorter experimental time; (2) milder experimental conditions, room temperature is sufficient, and the effect on DNA fragments is small, basically no DNA degradation is caused, and longer DNA fragments, up to 10kb in length, can be retained; the library constructed in this way has more unique reads, that is, the library has higher complexity; (3) better sequencing quality, because TAPS technology converts methylated C bases into T bases, and methylated C bases account for a very small proportion in the genome, and the conversion has almost no effect on the base balance of the library, thereby improving the quality of base sequencing.
[0085] The present invention discloses a method for screening free DNA markers, comprising the following steps:
[0086] Collect verified positive experimental samples (positive samples) and control group samples (negative samples).
[0087] The genomic coordinates of all CpG sites in the human genome (theoretically, any version or patch can be used; in this project, the 13th patch version of GRCh38, i.e., GRCh38.p13) are retrieved. The first CpG site in the positive experimental sample is used as a candidate segment. For each candidate segment, its next adjacent CpG site is examined. If the interval between this CpG site and the last CpG site in the candidate segment is within a specified range (e.g., 10 bases), the CpG site is merged into the current candidate segment, and further CpG sites are examined. If the interval is greater than the specified range, the current candidate segment is examined. If the number of CpG sites it contains reaches a preset threshold (e.g., 3, 4, or 5), the candidate segment is marked as a formal segment; otherwise, the candidate segment is discarded. After the candidate segment is examined, the next adjacent CpG site is used as the starting point for a new candidate segment, and this process is repeated until all CpG sites have been examined. The CpG segmentation method of the present invention can achieve the following: 99% of the segments are less than 63 bp, with a median of 17 bp, which is very suitable for experimental verification and kit design in free DNA; each segment contains 3 or more CpG sites, which can make up for the shortcomings of sequencing data depth and improve the reliability of markers; it is not based on known functional annotation elements such as CpG islands, and can identify potential markers located in enhancer regions (with better tissue specificity than gene promoters).
[0088] For all experimental samples, for all CpG segments identified above, the methylation quantitative values for all CpG sites within the segment were averaged as the methylation level for that segment. Averaging can be performed using a variety of methods, such as the arithmetic mean. Because the sequencing depth of each methylation site in the data varies, this project employed a weighted average, weighted by sequencing depth. This method is the most commonly used average calculation method in methylation analysis. Other methods include the arithmetic mean or median of the methylation levels for all CpG sites.
[0089] For each CpG segment, if the number or proportion of its methylation levels in all control groups is less than a preset threshold a (for example, 10%) is not less than x (for example, 90%), and the number or proportion of samples in the disease group that are greater than another preset threshold b (for example, 20%, which may be the same as a) is greater than a preset threshold y (for example, 25%), then the CpG segment is regarded as a candidate marker; or if the number or proportion of its methylation levels in all control groups is greater than a preset threshold aa (for example, 90%) is not less than a preset threshold xx (for example, 90%), and the number or proportion of samples in the disease group that are less than another preset threshold bb (for example, 80%, which may be the same as aa) is greater than a preset threshold yy (for example, 25%), then the CpG segment is regarded as a candidate marker.
[0090] Therefore, by setting a preset threshold, several CpG segment screening results of different grades can be determined, and their detection sensitivity also varies with the set preset threshold.
[0091] For example, the preset threshold a is 10%, the preset threshold x is 90%, the preset threshold b is 20%, the preset threshold y is 25%, the preset threshold aa is 90%, the preset threshold xx is 90%, the preset threshold bb is 80%, and the preset threshold yy is 25%.
[0092] Or for example, the preset threshold a is 5%, the preset threshold x is N-2 (N is the number of samples in the control group), the preset threshold b is 20%, the preset threshold y is 33%, the preset threshold aa is 95%, the preset threshold x is N-1 (N is the number of samples in the control group), the preset threshold bb is 80%, and the preset threshold yy is 33%. Preferably, the present invention adopts a scheme in which the preset threshold a is 5%, the preset threshold x is N-1 (N is the number of samples in the control group), the preset threshold b is 5%, the preset threshold y is 25%, the preset threshold aa is 95%, the preset threshold xx is N-1 (N is the number of samples in the control group), the preset threshold bb is 95%, and the preset threshold yy is 25%.
[0093] In the above scheme, the marker identification parameters of the present invention are not based on statistical tests. This is due to the difference in emphasis between disease diagnosis and biological research. Statistically different sites are often not effective as markers, especially sites where the control group is not very clean, such as an average of 50% for the control group and an average of 70% for the disease group. This may be statistically significant, but it is basically unusable for diagnosis, especially when the disease signal is weak and the source is unclear (such as early-stage disease). The present invention does not pursue a single marker to achieve a good effect (which often leads to over-optimization (over-fitting)), but identifies a certain amount of potential markers to enhance anti-interference ability (such as a single marker is susceptible to primer bias, individual SNPs, etc.). The present invention can also be further combined with data processing technology, such as machine learning techniques such as random forests, support vector machines, neural networks, gradient boosting decision trees, etc., or with simpler methods such as arithmetic mean, to achieve a better diagnostic effect.
[0094] In the above scheme, the positive test sample is, for example, a blood or body fluid sample, such as a peripheral blood sample, from an AD patient group whose cerebral cortical Aβ plaque pathology is positive as determined by Aβ PET. The control sample can be a blood or body fluid sample from an AD patient group whose cerebral cortical Aβ plaque pathology is negative as determined by Aβ PET. Furthermore, the blood or body fluid sample is preferably obtained non-invasively and tested within 6 hours of ex vivo to ensure the integrity of the cell-free DNA in the test sample.
[0095] In the above scheme, the specific methylation sequencing method can be, for example, the above-mentioned well-known sequencing methods, such as WGBS technology, RRBS technology, EM-seq sequencing method and TAPS technology, or digital PCR, real-time fluorescence quantitative PCR and methylated DNA specific recognition enzyme binding technology can be used to measure the methylation degree.
[0096] In the above scheme, for example, the following data processing technologies may also be considered:
[0097] The mean methylation levels of all markers included in the analysis;
[0098] A statistical value of all markers included in the analysis, such as Q1, median, Q3, etc.;
[0099] The proportion of non-zero values among all markers included in the analysis;
[0100] Develop classification methods using machine learning techniques, including but not limited to decision trees, random forests, support vector machines, neural networks, gradient boosted decision trees, and more.
[0101] The method for screening free DNA markers of the present invention can screen the positions of all CpG sites in the human genome GRCh38.p13, or can screen only some of the positions.
[0102] Through screening, the present invention can obtain a series of candidate markers. Each candidate marker, as an independent gene fragment, can be used as a free DNA marker to accurately determine whether the corresponding gene fragment is methylated. Therefore, each candidate marker has its unique use value and protection value. The present invention also uses it as a composition of gene fragments for special protection.
[0103] Such gene fragment compositions include, for example, the nucleotide sequences shown in Table 1.
[0104] Table 1 Composition of partial gene fragments screened by the method of the present invention
[0105] When it is used for free DNA markers to perform methylation gene detection machine sequencing as described above, it has the advantages of high resolution and good accuracy. The cfDNA methylation markers screened by the present invention combined with data processing technology can effectively distinguish healthy controls from early AD patients with an accuracy rate of more than 92%, which is significantly better than the currently commercialized plasma protein markers.
[0106] The present invention also discloses a marker obtained by screening according to the above screening method.
[0107] The present invention also discloses a primer, a primer-amplified strand, or an enriched cell-free DNA sample composition obtained by PCR amplification, replication, conversion, and / or transformation based on the markers described above. The cell-free DNA sample composition, for example, includes a cell-free DNA molecule connected to an adapter, an internal standard control composition including the amplicon, and a streptavidin carrier, and the cell-free DNA sample composition contains the candidate markers described above.
[0108] The present invention also discloses a methylation machine sequencing system, which comprises the nucleic acid library of candidate markers as described above.
[0109] In summary, the present invention discloses and claims protection for a composition of gene fragments obtained by screening based on the above-mentioned screening method for free DNA markers, as well as primers, primer amplification chains, and sample compositions formed by PCR amplification, replication, conversion, and transformation of these gene fragment compositions, a nucleic acid library with unique markers formed by incorporating the above-mentioned gene fragment composition into a nucleic acid library, and a nucleic acid sequencing system using the same, such as an Illumina nucleic acid library, and an Illumina nucleic acid sequencing system using such a nucleic acid library.
[0110] The present invention will be further described below through specific examples. It should be noted that the following examples are only for illustration and are not intended to limit the present invention.
[0111] Example
[0112] 1. Patient recruitment and blood sample processing
[0113] All patients included in the analysis underwent Aβ PET to determine whether their cortical Aβ plaque pathology was negative (control group) or positive (AD patient group). Blood processing was routine: peripheral blood (e.g., 6 ml) was collected from a vein using a blood collection tube containing an anticoagulant (e.g., EDTA). The blood was centrifuged at a low temperature (e.g., 4°C) and a moderate speed (e.g., 1600 g acceleration) for a period of time (e.g., 15 minutes). The supernatant was then aspirated and transferred to a new centrifuge tube. The blood was then centrifuged again at a low temperature (e.g., 4°C) and a high speed (e.g., 16,000 g acceleration) for a period of time (e.g., 15 minutes). The supernatant (i.e., plasma) was aspirated and transferred to a new centrifuge tube or cryovial. Plasma can be stored in an ultra-low temperature freezer (e.g., -80°C) until use.
[0114] 2. EM-seq data analysis and marker identification
[0115] Free DNA was extracted using a free DNA extraction kit (commercially available), and the free DNA was processed and library built using EM-seq technology. The library was then sequenced using second-generation sequencing technology (such as the Illumina Nova Seq 6000 sequencer) to obtain the methylation level of each CpG site. The sequencing data was analyzed using bioinformatics methods, including accusation, alignment, and methylation identification, to obtain the methylation level of each CpG site. The present invention uses Msuite2 software for data analysis, which includes a data quality control module that can remove low-quality sequencing data and allows the removal of some sequences during analysis to reduce the impact of free DNA end defects (such as jagged ends) and improve the accuracy of the results.
[0116] The positions of all CpG sites in the human genome (any major version, any patch version can be used, the 13th patch version of GRCh38, i.e. GRCh38.p13, is used in the present invention, and the older version GRCh 37 or the latest version T2T CHM13 v2.0 / hs1, Hans1, etc. can also be used) are taken out. The first CpG site is used as a candidate segment. For a candidate segment, if the interval between its next adjacent CpG and the last CpG in the segment is within a specified range (e.g., 10 bases), it is merged into the current candidate segment and the subsequent CpG sites are continued to be investigated; if the interval is greater than the specified range, the current candidate segment is investigated. If the number of CpG sites it contains reaches a preset threshold (e.g., 3), the candidate segment is marked as a formal segment, otherwise it is discarded, and the next CpG site is used as a new candidate segment to start, and the process is repeated until all CpG sites are investigated.
[0117] For all experimental samples, for all CpG segments identified above, the methylation quantitative values for all CpG sites within the segment were averaged as the methylation level for that segment. Averaging can be performed using various methods, such as the arithmetic mean. Because the sequencing depth of each methylation site in the data varied, this project employed a weighted average, which is the most commonly used method for calculating averages in methylation analysis.
[0118] For each CpG segment, if its methylation level in all control groups is less than a preset threshold a (=10%), and the proportion of samples in the disease group that are greater than another preset threshold b (=20%) is greater than a preset threshold c (=25%), then the CpG segment is regarded as a candidate marker; or if its methylation level in all control groups is greater than a preset threshold aa (=90%), and the proportion of samples in the disease group that are less than another preset threshold bb (=80%) is greater than a preset threshold cc (=25%), then the CpG segment is regarded as a candidate marker.
[0119] 3. AD-related biomarkers
[0120] Through experimental screening, 513 markers were obtained as shown in Table 1. The sequences and positional information of these markers in GRCh 38 are provided. Note that each marker has a different length and contains a different number of CpG groups (all containing at least 3). For each marker, each CpG site can also be considered a separate marker.
[0121] 4. The 513 markers measured were reused in the above-mentioned control group (healthy elderly) and AD patient group (early AD patients), and the marker sites of the present invention were determined using the whole-genome methylation sequencing method to verify the experimental results. Among them, two sample sets were collected, as shown in Appendix 1 and Appendix 2. Sample set 1, which served as the discovery set, tested 19 groups of healthy elderly people and 24 groups of early AD patients for 513 markers (due to space limitations, only 19 groups are posted in Appendix 2); as shown in Appendix 3 and Appendix 4, Sample set 2, which served as the validation set, tested 17 groups of healthy elderly people and 13 groups of early AD patients for 513 markers.
[0122] 5. Graphical display of some of the above test results
[0123] Figure 1 is a dot-line plot comparing the free DNA methylation rate (methylation level) of the marker SEQ ID No. 21 in an early-stage AD patient (asterisk) and a healthy elderly control (×) (each dot in the figure represents a CpG site). As can be seen from the figure, this marker can significantly distinguish healthy elderly people from early-stage AD patients, thereby playing a role in early detection and early warning, improving the recognition rate of initial diagnosis, and facilitating early intervention treatment and rehabilitation for patients.
[0124] Figures 2A and 2B show the distribution of the arithmetic mean of the methylation levels of all markers in healthy elderly individuals (left) and AD patients (right), respectively, based on Aβ PET imaging, from two independent datasets (each dot represents a patient). Sample Sets 1 and 2 were defined as described above. In these two datasets, using the average free DNA methylation level across all sites as a predictive index, we found that it could effectively distinguish healthy elderly individuals from early AD patients, with the following accuracy rates:
[0125] Sample set 1: sensitivity = 100%, specificity = 100%;
[0126] Sample set 2: Sensitivity = 92.3%, Specificity = 94.1%.
[0127] The present invention can also use any number of sites for prediction. For example, FIG1 uses one site, and FIG2 uses all sites. Alternatively, 2 to 512 sites can be randomly selected to calculate the average.
[0128] In addition, it should be noted that some values in the two data sets are "NA", which means that the sequencing depth of the corresponding samples at this site is low. This site will be ignored when calculating indicators such as the average value and will not be included in the average calculation.
[0129] Figure 3 is a performance comparison diagram (ROC curve diagram) of the markers of the present invention and other AD markers (Aβ 42 / Aβ 40, p-Tau181, NfL, GFAP) in the classification of healthy controls and AD patients. Among them, DNA methylation-dataset 1 (AUC=1), DNA methylation-dataset 2 (AUC=0.96), plasma p-Tau181 (AUC=0.89), plasma NfL (AUC=0.86), plasma GFAP (AUC=0.76), plasma Aβ 42 / Aβ 40 (AUC=0.80). Through ROC analysis, it can be seen that the scheme of the present invention (①, ②) has better predictive accuracy in early AD patients than existing plasma protein markers (③~⑥).
[0130] Figure 4 compares the performance of the cell-free DNA markers of the present invention and other AD markers (Aβ42 / Aβ40, p-Tau181, NfL, and GFAP) in predicting the intensity of Aβ accumulation signals in the patient's brain. A comparison between A and other markers B through E in Figure 4 shows that the cell-free DNA methylation markers of the present invention have a better correlation with brain Aβ signal intensity than plasma protein markers, indicating that the markers of the present invention have a better accuracy rate in dynamic monitoring.
[0131] The specific embodiments described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above are only specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for screening free DNA markers, characterized in that: The following steps are involved: For a gene sequence, the first CpG site is used as a candidate segment; for any candidate segment, the next adjacent CpG site is examined, and if the interval between the CpG site and the last CpG in the candidate segment is less than or equal to a first preset threshold, the CpG site is merged into the current candidate segment, and the subsequent CpG sites are continuously examined until the interval condition is not met; Inspect the current candidate segment, and if the number of CpG sites contained in the segment reaches a second preset threshold, mark the candidate segment as a formal segment, otherwise discard the candidate segment; After the candidate segment is examined, the next adjacent CpG site is used as a new candidate segment to start, and the above process is repeated until the preset end condition is reached; Collect verified positive experimental samples and control group samples; The positive experimental samples and the control group samples were measured by DNA methylation determination method: For all samples, for all formal segments determined above, the methylation quantitative values of all CpG sites contained therein are averaged as the methylation level of the formal segment; For each formal segment, if the number or proportion of samples in the control group whose methylation levels are less than a preset threshold a is not less than x, and the number or proportion of samples in the positive experimental samples whose methylation levels are greater than another preset threshold b is greater than a preset threshold y, then the formal segment is taken as a candidate marker; or For each formal segment, if the number or proportion of samples in the control group whose methylation levels are greater than the preset threshold aa is not less than the preset threshold xx, and the number or proportion of samples in the positive experimental samples whose methylation levels are less than another preset threshold bb is greater than the preset threshold yy, then the formal segment is also taken as a candidate marker.
2. The screening method according to claim 1, characterized in that The positive experimental sample is a blood or body fluid sample whose cerebral cortex Aβ plaque pathology is positive as determined by Aβ PET; preferably a peripheral blood sample; The control group samples are blood or body fluid samples whose cerebral cortex Aβ plaque pathology is determined to be negative by Aβ PET; Preferably, the blood or body fluid sample is obtained by a non-invasive method; further preferably, the sample is a peripheral blood sample obtained by a non-invasive method and within 6 hours of ex vivo.
3. The screening method according to claim 1, characterized in that The DNA methylation determination method is selected from WGBS technology, RRBS technology, EM-seq sequencing method, TAPS technology, digital PCR, real-time fluorescence quantitative PCR and methylated DNA specific recognition enzyme binding technology; The first preset threshold is 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90 or 100; The second preset threshold is 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10.
4. The screening method according to claim 1, characterized in that The method for calculating the average value is selected from arithmetic mean, median average, and weighted average.
5. The screening method according to claim 1, characterized in that The preset threshold a is 10%, the preset threshold x is 90%, the preset threshold b is 20%, the preset threshold y is 25%, the preset threshold aa is 90%, the preset threshold xx is 90%, the preset threshold bb is 80%, and the preset threshold yy is 25%; or The preset threshold a is 5%, the preset threshold x is N-2, the preset threshold b is 20%, the preset threshold y is 33%, the preset threshold aa is 95%, the preset threshold xx is N-1, the preset threshold bb is 80%, and the preset threshold yy is 33%; wherein N is the number of samples in the control group; or The preset threshold a is 5%, the preset threshold x is N-1, the preset threshold b is 5%, the preset threshold y is 25%, the preset threshold aa is 95%, the preset threshold xx is N-1, the preset threshold bb is 95%, and the preset threshold yy is 25%; wherein N is the number of samples in the control group.
6. The screening method according to claim 1, characterized in that In the step of determining candidate markers based on formal segmentation, machine learning technology is also used, preferably random forest, support vector machine, neural network, and gradient boosting decision tree methods to optimize the screening results.
7. The screening method according to claim 1, characterized in that The positions of all or part of the CpG sites on all chromosomes in the human genome are screened; The human genome data are from GRCh38, GRCh37, GRCh36, T2T CHM13v2.0 / hs1 or Han1.
8. The candidate marker obtained by screening according to any one of claims 1 to 7; Preferably, the candidate marker comprises the nucleotide sequence as described in SEQ ID No.1 to SEQ ID No.
513.
9. Primers, primer amplification chains or sample compositions obtained by PCR amplification, replication, conversion and / or transformation based on the candidate marker according to claim 8. 10 . A methylation machine sequencing system, comprising the nucleic acid library of candidate markers according to claim 8 .
Citation Information
Patent Citations
Method for screening prognostic markers of DNA methylation in acute myeloid leukemia
CN109852672A
Marker screening method based on methylation sequencing and cancer detection method and device
CN116312739A
System and method for screening large-fragment methylation markers
CN117059163A
Screening method of free DNA marker, DNA marker and application of DNA marker
CN118186057A
Cell-free DNA hydroxymethylation profiles in the evaluation of pancreatic lesions
US20200123616A1