Methods for generating double-stranded DNA libraries and sequencing methods for identifying methylated cytosine.
By linking adaptors to double-stranded DNA molecules and transforming them with cytosine, followed by sequencing with aptamers and hairpin sequences, the ambiguity and error problems in DNA methylation analysis in existing technologies have been solved, enabling efficient and low-cost construction of double-stranded DNA libraries and methylation detection.
Patent Information
- Application Number
- CN201580012375.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2014-01-07
- Filing Date
- 2015-01-07
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2035-01-07
AI Technical Summary
Existing DNA methylation analysis methods cannot directly analyze the original material, requiring replication and amplification, which leads to ambiguity, insufficient coverage, computational complexity and high cost, and makes it difficult to simultaneously detect primary sequence and epigenetic modifications.
By linking double-stranded DNA adaptors to the ends of DNA molecules, unmethylated cytosine is converted into uracil, and sequencing is performed using complementary strands to ensure sequence fidelity and error control. Aptamers and hairpin sequences are used for amplification and sequencing.
It enables efficient and low-cost construction of double-stranded DNA libraries, which can simultaneously detect sequence variations and methylation modifications, reduce the demand for starting materials, improve sequencing quality and coverage, and reduce errors and biases.
Smart Images

Figure CN106103743B_ABST
Abstract
Description
Invention Field
[0001] This invention relates to methods for determining the sequence of a population of double-stranded DNA molecules and for identifying methylated cytosine in a population of double-stranded DNA molecules. The invention also relates to aptamers and kits for synthesizing said aptamers, as well as double-stranded DNA libraries that can be used in the methods of the invention. Background of the Invention
[0003] The analysis of the primary structure of nucleic acids (such as DNA and RNA), including epigenetic modifications (i.e., DNA methylation), can be achieved using various techniques commonly referred to as "sequencing".
[0004] All currently available methods do not directly analyze the original material. They require processing or transformation of the original template, generation of replicas, and often amplification of the replicas. The resulting replicas (named genomic libraries) are suitable for sequencing using one or more of the currently available sequencing technologies (such as Illumina, Roche, or IonTorrent sequencing platforms).
[0005] Sequencing can be performed at a small scale (analyzing selected fragments) or at a high throughput (also known as genome-scale) (analyzing the entire or a large portion of the overall material). The length of the fragments that can be analyzed depends on the sequencing method used. Current sequencing technologies evaluate DNA strands separately for genome-scale and large-scale sequencing of specific loci.
[0006] The current gold standard for assessing DNA methylation involves chemically converting nucleic acids with bisulfite. This leads to ambiguity because unmethylated cytosine is converted to uracil and appears as thymine, making them indistinguishable from actual thymine in every sequencing method. This information reduction represents a challenge for genome-scale methods because there are still some unresolved drawbacks that limit their application, such as:
[0007] 1) Independent methods must be used to determine primary sequences (i.e., for detecting mutations or genetic variants) and epigenetic modifications (i.e., methylation of cytosine);
[0008] 2) The resulting ambiguity limits efficiency (a large proportion of sequence reads are abandoned due to ambiguity) and coverage (some areas cannot be analyzed), and involves a demanding computational process;
[0009] 3) A large amount of starting material is required for high-coverage studies;
[0010] 4) Uncontrolled bias limits quantification; and
[0011] 5) Sequencing errors are almost impossible to detect by the system.
[0012] Another method is the so-called hairpin-bisulfite PCR method (see Laird et al., 2004, Proc. Natl. Acad. Sci. USA 101, 204-209; Riggs and Xiong, 2004, Proc. Natl. Acad. Sci. USA 101, 4-5). In this method, the two complementary strands are covalently linked by a hairpin loop sequence before bisulfite treatment. However, this method is only suitable for specific double-stranded molecules and is not suitable for sequencing populations of double-stranded DNA molecules, and particularly not suitable for identifying methylated cytosine in populations of double-stranded DNA molecules.
[0013] Therefore, it is of interest to develop additional methods for determining the sequence of populations of double-stranded DNA molecules and, in particular, for identifying methylated cytosine in populations of double-stranded DNA molecules, methods that can overcome all or some of the disadvantages mentioned above associated with existing methods.
[0014] WO2010 / 048337 discloses a method for identifying methylated cytosine, the method comprising the steps of: generating a complementary copy of a template nucleic acid using a bisulfite-resistant cytosine analog; optionally pairing the template nucleic acid and the complementary copy; converting non-methylated cytosine residues in the complementary copy and the template nucleic acid to uracil residues; and determining the nucleotide sequences of the bisulfite-converted template nucleic acid and the unconverted complementary copy. However, because both the bisulfite-converted template nucleic acid and the unconverted complementary copy are rich in methylated cytosine, these chains are difficult to process.
[0015] This invention solves these problems. Invention Overview
[0017] This invention relates to a method for identifying methylated cytosine in a population of double-stranded DNA molecules, the method comprising the following steps:
[0018] (i) Connecting a double-stranded DNA adapter to at least one end of the strands of a plurality of double-stranded DNA molecules, and pairing the strands of the plurality of double-stranded DNA molecules to provide a plurality of paired adapter-modified DNA molecules.
[0019] (ii) Converting any (unmethylated) cytosine in the paired, adaptor-modified DNA molecule into uracil in the paired, adaptor-modified DNA molecule;
[0020] (iii) Using nucleotides A, G, C and T and primers to provide complementary strands of paired and transformed adaptor-modified DNA molecules, wherein the sequences of the primers are complementary to at least a portion of the double-stranded adaptor to provide partially transformed paired double-stranded molecules;
[0021] (iv) Optionally, the partially converted paired double-stranded DNA molecules obtained in step (iii) are amplified to provide amplified paired double-stranded DNA molecules.
[0022] (v) Sequencing the paired DNA molecules obtained in step (iii) or step (iv).
[0023] If cytosine is present in one of the strands of the paired double-stranded DNA molecule obtained in step (iii) or step (iv) and guanine is present at the corresponding position in the other strand of the paired double-stranded DNA molecule, then the presence of methylated cytosine at the given position is determined, and / or if uracil or thymine is present in one of the strands of the paired double-stranded DNA molecule obtained in step (iii) or step (iv) and guanine is present at the corresponding position in the other strand of the paired double-stranded DNA molecule, then the presence of unmethylated cytosine at the given position is determined. Attached Figure Description
[0024] Figure 1 A schematic diagram illustrating one embodiment of the method of the present invention is shown. Ligation Step [Step (i)]. A genomic fragment (black line; B) from the sample preparation step, having overhanging ends (A and C), is ligated to two molecules: dsDNA (aptamer; D) and a hairpin (E). Capture Step. A biotin-labeled probe (F) hybridizes with the hairpin to remove the ligation product without the hairpin. Bisulfite and Extension Step [Step (iii)]. The ligation product is treated with bisulfite and loses complementarity (glow). This allows primer (H) to initiate polymerase extension (dashed line; I), followed by an amplification step (not shown). An exemplary sequence fragment of genomic fragment B is shown. The nucleotide sequence C*GTTGGAA and its complementary sequence TTCCAAC*G are treated with bisulfite, and TTCCAAC*G is converted to TTUUAAC*G. After the amplification step, the nucleotide sequences CGTTAAAA and TTCCAACG are obtained. C*: methylated cytosine.
[0025] Figure 2 A schematic diagram showing an extended step [step (iii)] of one embodiment of the method of the present invention. It has only one first aptamer (DBE and EBD, as shown below). Figure 1 The ligation product (mentioned in the text) will be extended with primer (H) to obtain the synthetic chain (I).
[0026] Figure 3 1. A schematic diagram illustrating one embodiment of the method of the present invention, wherein aptamer molecules are provided in a manner fixed in a support. 2. Distribution of aptamers (A) on a solid surface. 3. Ligation of genomic fragments (B). Only one A aptamer can be ligated to each genomic fragment. 4. Ligated fragments. 5. Ligation of hairpin aptamers (C) to the free end of the genomic fragment. 6. Bisulfite conversion and loss of complementarity. 7. Elongation step (step iii). The first polymerase elongation using primer (D) is shown to obtain the synthetic strand (E).
[0027] Figure 4 1. A schematic diagram illustrating one embodiment of the method of the present invention, wherein the aptamer molecule is provided in a manner immobilized in a support. 1. A genomic fragment (B) linked to a first aptamer molecule (A). 2. A hairpin aptamer (C) is linked to the free end of the genomic fragment. 3. Bisulfite conversion and loss of complementarity. 4. Primer (D) hybridizes with a portion of the aptamer molecule sequence. 5, 6, 7. Elongation steps (step iii). The first polymerase extension with primer (D) is shown to obtain the synthetic strand (E). 8. The template remains attached to the solid surface, and the extension product is released into the supernatant. The released molecule can be amplified with primer (F) (step iv). References to letters A, B, and C are shown in [reference needed]. Figure 3 Same as above.
[0028] Figure 5 A schematic diagram showing one embodiment of the method of the present invention. Ligation Step [Step (i)]. A genomic fragment (black line) from the sample preparation step is ligated to two Y-aptamers, each Y-aptamer being formed by a first DNA strand (A) and a second DNA strand (B), the second DNA strand being formed by a hairpin loop (C), and the first segment being located at the 3' end (D) of the 3' region. Extension Step. A synthetic sequence is generated using a hairpin as a primer for polymerase (dashed line, F). Bisulfite (Step ii). The molecule obtained after the extension step is bisulfite-treated, and the complementarity of the strands is lost. Complementary Strand Generation Step (Extension Step) [Step (iii)]. Primer (G) is added for the first round of amplification (dotted line; H).
[0029] Figure 6 A schematic diagram illustrating another embodiment of the method of the present invention is shown. The ligation step [step (i)] and the extension step [step (iii)] are as described above. A first round of amplification step [step (iv)] is shown, wherein a primer complementary to a portion of the complementary sequence of the first DNA strand of the aptamer molecule (G) or a primer complementary to a specific sequence complementary to the sequence of the genomic fragment used to generate the library (J) of the present invention is used. Primer pairs (G, I) or (J, K) can be used for second and subsequent rounds of amplification.
[0030] Figure 7 A schematic diagram illustrating one embodiment of the method of the present invention. Ligation Step [Step (i)]. A genomic fragment (black line) from the sample preparation step is ligated to two Y-aptamers, each Y-aptamer being formed from a first DNA strand (A) and a second DNA strand (B). An extension primer (D) hybridizes with the second strand of the Y-aptamer molecule to generate an overhanging end compatible with a hairpin aptamer (C). Polymerase extension is performed using the extension primer (D) to obtain the synthetic strand (dashed line).
[0031] Figure 8 A schematic diagram illustrating one embodiment of the method of the present invention is shown, wherein a hairpin aptamer (C) and an extension primer (D) are provided in the form of a complex. Ligation Step [Step (i)]: A genomic fragment (black line) from the sample preparation step is ligated to two Y-aptamers, each Y-aptamer being formed from a first DNA strand (A) and a second DNA strand (B). The complex formed by the hairpin aptamer (C) and the extension primer (D) hybridizes with the second strand of the Y-aptamer molecule and is used for polymerase extension to obtain the synthetic strand (dashed line).
[0032] Figure 9 The diagram illustrates two further embodiments of the method of the present invention, wherein the hairpin aptamer and the elongation primer (A or B) are provided in the form of a complex.
[0033] Figure 10 A schematic diagram illustrating one embodiment of the method of the present invention is provided in the form of a hairpin aptamer (hairpin sequence or hairpin molecule) (F) and an extension primer (E). Fragmentation and ligation steps: A genomic fragment (black line) is bound to a hemiaptamer molecule comprising a first DNA strand (A) and a second DNA strand (C), and having a combined sequence (B) in the first DNA strand. Substitution step: The second DNA strand (C) of the hemiaptamer is replaced with an alternative second strand (D). Gap filling step: A gap is filled between the 5' end of the alternative second strand and the 3' end of the DNA fragment. The complex formed by the hairpin aptamer (F) and the extension primer (E) hybridizes with the alternative second strand of the Y-aptamer molecule and is used for polymerase extension to obtain a synthetic strand (dashed line).
[0034] Figure 11 The diagram shows another embodiment of the method of the present invention, wherein an alternative second strand (D), a hairpin aptamer (F), and an elongation primer (E) are provided in the form of a complex.
[0035] Figure 12The diagram illustrates the amplification steps of products from several embodiments of the method of the present invention. It shows the distribution of the original sequence (A) and synthesized sequence (B) of each amplified product linked by hairpin aptamers (C).
[0036] Figure 13 The diagram illustrates several embodiments of the method of the present invention, wherein aptamers (C and D) contain different combinations of sequences. Combination barcodes (YY, XX, respectively) allow for unique labeling of the molecules. After the entire process, the complementary strands that were initially joined together will share the same two barcodes. This allows for tracking the two strands (A and B) of each double-stranded DNA fragment.
[0037] Figure 14 A schematic diagram showing one embodiment of the method of the present invention, wherein the aptamers are Y-aptamers. Ligation Step [Step (i)]. A genomic fragment (black line; A, B) from the sample preparation step is ligated to two Y-aptamers (C, D), each Y-aptamer being formed from a first DNA strand (grey) and a second DNA strand (white), wherein the aptamers contain different dsDNA combination sequences (XX and YY, respectively). Bisulfite Step [Step (ii)]. The molecules obtained after the ligation step are bisulfite treated, and the complementarity of the strands is lost (glow). Complementary Strand Generation Step (Extension Step) [Step (iii)]. Primers (E) are added for the first round of amplification (dotted line). Combination Barcoding Allows for Unique Labeling of Molecules. After the entire process, those complementary strands that were initially joined together will share the same two barcodes. This allows for tracking of the two strands (A and B) of each double-stranded DNA fragment.
[0038] Figure 15 A schematic diagram of a method for generating DNA Y-aptamers containing a composite sequence is shown. Hybridization step: A first single-stranded polynucleotide (A) is contacted with a second single-stranded polynucleotide (B), wherein the second polynucleotide has a composite sequence (C) and is reversibly blocked at its 3' end (black triangle; D). Elongation step: The 3' end of the first polynucleotide is extended to produce a sequence (E) complementary to the 5' region of the second polynucleotide. Unblocking step: The 3' end of the second polynucleotide is unblocked (white triangle).
[0039] Figure 16 A schematic diagram showing exemplary aptamers containing combined sequences used in the methods of the present invention and their synthesis methods. A process for generating different Y-aptamers according to several embodiments of the present invention. Invention Details
[0041] This invention relates to a method for identifying methylated cytosine in a population of double-stranded DNA molecules. In any of the described embodiments, this method ensures sequence fidelity and improves sequencing quality because both strands of the same DNA molecule are read simultaneously, and error and bias control are exhaustive.
[0042] Furthermore, due to more accurate sequencing, less coverage is required to obtain reliable reads, and less starting material is needed. Specifically, double-stranded DNA libraries produced by the methods of this invention can be generated from small amounts of DNA and a wide range of sample preparation sources (including those involving DNA fragmentation).
[0043] The method of the present invention in any of the described embodiments provides the further advantage that the sample used as a DNA template can be preserved during the method and can be recovered, stored, and amplified and sequenced multiple times under different conditions without depleting the sample. In particular, the aptamers and / or hairpin sequences and / or barcode sequences (as applicable) used in the method of the present invention in any of the described embodiments may have unique barcodes (also known as barcode sequences, combinatorial sequences, or combinatorial barcodes) and functional groups for sample identification to allow recovery of the original DNA template after the elongation or amplification steps. The barcodes may also exist as separate molecules, as described below.
[0044] The method of this invention is particularly useful for methylation sequence sequencing in all its embodiments. The double-stranded DNA library produced by the method of this invention retains well-defined DNA sequence and DNA methylation information, thereby allowing for the simultaneous detection of sequence variations (including polymorphisms and mutations) and DNA methylation modifications. In particular, the method of this invention allows for the determination of DNA methylation symmetry at the genome scale due to the parallel analysis of the two strands. By reading both strands, the sequencing process can be monitored, and errors arising in each individual sequence read can be corrected, thus obtaining more reliable information about both the genome and the methylome.
[0045] Furthermore, quantitative results regarding sequence variants (including polymorphisms and mutations) and DNA methylation modifications can be obtained by introducing combinatorial barcodes into the DNA template. The barcodes allow for monitoring biases introduced during sample processing (i.e., heterogeneous degradation of DNA) and amplification (i.e., different amplification efficiencies of sequence variants) for each library. To achieve this, the present invention provides a method for synthesizing combinatorial barcode DNA aptamers with ultra-high efficiency.
[0046] Furthermore, this invention provides for the generation of a library and the sequencing of methylated sequences, wherein the aptamers used contain unique combinatorial barcodes that allow tracking of the sense and antisense strands of the original DNA molecule. In summary, compared to methods using existing techniques, the entire method for obtaining a DNA library and sequencing it using the method of this invention is less labor-intensive (both manual and computer-based) and less expensive. It allows for the identification of methylated cytosine in both strands of the original double-stranded DNA molecule, preferably genomic DNA.
[0047] This invention relates to a method for identifying methylated cytosine in a population of double-stranded DNA molecules, the method comprising the following steps:
[0048] (i) Connecting a double-stranded DNA integrator to at least one end of the strands of a plurality of double-stranded DNA molecules and pairing the strands of the plurality of double-stranded DNA molecules to provide a plurality of paired integrator-modified DNA molecules;
[0049] (ii) Converting (unmethylated) cytosine present in both strands of the paired, adaptor-modified DNA molecules into uracil in the paired, adaptor-modified DNA molecules;
[0050] (iii) Using nucleotides A, G, C and T and primers to provide complementary strands of paired and transformed adaptor-modified DNA molecules, wherein the sequences of the primers are complementary to at least a portion of the double-stranded adaptor (as obtained after transformation step (ii)) to provide partially transformed paired double-stranded molecules.
[0051] (iv) Optionally, the partially converted paired double-stranded DNA molecules obtained in step (iii) are amplified to provide amplified paired double-stranded DNA molecules;
[0052] (v) Sequencing the paired DNA molecules obtained in steps (ii), (iii), or (iv) (preferably steps (iii) and / or (iv)).
[0053] If cytosine is present in one of the strands of the paired double-stranded DNA molecule obtained in step (iii) or step (iv) and guanine is present at the corresponding position in the other strand of the paired double-stranded DNA molecule, then the presence of methylated cytosine at the given position is determined; or if uracil or thymine is present in one of the strands of the paired double-stranded DNA molecule obtained in step (iii) or step (iv) and guanine is present at the corresponding position in the other strand of the paired double-stranded DNA molecule, then the presence of unmethylated cytosine at the given position is determined.
[0054] The method of the present invention allows for the acquisition of a double-stranded DNA library, wherein the original sense and antisense strands of the DNA molecule can be physically bound (if pairing occurs via hairpin molecules, as described below) after steps (i), (ii), (iii) and optionally (iv). Figure 1 A schematic diagram of the method of the present invention is shown in the figure.
[0055] As used herein, the term “DNA library” may refer to a collection of DNA fragments that have been ligated with aptamer molecules to identify and isolate DNA fragments of interest.
[0056] As used herein, the term "double-stranded DNA library" can refer to a library containing two strands (i.e., a sense strand and an antisense strand) of a DNA molecule, which are physically linked by one of their ends and form part of the same molecule. The strands of the double-stranded DNA molecules in a DNA library may also not be physically linked by one of their ends. As described below, they can be paired by the presence of a barcode sequence. The double-stranded DNA library of the method of the present invention is not a circular library. The original strands of the DNA molecule can be physically linked by a loop by one of their ends, thus forming a double-stranded structure between the sense and antisense strands. Each molecule in the double-stranded DNA library can also be in a linear conformation when the complementarity between the sense and antisense strands of the DNA molecule is partially or completely lost. Additionally, the original strands of the DNA molecule may not be physically linked by one of their ends, but rather paired by the presence of at least one barcode sequence.
[0057] The method of this invention requires a population or multiple double-stranded DNA molecules. As used herein, "a population or multiple double-stranded DNA molecules" refers to a collection of double-stranded DNA molecules, which may be, but are not limited to, genomic DNA (nuclear DNA, mitochondrial DNA, chloroplast DNA, etc.), plasmid DNA, or double-stranded DNA molecules obtained from single-stranded nucleic acid samples (e.g., DNA, cDNA, mRNA). In one embodiment, the population is formed from DNA fragments.
[0058] Preferably, the multiple double-stranded DNA molecules are genomic DNA. This can be a whole genome or a reduced representation of the genome. The DNA can be obtained, for example, by enrichment or by chromatin immunoprecipitation (ChiP).
[0059] The term "genomic DNA" refers to the heritable genetic information of an organism. Genomic DNA includes nuclear DNA (also known as chromosomal DNA), as well as DNA from plastids (e.g., chloroplasts) and other organelles (e.g., mitochondria). The term "genomic DNA" as used herein includes genomic DNA containing sequences complementary to those described herein.
[0060] Preferably, the multiple double-stranded DNA molecules are DNA fragments. DNA fragmentation is achieved by any suitable method, including but not limited to mechanical stress (sonication, atomization, cavitation, etc.), enzymatic fragmentation (digestion with restriction endonucleases, nicking endonucleases, exonucleases, etc.), and chemical fragmentation (dimethyl sulfate, hydrazine, NaCl, piperidine, acids, etc.). In principle, there is no limitation on the length of the fragmented DNA fragments, but a narrow length range is preferred. A suitable fragment size can be selected before step (i) of the first method of the invention. The optimal length ultimately depends on the available sequencing method. In a more preferred embodiment, the double-stranded DNA molecules are genomic DNA fragments.
[0061] The multiple double-stranded DNA molecules used in step (i) can be obtained as follows:
[0062] a) Provide a population of double-stranded DNA molecules derived from genomic DNA;
[0063] b) Separate double-stranded molecules derived from genomic DNA to provide single-stranded DNA molecules derived from genomic DNA;
[0064] c) Use nucleotides A, G, C, and T to provide complementary strands of single-stranded DNA molecules derived from genomic DNA to obtain the double-stranded DNA molecules used in step (i).
[0065] Preferably, the plurality of double-stranded DNA molecules linked to the adaptor comprise DNA molecules containing (unmethylated) cytosine on both strands and / or without methylated cytosine in one of the strands.
[0066] Typically, processing the ends of a population of double-stranded DNA molecules allows the sample to enter a specific process on the sequencing platform.
[0067] Optionally, the double-stranded DNA adaptor may contain a "cleavage site" (e.g., a "restriction site," i.e., a sequence of oligonucleotides recognized by restriction enzymes). The "cleavage site" increases the ways to adapt the final elements of the library to the needs of different sequencing platforms. While this adaptation can be achieved through the specific design of the double-stranded DNA adaptor (by introducing sequences compatible with platform reagents, such as sequencing primers), the cleavage site allows for modularity to add barcodes or aptamers for multiplexing (mixing samples from different sources) or to meet the needs of any platform used for large-scale sequencing (or also eliminate potentially unnecessary nucleotides). The "cleavage site" is a specific sequence that allows a known target to be present at the edge of multiple paired adaptor-modified DNA molecules (the library of paired adaptor-modified DNA molecules obtained in step (i), or the library of paired and transformed adaptor-modified DNA molecules obtained as in step (iii) (and optionally in step (iv)). The "cleavage site" can be ligated to multiple double-stranded DNA molecules before or after the step of ligating the adaptor and / or hairpin and / or barcode sequences. As mentioned above, the "cleavage site" may already be included in the aptamer and / or hairpin and / or barcode sequence. This allows all fragments to be cleaved and aptamers to be correctly ligated (thus, unwanted adaptors and / or hairpin and / or barcode sequences can be removed to improve sequencing efficiency).
[0068] Preferably, the ends of the double-stranded DNA molecules used in step (i) are repaired before step (i).
[0069] As used herein, the term "end repair" refers to the conversion of DNA fragments containing damaged or incompatible 5'- and / or 3'-hanging ends into blunt-ended DNA containing 5'-phosphate and 3'-hydroxyl groups. Blunt-end treatment of DNA can be achieved by enzymes including, but not limited to, T4 DNA polymerase (which has 5'→3' polymerase activity to fill 5'-hanging DNA ends) and the Klenow fragment of E. coli DNA polymerase I (which has 3'→5' exonuclease activity to remove 3'-hanging ends). For efficient phosphorylation of DNA ends, any enzyme capable of adding 5'-phosphate to the ends of unphosphorylated DNA fragments can be used, including but not limited to T4 polynucleotide kinase.
[0070] Preferably, the method of the present invention further includes a step of adding dA tails to the DNA molecule after the end repair step.
[0071] As used herein, the term "dA tailing" refers to the addition of an A base to the 3' end of a blunt-ended phosphorylated DNA fragment. This treatment generates compatible overhangs for subsequent ligation. This step can be performed using, for example, the Klenow fragment of E. coli DNA polymerase I, via methods known to those skilled in the art.
[0072] Multiple double-stranded DNA molecules, which can be used as starting materials in the methods of the present invention, can also be synthesized from single-stranded DNA or cDNA. A population of double-stranded DNA molecules can be obtained from cDNA. Double-stranded DNA can also be obtained from mRNA (e.g., from viral RNA) using methods known in the art, which include isolating mRNA, reverse transcribing RNA to produce single-stranded cDNA, and processing the single-stranded DNA to obtain double-stranded DNA.
[0073] Samples used to obtain multiple double-stranded DNA molecules can be from biological or environmental sources. Biological samples include, but are not limited to, animal or human samples, liquid and solid food and feed products (dairy, vegetables, meat, etc.). Preferred biological samples include, but are not limited to, any biological fluid, cell, tissue, organ, or part thereof containing DNA or mRNA. Biological samples may include proliferative cells, such as cells from the colon, rectum, mammary gland, ovary, prostate, kidney, lung, blood, brain, or other organs or tissues. Any organism can be used as a source, and includes, but is not limited to, bacteria, fungi, viruses, plants, animals such as humans, non-human primates, reptiles, insects, birds, worms, fish, mammals, livestock, and farm animals (cattle, horses, pigs, sheep, goats, dogs, cats, rodents, etc.). Environmental samples include, but are not limited to, surface materials, soil, water, and industrial samples, as well as samples obtained from food and dairy processing equipment. The analyzed sample may originate from a single source (e.g., a single organism, tissue, cell, etc.) or a combination of nucleic acids from multiple organisms, tissues, or cells.
[0074] Step (i)
[0075] In the first step, the method for identifying methylated cytosine in the population of double-stranded DNA molecules of the present invention involves linking a double-stranded DNA initiator to at least one end of the strands of a plurality of double-stranded DNA molecules. Preferably, the double-stranded DNA initiator can be linked to one end of the strands of the plurality of double-stranded DNA molecules. Alternatively, the double-stranded DNA initiator can be linked to both ends of the strands of the plurality of double-stranded DNA molecules.
[0076] The terms "aptamer" and "adaptor" are used interchangeably and refer to oligonucleotides or nucleic acid fragments or segments that can be linked to nucleic acid molecules of interest.
[0077] The "aptamer molecule" of the method of the present invention is a double-stranded DNA molecule having an end compatible with the end of a double-stranded DNA. The aptamer molecule can be formed from a substantially complementary first DNA strand and a second DNA strand. The aptamer molecule can be a Y-aptamer, wherein the 3' region of the first DNA strand and the 5' region of the second DNA strand form a double-stranded region through sequence complementarity, and wherein the 5' region of the first DNA strand and the 3' region of the second DNA strand are not complementary.
[0078] In one embodiment, at least one portion of the double-stranded adaptor has a sequence common to all double-stranded adaptors used in step (i). In this case, the same primers used for amplifying the complementary strand of the adaptor-modified DNA molecule in step (iv) and / or for generating the paired and transformed adaptor-modified DNA molecule in step (iii) can be used.
[0079] Optionally, the aptamer contains a unique composite barcode (also referred to as a "composite sequence," "barcode," "barcode sequence," or "composite marker") that allows for sample identification, multiplexing, pairing, and quantitative analysis. The constructs obtained by the method of this invention have barcodes that allow the generation of unique identifiers associated with the initial construct, thus providing the ability to distinguish constructs. These unique identifiers allow for the identification of specific constructs containing the identifier and its descendants. Each unique identifier is associated with a single molecule in the starting sample. Therefore, it is assumed that any amplification products of the initial single molecule carrying the unique identifier are descent-wise identical. The composite barcode also allows for the quantification of the percentage of individual sequences within the sample and can be used to monitor bias and error control during amplification steps.
[0080] Barcode sequences add the feature of "bias control." When amplification occurs, some fragments may be selectively amplified for various reasons. This undesirable effect is a major problem for quantification purposes, which are crucial in many sequencing applications, particularly for the analysis of DNA methylation status (because each allele in each cell can have a different methylation status, and even samples can have heterogeneous compositions, making quantification and bias control essential for most applications). Therefore, the presence of at least one barcode sequence allows for bias control. Since each double-stranded DNA molecule from multiple double-stranded DNA molecules can have one or more different barcode sequences, bias control can be implemented, and selective amplification of a given double-stranded or single-stranded DNA molecule can be detected.
[0081] Currently, sequencers have an error rate that requires assumptions. Most of these errors cannot be described and remain hidden in the final results. This negatively impacts subsequent processing and analysis of the results. The method of this invention provides up to four sources of information for each nucleotide (given the upper and lower strands of dsDNA and, depending on the specific case, their respective synthetic complementary strands), thereby allowing verification of each nucleotide read, since all reads must be consistent. Therefore, the method of this invention allows for the detection and even correction of errors in sequencing (both for primary sequencing and for cytosine methylation analysis).
[0082] Preferably, aptamer molecules and / or hairpin sequences and / or barcode sequences are provided as molecular libraries, wherein each member in the library can be distinguished from other members by combination sequences within the sequence, as described below.
[0083] As used herein, the terms “library of aptamer molecules and / or hairpin sequences and / or barcode sequences” and / or “combination markers” refer to a collection of aptamer molecules and / or hairpin sequences and / or barcode sequences, wherein each member of the collection is distinguishable from other members by combination sequences within the aptamer and / or hairpin sequences and / or barcode sequences.
[0084] The terms “combined sequence,” “barcode sequence,” “barcode,” and “combined barcode” are used interchangeably throughout this specification and refer to a unique identifier for a single aptamer / hairpin sequence or a separate DNA molecule (the barcode sequence itself, not belonging to the aptamer and / or hairpin sequence). Preferably, the barcode sequence is contained within the aptamer and / or hairpin sequence. In one embodiment, the combined sequence within the aptamer / hairpin sequence is a degenerate nucleic acid sequence. The combined sequence can contain any nucleotides, including adenine, guanine, thymine, cytosine, methylated cytosine, and other modified nucleotides. Preferably, the number of nucleotides in the combined sequence is designed such that the number of potential and actual sequences represented by the combined sequence is greater than the total number of aptamers in the library. The combined sequence can be located in any region of the aptamer / hairpin sequence. Preferably, it is located within a double-stranded region of the aptamer / hairpin sequence.
[0085] Optionally, the aptamer / hairpin sequence / barcode sequence is incorporated with a base labeled with a second member of the binding pair as described below, which allows for the recovery of the original DNA template after the elongation or amplification steps. This provides the advantage that the sample used as the DNA template is preserved during the process, and that the original DNA template formed by the sense and antisense strands can be recovered, stored, and subjected to multiple amplifications and sequencing under different conditions without depleting the sample. Figure 1 A schematic diagram of the method of the present invention is shown in the figure.
[0086] In the first step (i), the method for identifying methylated cytosine in the double-stranded DNA molecular population of the present invention further relates to... pair Multiple strands of double-stranded DNA molecules are used to provide multiple paired, adaptor-modified DNA molecules.
[0087] The first step of the method of the present invention, the "pairing step," can be performed by covalently coupling one or more strands of a double-stranded DNA molecule with a hairpin sequence (also known as a "hairpin molecule" or "hairpin aptamer"). The first step of the method of the present invention, the "pairing step," can be performed by using a barcode sequence. The first step of the method of the present invention, the "pairing step," can be performed by using both a hairpin sequence and a barcode sequence.
[0088] For example, the "pairing step" in the first step of the method of the present invention can be performed via... hairpin sequence Implementation. A hairpin sequence may include a hairpin loop region and a double-stranded region, wherein the double-stranded region contains ends compatible with the ends of a double-stranded DNA molecule (and / or the ends of barcode sequences, if they are already linked to the DNA strand). Therefore, the pairing step can be performed by covalently coupling one or more strands of a double-stranded DNA molecule with the hairpin sequence. The hairpin sequence may also contain one or more barcode sequences.
[0089] In this case, the following double-stranded DNA library is obtained, wherein the original sense and antisense strands of the DNA molecule are physically bound (see example...). Figure 1 ).
[0090] The terms "pairing sequence" or "pairing molecule" may be used in the context of this invention to refer to a sequence suitable for pairing the strands of one or more double-stranded DNA molecules. For example, a "pairing sequence" may refer to a hairpin sequence and / or one or more barcode sequences. As used in the methods of this invention, the term "hairpin sequence" (or "hairpin molecule" or "hairpin aptamer") refers to a double-stranded body formed from a single-stranded nucleic acid that folds itself to form a double-stranded region maintained by base pairing between complementary base sequences on the same strand. The hairpin molecule also includes a hairpin loop region formed by unpaired bases. The hairpin sequence is located at the opposite end of the double-stranded DNA molecule relative to the position of the double-stranded DNA adaptor in the double-stranded DNA molecule.
[0091] Optionally, the "pairing step" of the first step (step (i)) of the method of the present invention can be performed via... Use barcodes (Also known as "barcode sequence", "combined sequence" and / or "combined barcode" and / or "barcode" and / or "combined mark", as described above) This is achieved. Therefore, the pairing step can be performed using a barcode sequence.
[0092] The “combination sequence,” “combination barcode,” “barcode sequence,” or “barcode” used to pair multiple double-stranded DNA molecules can be located in the adaptor and / or, if present, in the hairpin sequence. The barcode can be a separate double-stranded DNA molecule, which can be attached to one or both ends of the double-stranded DNA molecule. For example, the barcode can be attached to one or both ends of the double-stranded DNA molecule before the adaptor and / or hairpin sequence is attached to it (and in this case, the adaptor and / or hairpin sequence can be attached to the barcode). For example, the barcode can be attached to one or both ends of the double-stranded DNA molecule after the adaptor is attached to it, and can thereby be attached to the adaptor.
[0093] Alternatively, a pairing step can be performed using both hairpin sequences and barcodes.
[0094] The adaptor, hairpin sequence, and / or barcode sequence may contain unmethylated cytosine and / or methylated cytosine. The adaptor, hairpin sequence, and / or barcode sequence may not contain unmethylated cytosine. For example, the adaptor, hairpin sequence, and / or barcode sequence may not contain cytosine. For example, the adaptor, hairpin sequence, and / or barcode sequence may contain methylated cytosine, but they do not contain unmethylated cytosine. For example, the adaptor, hairpin sequence, and / or barcode sequence may contain methylated cytosine and unmethylated cytosine.
[0095] If the adapter contains unmethylated cytosine, these unmethylated cytosines (in step (ii) in any embodiment of the method of the invention) are treated equally with the following reagent and thus equally converted to bases (preferably uracil) that are detectably different from cytosine in terms of hybridization characteristics, said reagent allowing the conversion of unmethylated cytosine to bases (preferably uracil) that are detectably different from cytosine in terms of hybridization characteristics. Therefore, the primers used in step (iii) (and optionally in step (iv)) should contain a sequence complementary to at least a portion of the double-stranded adapter after such conversion.
[0096] As used in this article, the term "terminus" refers to a region of the sequence at or near any end of a nucleic acid sequence.
[0097] As used herein, the term "compatible" means that the two strands of one of the ends of an aptamer molecule can be ligated to one or both ends of a double-stranded DNA molecule used as starting material. Compatible ends include blunt DNA ends and sticky ends with complementary overhanging ends. Two compatible ends can preferably be ligated together without any gaps or mismatches and can be ligated to produce a DNA sequence that often contains restriction endonuclease sites.
[0098] As used in this article, the term "blunt end" means that the two strands of double-stranded DNA are of the same length and end with a base pair (i.e., there are no unpaired bases and the strands do not overlap or dangle from each other).
[0099] The terms “sticky end,” “adhesive terminator,” and “overhang” are used interchangeably herein and refer to non-flat ends formed by multiple overhangs. An overhang is a segment of unpaired nucleotide at the end of a DNA molecule. These unpaired nucleotides can be in either strand, forming 3' or 5' overhangs. These overhangs are palindromic in most cases. The simplest case of an overhang is a single nucleotide. The single nucleotide is most commonly adenosine and is formed as a 3' overhang by some DNA polymerases. The product is ligated to a linear DNA molecule having a 3' thymine overhang. This facilitates the ligation of the two molecules by ligases because adenine and thymine form a base pair. In the first method of the invention, when the double-stranded DNA molecule used in step (i) is end-repaired and dA-tailed prior to step (i), the first and second aptamer molecules must have 3'-thymine overhangs. Longer overhangs are most often formed by restriction endonucleases. For example, restriction endonucleases can cleave two DNA strands four base pairs apart, creating a 4-base 5' overhang in one molecule and a complementary 5' overhang in the other. These ends are called sticky ends because they are easily ligated back together by ligases. Since different restriction endonucleases typically produce different overhangs, a DNA molecule can be cleaved with two different enzymes and then ligated to another DNA molecule with ends produced by the same enzyme. Because the overhangs must be complementary for the ligase to work, the two molecules can only be ligated in one direction.
[0100] The ligation step (i) is performed under conditions suitable for the ligation of adaptors and / or pairing molecules (hairpins and / or barcode sequences) to DNA molecules to produce multiple aptamer-modified DNA molecules.
[0101] As used herein, the term "ligation" refers to the formation of a covalent bond or link between the ends of two or more nucleic acids. The nature of the bond or link can vary widely, and ligation can be performed enzymatically or chemically. As used herein, ligation is typically performed enzymatically to form a phosphodiester link between the 5' terminal nucleotide of one nucleic acid and the 3' carbon of another nucleic acid. Suitable conditions for ligation are any conditions that allow for the production of a double-stranded DNA molecule ligated with one or two aptamers. Preferred conditions are those using DNA ligase, although procedures for ligation without the use of DNA ligase are also known.
[0102] Pairing can be performed before or after the connection, or simultaneously with the connection of the connector and / or hairpin sequence. Preferably, pairing is performed simultaneously with the connection of the connector and / or hairpin sequence.
[0103] The result of the first step (step (i)) of the method of the present invention is multiple DNA molecules.
[0104] In the context of the method of this invention, If the hairpin sequence enables the matching of step (i) of the method of the present invention, For steps (For example, by the presence of the hairpin sequence itself or by the presence of the hairpin sequence and one or more barcode sequences), then the DNA molecule obtained in step (i) can be (see Figure 1 ):
[0105] A) A double-stranded DNA molecule linked at one end to an aptamer molecule (optionally containing at least one barcode sequence) and at the other end to a second molecule, wherein the second molecule is a hairpin sequence (and optionally contains one or more barcode sequences) (these are so-called paired aptamer-modified DNA molecules);
[0106] B) A double-stranded DNA molecule linked to two aptamer molecules in the absence of hairpin sequences (at least one at each end);
[0107] C) A double-stranded DNA molecule linked to two molecules (at least one at each end), said molecules may be hairpin sequences and / or barcode sequences; and
[0108] D) Double-stranded DNA molecules that are not linked to any other molecules (original double-stranded DNA molecules, i.e., unmodified double-stranded DNA molecules).
[0109] In the context of the method of this invention, If the presence of one or more barcode sequences and the absence of hairpins In the case of a sequence, the pairing step of step (i) of the method of the present invention is implemented. Then the DNA molecule obtained in step (i) can be:
[0110] A) In double-stranded DNA molecules Each end A double-stranded DNA molecule (including at least one barcode sequence) linked to an aptamer molecule (these are so-called paired aptamer-modified DNA molecules);
[0111] B) At the ends of double-stranded DNA molecules Only one in A double-stranded DNA molecule linked to at least one aptamer molecule;
[0112] C) If this is the case, then it is linked to at least one barcode sequence, but In the absence of connective molecules In the case of a double-stranded DNA molecule being linked to each end of a double-stranded DNA molecule; and
[0113] D) Double-stranded DNA molecules that are not linked to any other molecules (original double-stranded DNA molecules, i.e., unmodified double-stranded DNA molecules).
[0114] If this is achieved through the presence of (one or more) barcode sequences and in the absence of hairpin sequences... Pairing steps of the method (i) Then, the paired aptamer-modified DNA molecules obtained after step (i) should be placed in a double-stranded DNA molecule (containing at least one barcode sequence for pairing of the two strands). Each end It is linked to an aptamer molecule.
[0115] A population of double-stranded DNA molecules can be pretreated with aptamers in step (i) under conditions suitable for ligation of aptamers to DNA molecules, thereby introducing sticky ends into the DNA molecules. Adaptamers with sticky ends can be obtained by digesting double-stranded DNA with a suitable restriction endonuclease or synthesized (e.g., by annealing single-stranded oligonucleotides).
[0116] Following the ligation of aptamers and / or pairing molecules (e.g., hairpin sequences and / or barcode sequences) in step (i), optionally, the following can be performed: Capture Steps (or “recovery step”, or “purification step”) to recover from the reaction mixture those molecules containing aptamers and / or hairpin sequences and / or barcode sequences (depending on the pairing molecules, as described above) according to section (A) above, i.e.: if pairing is achieved by the presence of at least a hairpin molecule, then for a double-stranded DNA molecule linked at one end to an aptamer molecule (optionally containing at least one barcode sequence) and at the other end to a second molecule, wherein the second molecule is a hairpin sequence (and optionally contains one or more barcode sequences); and if pairing is achieved in the absence of a hairpin sequence (by at least one barcode sequence only), then for a double-stranded DNA molecule (containing at least one barcode sequence) Each end A double-stranded DNA molecule linked to an aptamer molecule (“recovery step”). Therefore, the first step (i) of the method of the present invention may further involve recovering from the population of DNA molecules obtained in step (i) those molecules as aptamer-modified DNA molecules (molecules according to (A) above), said aptamer-modified DNA molecules containing an aptamer and / or hairpin sequence and / or barcode sequence at one or both ends of paired aptamer-modified DNA molecules.
[0117] The steps described allow for the separation of the paired, aptamer-modified DNA molecule (according to (A) above) obtained in step (i), containing aptamers and / or hairpin sequences and / or barcode sequences, from the remaining portion of the resulting DNA molecule (e.g., according to (B)-D) above) (not according to (A) above). The capture step can be performed, for example, by relying on probes or ligands that have affinity for the hairpin sequences and / or barcode sequences but not for the aptamer molecules, or by relying on probes or ligands that have affinity only for the aptamer molecules.
[0118] This provides the following advantages: the sample used as a DNA template is preserved during the process, and the original DNA template formed by the sense and antisense strands can be recovered, stored, and amplified and sequenced multiple times under different conditions without depleting the sample (steps iii and / or iv above). Figure 1 The diagram is shown in the image.
[0119] Preferably, a recovery step (which may also be referred to as a “capture step”, “purification step”, or “separation step”) is performed using a polynucleotide, the polynucleotide comprising a sequence and a purification tag that are complementary to at least a portion of an aptamer sequence and / or a hairpin sequence and / or a barcode sequence.
[0120] As used herein, the term "polynucleotide" refers to a single-stranded DNA or RNA molecule comprising a number of covalently bonded nucleotide monomers. Preferably, the polynucleotide has eight or more nucleotide monomers. In a preferred embodiment, the polynucleotide is a single-stranded DNA molecule having a length of at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 18, at least 20, at least 25, at least 30, at least 35, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100 or more nucleotides.
[0121] As used herein, the term "purification tag" refers to a component that enables the separation of the polynucleotide and the target sequence. Preferably, the DNA backbone of the polynucleotide contains one or more nucleotides coupled to the affinity purification tag. Preferably, the affinity purification tag can be a member of a binding pair. More preferably, the affinity purification tag is biotin, and the double-stranded DNA molecule is separated by affinity purification with avidin or streptavidin. These steps can be performed using magnetic beads, but are not limited thereto.
[0122] Prior to step (ii), the multiple paired, adaptor-modified DNA molecules generated in step (i) (as defined in (A) above, depending on the paired molecules in both cases) can be separated from the DNA molecules generated in step (i) according to (B)-D) as defined above, to generate a library of paired, adaptor-modified DNA molecules according to (A) above. Thus, the method of the present invention allows for the acquisition of double-stranded DNA libraries in which the original sense and antisense strands of the DNA molecules are paired.
[0123] As used herein, the term “DNA library” refers to a collection of DNA fragments that have been ligated with aptamer molecules to identify and isolate DNA fragments of interest.
[0124] In the context of this invention, the term "double-stranded DNA library" can refer to a library containing two strands (i.e., the sense strand and the antisense strand) of a DNA molecule physically linked by one of its ends (e.g., by a hairpin sequence) and forming part of the same molecule. A double-stranded DNA library is not a circular library. The original strands of the DNA molecule can be physically linked by a loop through one of its ends, thus forming a double strand between the sense and antisense strands (according to A above). Each molecule in a double-stranded DNA library can also be linear in conformation when the complementarity between the sense and antisense strands of the DNA molecule is partially or completely lost. Alternatively, in the context of the method of this invention, the term "double-stranded DNA library" can refer to a library in which the two strands of the DNA molecule are not physically linked by one of its ends, but are paired using, for example, at least one barcode sequence (according to A above).
[0125] The connection step in step (i) of the method of the present invention can be referred to as the "contact step".
[0126] Step (ii)
[0127] DNA methylation typically occurs at CpG sites (cytosine-phosphate-guanine sites, where cytosine is immediately followed by guanine in the DNA sequence). This methylation leads to the conversion of cytosine to 5-methylcytosine. The formation of Me(methyl)-CpG is catalyzed by the enzyme DNA methyltransferase. Approximately 80–90% of CpG sites in human DNA are methylated, but there are certain regions called CpG islands, which are rich in GC (composed of approximately 65% of CG residues), none of which are methylated. These are associated with the promoters of 56% of mammalian genes, including all widely expressed genes. 1% to 2% of the human genome consists of CpG clusters, and there is an inverse correlation between CpG methylation and transcriptional activity.
[0128] Methylation patterns are important in the study of some diseases. In normal tissues, gene methylation is predominantly located in coding regions, where CpG is depleted; while promoter regions are unmethylated, despite high density of CpG islands. However, in cancer, methylation imbalances exist, with genome-wide hypomethylation accompanied by localized hypermethylation and increased expression of DNA methyltransferases. The methylation status of some genes can be used as biomarkers for tumorigenesis.
[0129] In the second step (step (ii)), the paired, aptamer-modified DNA molecule population generated after step (i) is treated with a reagent that allows the conversion of unmethylated cytosine to a base (preferably uracil) that is detectably different from cytosine in terms of hybridization properties. Preferably, the primers used in steps (iii) (and optionally (iv)) are specific to the aptamer molecules that have been treated with the reagent. Figure 1 The schematic diagram illustrates this embodiment of the method of the present invention.
[0130] As used herein, the phrase "a base that is detectably different from cytosine in terms of hybridization characteristics" refers to a base that cannot hybridize with guanine in the complementary strand. Preferably, the base that is detectably different from cytosine is thymine or uracil, more preferably uracil.
[0131] The reagent used in this step must be capable of converting unmethylated cytosine into a base that is detectably different from cytosine in terms of hybridization characteristics, but cannot act on methylated cytosine. Examples of such reagents are, but are not limited to, bisulfites, metasulfites, and cytidine deaminases, such as activation-induced cytidine deaminase (AID). In a preferred embodiment, the reagent is a bisulfite. As used herein, the bisulfite ion has its usual meaning HSO3. -Typically, bisulfites are used as aqueous solutions of bisulfites, such as sodium bisulfite having the formula NaHSO3, or magnesium bisulfite having the formula Mg(HSO3)2. Suitable counterions for bisulfite compounds can be monovalent or divalent. Examples of monovalent cations include, but are not limited to, sodium, lithium, potassium, ammonium, and tetraalkylammonium. Suitable divalent cations include, but are not limited to, magnesium, manganese, and calcium. Treatment of DNA with bisulfite converts unmethylated cytosine bases to uracil, but leaves 5-methylcytosine bases unaffected. This conversion is performed using standard procedures (Frommer et al. 1992, Proc Natl Acad SciUSA, 89:1827-31; Olek, 1996, Nucleic Acid Res. 24:5064-6; EP 1394172). Methods for obtaining samples include those used for reduced-presentation bisulfite sequencing (RRBS).
[0132] Preferably, bisulfite is used to convert (unmethylated) cytosine to uracil in paired DNA molecules.
[0133] When the paired, aptamer-modified DNA molecule obtained in step (ii) is treated with a reagent capable of converting unmethylated cytosine into a base (preferably uracil, as described above) that is detectably different from cytosine in terms of hybridization properties, the complementarity between the sense and antisense strands of the original double-stranded DNA molecule is partially or completely lost. This promotes the annealing of primers used in subsequent steps. In particular, if one strand has unmethylated cytosine and no methylated cytosine, this also promotes the generation of the complementary strand in step (iii).
[0134] When nucleotides in one of the regions containing the sense or antisense strand pair with less than 100%, less than 99%, less than 95%, less than 90%, less than 80%, less than 70%, less than 60%, less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, less than 5%, less than 3%, less than 1%, less than 0.5%, or less than 0.1% of the aptamer-modified DNA molecule, the paired aptamer-modified DNA molecule is considered to have partially lost the complementarity between the sense and antisense strands of the original double-stranded DNA molecule. When nucleotides in one of the regions containing the sense or antisense strand pair with 0% of the nucleotides in the other region, complementarity is considered to have been completely lost.
[0135] In the specific case where the original double-stranded DNA molecule is fully methylated, the complementarity between the sense and antisense strands of the original double-stranded DNA molecule is not lost after treatment with a reagent capable of converting unmethylated cytosine to a base that is detectably different from cytosine in terms of hybridization properties (e.g., uracil). In the specific case where one strand has (unmethylated) cytosine and no methylated cytosine, optimally, complementarity is lost.
[0136] Step (iii)
[0137] In the third step, the paired and transformed aptamer-modified DNA molecule obtained in step (ii) is used as a template, under conditions allowing for strand synthesis, and a DNA strand is synthesized using primers whose sequences are complementary to at least a portion of the sequence of the first aptamer molecule. Thus, step (iii) of the method of the present invention uses nucleotides A, G, C, and T and primers to provide a complementary strand of a paired and transformed adaptor-modified DNA molecule, the sequences of which are complementary to at least a portion of a double-stranded adaptor to provide a partially transformed paired double-stranded molecule.
[0138] After treatment with, for example, bisulfite, the transformed and paired aptamer-modified DNA molecule obtained in step (ii) is used as a template, and a DNA strand is synthesized using primers whose sequences are complementary to at least a portion of the sequence of the aptamer molecule or at least a portion of the complementary sequence of the aptamer molecule (preferably, the primers used in step (iii) (and optionally (iv)) are specific to the aptamer molecule treated with the reagent used in step (ii) or its complementary sequence, as previously described), and optionally, the obtained product can be amplified. Uracil is recognized as thymine by Taq polymerase, and after the elongation (and optionally amplification step), the obtained product contains thymine at the sites of unmethylated cytosine in the DNA template and cytosine at the sites of 5-methylcytosine in the DNA template.
[0139] The term "synthetic DNA strand" refers to a DNA molecule synthesized under conditions that allow for strand synthesis, which is complementary to an aptamer-modified DNA molecule used as a template.
[0140] The term "template" refers to the DNA strand that sets the gene sequence for a new strand.
[0141] The phrase "under conditions that allow for strand synthesis" refers to conditions in which the hydrogen bonds between complementary bases in the regions containing the sense and antisense strands of the double-stranded DNA molecule used in step (i) are broken. These conditions are suitable for separating the regions containing the sense and antisense strands of the original double-stranded DNA molecule, and include, but are not limited to, conditions that allow for the use of a linearized form of the aptamer-modified DNA molecule obtained after step (i) in the case of pairing obtained by hairpin molecules, or conditions that allow for the use of isothermal techniques, such as strand displacement DNA polymerase.
[0142] Suitable conditions for separating the regions include denaturation of the two regions, for example, by heating the molecules to 94-98°C for 20 seconds to 2 minutes, causing the hydrogen bonds between complementary bases to break. Separation of the regions can also be achieved without heating the molecules using isothermal techniques, for example, by using strand-displacement DNA polymerases, such as, but not limited to, Phi29 DNA polymerase or large fragments of Bacillus stearothermophilus DNA polymerase.
[0143] Furthermore, when the paired, aptamer-modified DNA molecule obtained in step (i) is treated with a reagent capable of converting unmethylated cytosine into a base (preferably uracil) that is detectably different from cytosine in terms of hybridization properties, the complementarity between the sense and antisense strands of the original double-stranded DNA molecule is partially or completely lost. In particular, if one strand has unmethylated cytosine and no methylated cytosine, this can promote the synthesis of the complementary strand.
[0144] As used herein, the term "primer" refers to a short nucleic acid strand that is complementary to a sequence in another nucleic acid and serves as the starting point for DNA synthesis. Preferably, primers have at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 18, at least 20, at least 25, at least 30 or more bases in length.
[0145] The term "complementary" refers to the base pairing that allows for the formation of a duplex between nucleotides or nucleic acids, such as between the two strands of a double-stranded DNA molecule, between an oligonucleotide primer and a primer-binding site on a single-stranded nucleic acid, or between an oligonucleotide probe and its complementary sequence in a DNA molecule. Complementary nucleotides are generally A and T (or A and U), or C and G. Two single-stranded DNA molecules are considered substantially complementary when the nucleotides of one strand pair with approximately 60%, at least 70%, at least 80%, at least 85%, typically at least about 90% to about 95%, and even about 98% to about 100% of the nucleotides of the other strand in optimal alignment and comparison, and with appropriate nucleotide insertions or deletions. The degree of identity between two nucleotide regions is determined using algorithms executed in a computer and methods known to those skilled in the art. Preferably, the identity between two nucleotide sequences is determined using the BLASTN algorithm (BLASTN et al., Altschul, S. et al., NCBI NLM NIH Bethesda, Md. 20894, Altschul, S. et al., J., 1990, Mol. Biol. 215: 403-410).
[0146] Primers can hybridize with the sequence of the aptamer molecule (and preferably with the sequence obtained therefrom after treatment with the reagent in step (ii), preferably bisulfite) under low-tightness conditions, preferably medium-tightness conditions, and most preferably high-tightness conditions. The primers used in step (iii) and (if so) in step (iv) are specific to the aptamer molecule that has been treated with a reagent (e.g., bisulfite) that allows the conversion of unmethylated cytosine to bases that are detectably different from cytosine in terms of hybridization characteristics, as previously described.
[0147] "Hybridization" refers to the process by which two single-stranded polynucleotides non-covalently bind to form a stable double-stranded polynucleotide. "Hybridization conditions" typically include a salt concentration of about 1 M or less, more typically less than about 500 mM, and may be less than about 200 mM. "Hybridization buffer" is a buffered salt solution, such as 5% SSPE, or other such buffers known in the art. Hybridization temperatures can be as low as 5 °C, but are typically above 22 °C, and more typically above about 30 °C, and often exceed 37 °C. Hybridization is often performed under stringent conditions, i.e., conditions under which primers hybridize with their target sequence but not with other non-complementary sequences. Stringent conditions are sequence-dependent and vary under different conditions. For example, longer fragments may require higher hybridization temperatures for specific hybridization than shorter fragments. Because other factors can affect the stringency of hybridization, including the base composition and length of the complementary strand, the presence of organic solvents, and the degree of base mismatch, the combination of parameters is more important than the absolute measure of any single parameter. Generally, stringent conditions are chosen to be about 5 °C lower than the Tm of a particular sequence at specified ionic strengths and pH. Exemplary stringent conditions include a salt concentration of at least 0.01 M to no more than 1 M sodium ion concentration (or other salt) at a pH of about 7.0 to about 8.3 and a temperature of at least 25°C.
[0148] Therefore, transformed paired aptamer-modified DNA molecules are converted into double-stranded DNA molecules (one or, depending on the specific case, two separate single-stranded molecules, depending on the type of pairing (physical or via barcode sequence)) by methods known in the art (i.e., using DNA polymerase and dNTP extension / stretching) to provide partially transformed (the original strand has been affected by treatment with a reagent that allows the conversion of unmethylated cytosine to bases that are detectably different from cytosine in terms of hybridization properties (preferably uracil), but is not newly generated) paired double-stranded DNA molecules.
[0149] The paired and transformed aptamer-modified DNA molecules recovered in step (ii) may also include double-stranded DNA molecules linked to two hairpin molecules and / or two barcode sequences (e.g., if no recovery or purification step occurs). However, said molecules are not converted into double-stranded DNA molecules because the hairpin molecules / barcode sequences do not contain the target sequences of the primers used in step (iii). Figure 2 The method of the present invention is shown in the extension step, wherein only the ligation product having an aptamer attached to one end of a double-stranded DNA molecule is extended and amplified.
[0150] The constructs obtained after step (ii), or after step (iii), or after step (iv) form the double-stranded DNA library of the present invention and can be used for sequencing or in other conventional molecular biology techniques.
[0151] Step (iii) can also be called the "extension step".
[0152] Step (iv) (Also known as the "amplification step")
[0153] Optionally, the construct can be amplified to increase the amount of material used in the following steps. In a preferred embodiment, the double-stranded DNA molecule obtained in step (iii) is amplified using primers whose sequences are complementary to at least a portion of the aptamer region (the primers used in (iv) are specific to the aptamer molecules that have been treated with the reagents of step (ii), as described above).
[0154] Therefore, in the optional fourth step, the method of the present invention includes amplifying the partially transformed paired double-stranded DNA molecules obtained in step (iii) to provide amplified paired double-stranded DNA molecules.
[0155] DNA amplification allows the generation of multiple copies of the molecule by synthesizing a double-stranded DNA molecule in vitro. Any method for DNA amplification can be used. Preferably, it is performed by polymerase chain reaction. In another embodiment, it can be performed by real-time PCR using different probes (e.g., LightCycler, Taqman, Escorpio, Sunrise, Molecular Beacon, or Eclipse). Different amplification conditions can be used on aliquots of the same sample to overcome any possible bias.
[0156] Double-stranded DNA molecules obtained by the method of the present invention can be recovered from the reaction mixture (“recovery step” or “purification step”). Therefore, preferably, the DNA molecules obtained in step (iii) or, depending on the specific circumstances, step (iv) are recovered from the reaction mixture. More preferably, the recovery is carried out using the first member of the binding pair, wherein the primers used in step (iii) or, depending on the specific circumstances, step (iv) are modified with the second member of the binding pair.
[0157] Optionally, multiple original double-stranded DNA molecules can be recovered. The paired aptamer-modified DNA molecules obtained in step (i) serve as the original templates for the extension and amplification steps. The original templates are not destroyed during processing and can be preserved and reused or stored for subsequent processes. To achieve this, the original templates can be tagged with modified aptamer / hairpin / barcode sequences. Therefore, in a preferred embodiment, the paired aptamer-modified DNA molecules obtained in step (i) are recovered from the reaction mixture obtained after step (iii) or, depending on the specific circumstances, after step (iv). In a more preferred embodiment, the recovery is carried out using a first member of the binding pair, wherein the aptamer and / or hairpin and / or barcode sequence is modified with a second member of the binding pair.
[0158] As used herein, the term "reaction mixture" refers to the mixture obtained after steps (iii) and / or (iv). The reaction mixture is formed by a combination of reagents, paired aptamer-modified DNA molecules used as templates, non-reactive paired aptamer-modified DNA molecules, and reaction products, including molecules that form a double-stranded DNA library.
[0159] As used herein, the term "binding pair" refers to a pair formed by a first member and a second member, and includes any class of immune-type binding pairs, such as antigen / antibody (digoxigenin / anti-digoxigenin antibody) or hapten / anti-hapten systems; and also includes any class of non-immune-type binding pairs, including systems in which two components share natural affinity for each other but are not antibodies, such as biotin / avidin, biotin / streptavidin, folic acid / folate-binding protein, complementary nucleic acid segments, protein A or G / immunoglobulin; and covalently binding pairs that form covalent bonds with each other, such as thiol-reactive groups, including maleimide and haloacetyl derivatives, and amine-reactive groups, such as isothiocyanates, succinimide esters, and sulfonyl halides.
[0160] The sequences of primers used in step (iii) or step (iv), or the sequences of aptamers and / or hairpin sequences and / or barcode sequences, can be designed to incorporate bases labeled with a second member of the binding pair (e.g., digoxigenin, biotin, etc.). The incorporated labeled bases can be used to complex them with a first member of the binding pair (optionally with a support).
[0161] The primer pairs used in step (iii) and, depending on the case, in step (iv) are specific to the aptamer molecule after it has been treated with a reagent (e.g., bisulfite) that allows the conversion of unmethylated cytosine to a base that is detectably different from cytosine in terms of hybridization characteristics. As used herein, the term "specific" means that the primers can hybridize with the aptamer molecule only when it has been treated with a reagent that converts unmethylated cytosine to a base (e.g., uracil) that is detectably different from cytosine in terms of hybridization characteristics. Preferably, the primers cannot hybridize with the aptamer molecule before the conversion. If the aptamer molecule contains unmethylated cytosine, the primers used in steps (iii) and / or (iv) have an adenine base instead of a guanine base at the position where they pair with the unmethylated cytosine of the original aptamer molecule. The aptamer molecule may contain methylated or unmethylated cytosine. Optionally, to prevent specific portions of the sequences of the aptamer molecule, hairpin sequence, and / or barcode sequence from changing after treatment with the reagent, the sequence of the first aptamer molecule and the combined sequence within the preferred aptamer (if so) sequence may contain modified cytosine resistant to treatment with the reagent, which allows the conversion of unmethylated cytosine into bases that are detectably different from cytosine in terms of hybridization properties.
[0162] As used herein, the term "modified cytosine" refers to a cytosine base that has been obtained by substituting or adding one or more atoms or groups to obtain a modified cytosine base that cannot be converted to a base that is detectably different from cytosine in terms of hybridization properties by treatment with a reagent that converts unmethylated cytosine to a base that is detectably different from cytosine. Examples of modified cytosines suitable for aptamer sequences and preferably suitable for combinations of aptamer, hairpin, and / or barcode sequences of the present invention are, but not limited to, methylcytosine and 5-hydroxymethylcytosine. The modified cytosines are resistant to reagent treatment because they remain unchanged after treatment (e.g., methylcytosine) or because after bisulfite treatment they are converted to a base complementary to guanine and read as cytosine in polymerase base amplification and sequencing (e.g., 5-hydroxymethylcytosine, which is converted to methyl cytosine-5-sulfonate).
[0163] Subsequently, the paired DNA molecules obtained in step (iii) and / or step (iv) (and / or, optionally and less preferably, the paired DNA molecules obtained in step (ii)) are sequenced (see the section "Sequencing Steps" below in the specification).
[0164] If cytosine is present in one strand of the paired double-stranded DNA molecule obtained in step (iii) and / or step (iv) and guanine is present at the corresponding position in the other strand of the paired double-stranded DNA molecule, the presence of methylated cytosine at the given position is determined; or if uracil or thymine is present in one strand of the paired double-stranded DNA molecule obtained in step (iii) or step (iv) and guanine is present at the corresponding position in the other strand of the paired double-stranded DNA molecule, the presence of unmethylated cytosine at the given position is determined. This is further described in the section “Identification of Methylated Cytosine” below.
[0165] First embodiment of the method of the present invention
[0166] The methods of the present invention can be performed on a solid support. Specifically, double-stranded linkers, hairpin sequences, and / or barcode sequences can be provided immobilized in the support. When using a solid support, a high level of automation of the sequencer and a high degree of sample preservation (sample permanence) are expected.
[0167] Preferably, fixation can be achieved by attaching the end of one of the strands of the double-stranded adaptor, hairpin sequence, and / or barcode sequence to the support. Preferably, the end of one of the strands of the double-stranded adaptor, hairpin sequence, and / or barcode sequence is attached to the support. Therefore, when the double-stranded adaptor, hairpin sequence, and / or barcode sequence is fixed in the support, the sense of ligation can be guided, thereby ligating a specific aptamer / hairpin sequence or barcode sequence to a specific end of the DNA molecule.
[0168] If double-stranded aptamer molecules, hairpin sequences, and / or barcode sequences are immobilized in the support, the steps for recovering the original DNA template as described above are not required because the paired aptamer-modified DNA molecules obtained in step (i) remain attached to the support.
[0169] Furthermore, if aptamer molecules are immobilized in the support, then when pairing occurs through the presence of hairpin molecules, all paired aptamer-modified DNA molecules obtained in step (i) are double-stranded DNA molecules linked to an aptamer molecule at one end and to a hairpin sequence and optionally a barcode sequence at the other end. If pairing occurs in the absence of hairpin sequences, i.e., solely through barcode sequences, then the paired aptamer-modified DNA molecules obtained in step (i) are double-stranded DNA molecules linked to aptamer molecules at both ends, which additionally contain at least one barcode sequence. Therefore, it is not necessary to recover molecules containing an adaptor sequence at one end and a hairpin sequence and / or barcode sequence and / or adaptor sequence (as described above) from the molecular population obtained in step (i). Additionally, in this embodiment of the invention, where double-stranded adaptors, hairpin sequences, and / or barcode sequences are provided immobilized on the support, pairing in step (i) can be performed by distributing the double-stranded adaptors to predetermined positions on the support.
[0170] As used herein, the term "support" refers to any material configured to chemically bond with nucleic acids, including but not limited to plastics, latex, glass, metals (i.e., magnetized metals), nylon, cellulose nitrate, quartz, silicon, or ceramic articles. Preferably, the support is solid and may be generally spherical (i.e., beads) or may comprise standard laboratory containers such as microplates or surfaces.
[0171] As used herein, the term “fixation” refers to the association or binding between a molecule (e.g., an aptamer, a hairpin sequence, a barcode sequence) and a support in a manner that provides stable association under conditions of elongation, amplification, cleavage, and other processes as described herein. Such binding can be covalent or non-covalent. Non-covalent binding includes electrostatic, hydrophilic, and hydrophobic interactions. Covalent binding is the formation of covalent bonds characterized by shared electron pairs between atoms. Such covalent binding can be formed directly between the aptamer and the support, or through crosslinking, or by including specific reactive groups on the support or the aptamer, or both. Covalent attachment of the aptamer can be achieved using binding partners fixed to the support, such as avidin or streptavidin, and biotinylated aptamers to the non-covalent binding of avidin or streptavidin. Fixation can also involve a combination of covalent and non-covalent interactions.
[0172] Aptamers, hairpin sequences, and / or barcode sequences can be synthesized first and then attached to a support. Alternatively, aptamers, hairpin sequences, and / or barcode sequences can be synthesized directly on the support. Preferably, the immobilization is performed by covalently binding nucleotides at the end of one of the chains of the aptamer, hairpin sequence, and / or barcode sequence or the hairpin loop to the support.
[0173] Preferably, aptamers are attached to the support, but hairpin sequences and / or barcode sequences are not attached to the support. In the context of this embodiment, each aptamer molecule attached to the support is sufficiently separated from neighboring aptamer molecules to prevent a double-stranded DNA molecule from being linked to two of these aptamers. In step (i) of the method according to the invention, in one embodiment, multiple double-stranded DNA molecules are contacted with aptamer molecules attached to the support, and only one aptamer is able to link to each double-stranded DNA molecule. Thereafter, adaptors and / or hairpin sequences and / or barcode sequences can be linked to the free ends of the double-stranded DNA molecules.
[0174] Alternatively, hairpin sequences and / or barcode sequences can be attached to the support, but aptamer molecules cannot. In this context of this embodiment, each hairpin sequence and / or barcode sequence attached to the support is sufficiently separated from neighboring hairpin sequences and / or barcode sequences to prevent a double-stranded DNA molecule from being linked to two of these molecules. For example, multiple original double-stranded DNA molecules can be contacted with hairpin sequences and / or barcode sequences attached to the support, and only one hairpin sequence and / or barcode sequence can be linked to each double-stranded DNA molecule. The aptamer molecule is then ligated to the free end of the double-stranded DNA molecule (if the hairpin is attached to the support and ligated to the other end of the double-stranded DNA molecule). If pairing is achieved solely by barcode sequences (in the absence of hairpin sequences) and the barcode sequences are attached to the support, the adaptor should first be ligated to the attached barcode sequence. Then, the double-stranded DNA sequence should be ligated to the aptamer. Finally, another adaptor molecule is ligated to the free end of the double-stranded DNA molecule.
[0175] Alternatively, aptamer molecules and hairpin and / or barcode sequences can be attached to a support. Therefore, additional measures must be taken to avoid the double-stranded DNA molecule being linked to two identical molecules that are not two identical adaptors in cases where pairing is achieved solely by the presence of a barcode sequence (in the absence of a hairpin sequence). In such cases, the double-stranded DNA molecule is linked to two adaptors, with one adaptor attached to each end of the double-stranded DNA molecule.
[0176] The advantages of carrying out the method of the present invention on a solid support (e.g., by fixing connectors and / or hairpin sequences and / or barcode sequences in the solid support, as described above) are as follows:
[0177] - The direction of the connections can be controlled, thus preventing the formation of unwanted connections (e.g., molecules with two hairpin sequences). Furthermore, since the original multiple double-stranded DNA molecules are fixed to a solid support, there is no loss of the original material.
[0178] - Additionally, reactions can be performed in flow cells. After reactions (e.g., ligation, transformation, generation of complementary strands) have occurred, the original material can be stored and reused (because it is attached to the support). Flow cells can be integrated into NGS (next-generation sequencing) devices (and specific reactions, such as bridging amplification or sequencing reactions, can be performed on the flow cell itself), thereby allowing for automation and simplification of the process.
[0179] After the ligation of the double-stranded adaptor and the pairing of the strands of multiple double-stranded DNA molecules (which are at the end of the process of immobilization in the support, as described above), the (unmethylated) cytosine present in the two strands of the paired adaptor-modified DNA molecules is converted into bases (preferably uracil) in the paired adaptor-modified DNA molecules that are detectable to a different degree than cytosine, as described above (step ii).
[0180] Subsequently, nucleotides A, C, G, and T, and primers are used to provide complementary strands of paired and transformed aptamer-modified DNA molecules to provide partially transformed paired double-stranded molecules (step iii), wherein the sequence of the primers is complementary to at least a portion of the double-stranded linker (preferably, the primers are specific to the aptamer molecules that have been treated with the reagents of step (ii), as described above).
[0181] This step (step iii) of providing the complementary strand in this embodiment of the invention can be performed using paired and transformed aptamer-modified DNA molecules attached to a support, and under conditions that allow strand synthesis to occur.
[0182] Preferably, the primers used in step iii of this embodiment are not attached to the support. Figure 3 and 4 The diagram illustrates this embodiment, showing a population of aptamer molecules immobilized in a support. The template strand remains attached to the support, and the elongated product is released into the supernatant. Thus, in this embodiment, a double-stranded DNA library is released into the supernatant. Optionally, step (iv) (optionally, amplifying partially transformed paired double-stranded molecules) can be performed on the molecules in the supernatant before or after recovery from the reaction mixture. Such aptamers or aptamer-modified DNA molecules attached to the support can be released from the support at various stages of the method.
[0183] Optionally, the primers used in step (iii) may also be attached to the support. In this case, the template and the extension product remain attached to the support. Therefore, in this case, the double-stranded DNA library is attached to the support. The aptamers, primers, and / or aptamer-modified DNA molecules attached to the support may be released from the support at various stages of this embodiment of the method.
[0184] Optionally, step (iv) can be performed on molecules attached to the support or on molecules that have been released from the support in the supernatant.
[0185] Preferably, the primers used in the optional amplification step (iv) are not attached to the support. However, one or both primers used in step (iv) may be attached to the support. When both primers used in step (iv) are attached to the support, bridging amplification (which may be isothermal) can occur. This bridging amplification can induce polymerase cloning, i.e., clonal clustering of amplicones, which is suitable for sequencing procedures, particularly for NGS (Next Generation Sequencing).
[0186] This implementation allows for the recovery of the template bound to the support, thus improving sample permanence. The support attached to the template can be used to store the sample. The template can be used in different amplifications under different conditions to overcome any possible bias. The recovery of the original template does not require a recapture step based on the binding pair. Therefore, this implementation is particularly suitable for samples with limited amounts of material. Although a recapture step is not required, DNA molecules obtained in step (i) or (ii) can be released from the support and recovered from the reaction mixture obtained after step (iii) or, depending on the specific circumstances, after step (iv). Preferably, the recovery from the reaction mixture is carried out using the first member of the binding pair, wherein the aptamer and / or hairpin sequence and / or barcode sequence are modified with the second member of the binding pair.
[0187] The DNA molecules obtained in step (i) can also be used for sequencing.
[0188] Preferably, the double-stranded DNA molecule used in step (i) is a genomic DNA fragment. Optionally, the double-stranded DNA molecule used in step (i) is end-repaired before step (i), preferably further including a dA-tailing step after the end-repair step. Aptamer molecules and / or hairpin molecules and / or barcode molecules can be provided in the form of a molecular library, wherein each member in the library is distinguishable from other members by a combination of sequences within the molecular sequence. After step (i), the population of aptamer-modified DNA molecules is treated with a reagent that allows the conversion of unmethylated cytosine into bases that are detectably different from cytosine in terms of hybridization properties, and wherein the primers used in steps (iii) and optionally (vi) of the method of the present invention are specific to the aptamer molecules that have been treated with the reagent (step (ii) of the method of the present invention). The combined sequences within the aptamer sequence and / or hairpin sequence and / or barcode sequence may contain modified cytosine resistant to treatment with reagents that allow the conversion of unmethylated cytosine into bases that are detectably different from cytosine in terms of hybridization properties. They may be unmethylated cytosine-free. Optionally, the DNA molecules obtained in step (iii) of the second method of the invention, or, depending on the specific circumstances, step (iv), are recovered from the reaction mixture, preferably using the first member of the binding pair, wherein the primers used in step (iii) of the method of the invention, or, depending on the specific circumstances, step (iv), are modified with the second member of the binding pair. The population of double-stranded DNA molecules may be treated with aptamer molecules prior to step (i) under conditions suitable for the ligation of aptamer molecules to DNA molecules, thereby introducing sticky ends into the DNA molecules.
[0189] Second embodiment of the method of the present invention
[0190] In a second embodiment of the method of the present invention, the double-stranded DNA aptamer connected to at least one end of the strands of a plurality of double-stranded DNA molecules (step (i) of the method of the present invention) has a “Y” form and is referred to as a “Y aptamer”.
[0191] The terms “Y-aptamer” and “Y-adaptor” are used interchangeably, and in the context of this embodiment, they refer to an aptamer formed by two DNA strands, wherein the 3' region of the first DNA strand and the 5' region of the second DNA strand form a double-stranded region through sequence complementarity, wherein the ends of the double-stranded region formed by the 3' region of the first DNA strand and the 5' region of the second DNA strand of the Y-aptamer are compatible with the ends of a double-stranded DNA molecule. As used herein, the term “3' region” refers to the region of the nucleotide chain that includes the 3' end of said chain.
[0192] As used herein, the term "3' end" refers to the end of a nucleotide chain where a hydroxyl group with a third carbon atom is located in the sugar ring of the deoxyribose at its end.
[0193] As used herein, the term "5' region" refers to the region of a nucleotide chain that includes the 5' end of the chain.
[0194] As used herein, the term "5' end" refers to the end of a nucleotide chain where the deoxyribose sugar ring at its end has a fifth carbon.
[0195] As used in this article, the term "sequence complementarity" refers to the shared property between two nucleic acid sequences such that when they are antiparallel aligned, the nucleotide bases at each position will be complementary.
[0196] Multiple paired, adaptor-modified DNA molecules obtained according to step (i) of the method according to this second embodiment of the invention are obtained as follows:
[0197] (a) Connecting a DNA Y-aptamer to each end of a plurality of double-stranded DNA molecules, the aptamer comprising a first DNA strand and a second DNA strand.
[0198] The 3' region of the first DNA strand and the 5' region of the second DNA strand form a double-stranded region through sequence complementarity.
[0199] The ends of the double-stranded region, formed by the 3' region of the first DNA strand and the 5' region of the second DNA strand of the Y-aptamer, are compatible with the ends of the double-stranded DNA molecule.
[0200] (b) For each strand of the DNA molecule obtained in step (a), using each strand of the DNA molecule obtained in step (a) as a template, a complementary strand is synthesized by polymerase extension from the 3' end of the second DNA strand of the Y-aptamer molecule, thereby pairing each strand of the DNA molecule obtained in step (a) with its synthetic complementary strand to provide a plurality of paired adaptor-modified DNA molecules.
[0201] Preferably, the 3' region of the second DNA strand of the Y-aptamer forms a hairpin loop through hybridization between a first segment and a second segment within the 3' region, wherein the first segment is located at the 3' end of the 3' region of the second DNA strand, and the second segment is located near the 5' region of the second DNA strand.
[0202] Optionally, the 3' region of the second DNA strand of the Y-aptamer does not form a hairpin loop through hybridization between the first and second segments within the 3' region.
[0203] The pairing step of step (i) of the method of the present invention according to this embodiment can occur physically (by the presence of a hairpin sequence in the Y-connector, which allows the original DNA strand to physically pair with its synthetic complementary strand), or by the presence of a barcode sequence in the Y-connector (in the double-stranded region and / or in any single-stranded region, or in all three regions), and in the absence of a hairpin sequence (which allows the original DNA strand to not physically bind with its synthetic complementary strand, but to pair by the presence of at least one barcode sequence), or by both of the above (hairpin sequence and one or more barcode sequences).
[0204] Therefore, in One aspect In this process, the 3' region of the second DNA strand of the Y-aptamer can form a hairpin loop through hybridization between a first segment and a second segment of the 3' region, the first segment being located at the 3' end of the 3' region of the second DNA strand, and the second segment being located near the 5' region of the second DNA strand. In this aspect, each strand of the DNA molecule obtained in step (a) is physically paired with its synthetic complementary strand (by means of at least one hairpin molecule) to provide a paired, adaptor-modified DNA molecule. Of course, a barcode sequence can also be present in the Y-adaptor.
[0205] As used herein, the term "hairpin loop" refers to a region of DNA formed by unpaired bases that occurs when a DNA strand folds and forms a base pair with another part or segment of the same strand.
[0206] As used in this article, the term "hybridization" refers to the non-covalent binding of two single-stranded polynucleotides or two regions of the same strand to form a stable double-stranded polynucleotide.
[0207] Therefore, this aspect allows for the acquisition of double-stranded DNA libraries in which the original sense and antisense strands of a DNA molecule are physically bound to each other. Each original strand of the DNA molecule is physically bound to a complementary strand obtained through synthetic extension. Figure 5 and Figure 6 A schematic diagram of this implementation is shown in the figure.
[0208] As used herein, the term “DNA library” refers to a collection of DNA fragments that have been ligated with aptamer molecules to identify and isolate DNA fragments of interest.
[0209] In the context of this invention, the term "double-stranded DNA library" refers to a library containing one of the original strands of a DNA molecule physically linked through one of its ends to a complementary strand obtained through synthetic extension. The double-stranded DNA library of the third method of this invention is not a circular library. The original DNA strand of the DNA molecule and its synthetic complementary strand are physically linked through a loop via one of its ends, thus forming a double strand between them (see [link to third method]). Figure 12In this aspect of a second embodiment of the method of the invention, each original DNA strand of the DNA molecule is paired with its synthetic complementary strand via a hairpin loop, thus physically paired. Each molecule in the double-stranded DNA library may also be in a linear conformation when the complementarity between the two strands is partially or completely lost.
[0210] In another aspect of a second embodiment of the method of the invention, pairing is achieved by the presence of at least one barcode sequence in the Y-adaptor (in any of its single-stranded regions, and / or in its double-stranded regions, or in both single-stranded and double-stranded regions). According to this aspect, the 3' region of the second DNA strand of the Y-aptor does not form a hairpin loop through hybridization between the first and second segments within the 3' region. In this case, each original DNA strand of the DNA molecule pairs with its synthetic complementary strand by the presence of a barcode sequence in the double-stranded region or the single-stranded region of the Y-aptor, or by the presence of a barcode sequence anywhere else on the original DNA strand.
[0211] Of course, the pairing of the original DNA strand of a DNA molecule with its synthetic complementary strand can be achieved both physically (by the presence of a hairpin loop in the Y aptamer, as described above) and by the presence of one or more barcode sequences.
[0212] Y aptamers can contain one or more barcode sequences in their double-stranded DNA region. This will provide at least... Original double Pairing between each original DNA strand of the strand DNA molecule .
[0213] In this case, the original sense and antisense strands of the DNA molecule are identified by a combination marker (“barcode sequence” or “combination barcode”), specifically, each of the sense and antisense strands is linked to two combination sequences. Since these two combination sequences are identical for both the sense and antisense strands, the two strands can be tracked during the process. After the entire process, the complementary strands that were initially together will share the same two combination sequences. This allows for the tracking of the two strands of each double-stranded DNA fragment initially used in step (i) of the method of the present invention. Figure 13 An embodiment of the method of the present invention is shown, wherein combined notation is used.
[0214] Alternatively / in addition, the Y-aptamer may contain one or more barcode sequences in the 5' region of the first DNA strand and / or the 3' region of the second DNA strand (and / or in the double-stranded region) of the Y-aptamer formed by two DNA strands. Therefore, the barcode sequence may be located in the single-stranded region of the Y-aptamer molecule and / or in the double-stranded region of the Y-aptamer. In this case, each original DNA strand is then paired with its synthetic complementary strand.
[0215] Preferably, the DNA Y adaptor has a first barcode sequence in the double-stranded region and / or a second barcode sequence in the 3' region of the second DNA strand of the Y-aptamer. Optionally, the DNA Y adaptor has a first barcode sequence in the double-stranded region and / or a second barcode sequence in the 5' region of the first DNA strand of the Y-aptamer. Optionally, the DNA Y adaptor has a first barcode sequence in the double-stranded region and / or a second barcode sequence in the 3' region of the second DNA strand of the Y-aptamer and / or a third barcode sequence in the 5' region of the first DNA strand of the Y-aptamer.
[0216] Preferably, the DNA Y-adaptor has a restriction site in the 5' region of the first DNA strand of the Y aptamer.
[0217] When each original DNA strand of a paired double-stranded DNA molecule is paired with its synthetic complementary strand, this is called "double pairing" (i.e., the strand that is paired with both its original complementary strand and its synthetic complementary strand simultaneously). Double pairing provides intrinsic confirmation of each nucleotide readout by allowing comparison of four different sources of molecular information (the upper and lower strands of a given dsDNA molecule and their respective synthetic complementary strands), which further enhances the reliability of the results. Additionally, this allows for the evaluation of both the upper and lower strands of the original double-stranded DNA molecule, and thus the analysis of hemimethylation at the genome scale. Preferably, the plurality of paired, adaptor-modified DNA molecules obtained after step (i) according to this embodiment are double-paired, as described above.
[0218] Preferably, the double-stranded DNA molecule used in step (a) is Genomic DNA Fragments. Preferably, the genomic DNA fragments used in pairing step (a) provide multiple paired genomic DNA fragments. As described above, this pairing is preferably performed using barcode sequences.
[0219] Preferably, the double-stranded DNA molecules used in step (a) are end-repaired before step (a), and preferably further include a step of adding dA tails to the DNA molecules after the end-repair step.
[0220] The ligation step (a) is performed under conditions suitable for the ligation of the Y-aptamer to the two ends of the double-stranded DNA molecule.
[0221] The result of step (a) is a plurality of DNA molecules containing Y-aptamers. The molecules are paired (by hairpin sequences and / or by barcode sequences) double-stranded DNA molecules having one Y-aptamer attached to each end of the molecule.
[0222] Preferably, in this embodiment, the DNA molecules containing Y-aptamers obtained in step (b) pretreatment step (a) are obtained under conditions suitable for separating the strands of DNA molecules containing Y-aptamers.
[0223] Suitable conditions for separating the strands of a DNA molecule containing a Y-aptamer can be, but are not limited to, the following, including, for example, denaturing the two strands by heating the molecule to 94-98°C for 20 seconds to 2 minutes, causing the hydrogen bonds between complementary bases to break and producing single-stranded DNA molecules. Strand separation can also be achieved without heating the molecule using isothermal techniques, for example, by using a strand-displacement DNA polymerase, such as, but not limited to, a large fragment of Phi29 DNA polymerase or Bacillus stearothermophilus DNA polymerase.
[0224] After aptamer ligation in step (a), each strand of the DNA molecule obtained in step (a) is used as a template to convert each strand of the DNA molecule obtained in step (a) into a paired double-stranded DNA molecule by polymerase extension from the 3' end of the second DNA strand in the Y-aptamer molecule (step (b) above).
[0225] In the context of this embodiment, the phrase "converting each strand into a paired double-stranded DNA molecule" refers to synthesizing DNA strands complementary to each other, where the two strands are paired. This can be physically achieved. pair (That is, when the 3' region of the second DNA strand of the Y-aptamer forms a hairpin loop through hybridization between the first segment and the second segment within the 3' region, the first segment is located at the 3' end of the 3' region of the second DNA strand, and the second segment of the second DNA strand is located near the 5' region), a double-stranded DNA conformation is produced, in which a single DNA strand folds itself.
[0226] Optionally, when the 3' region of the second DNA strand of the Y-aptamer No When a hairpin loop is formed by hybridization between the first and second segments within the 3' region, it can be achieved by introducing a barcode sequence into at least one strand of the Y-aptamer, i.e., in a double-stranded region or a single-stranded region, preferably in a single-stranded region. pair .
[0227] As described above, pairing can be achieved through both the physical connection of the original DNA strand with its synthetic complementary strand and the presence of one or more barcode sequences.
[0228] As used herein, "polymerase elongation" refers to the synthesis of a complementary strand by a DNA polymerase, which adds free nucleotides to the 3' end of the second DNA strand in a Y-aptamer molecule. The Y-aptamer molecule can act as a primer for the elongation step. During this step, the temperature is selected based on the optimal temperature for the specific DNA polymerase used.
[0229] Preferably, nucleotides A, G, C, and T are used for step (b). Optionally, methylated cytosine can be used for the elongation step, but instead, unmethylated cytosine is used to address an important issue:
[0230] 1-Control of bisulfite conversion. The bisulfite conversion efficiency is variable, meaning that not all cytosine is successfully converted (preferably to uracil (U)). Furthermore, it is important to assess this efficiency in each experiment and even in each molecule within a given experiment.
[0231] Using methylated cytosine (C) on the new chain will not allow for the assessment (or control) of the extent of C>U conversion, because each C will be methylated and therefore read as C. In contrast, when using unmethylated C, the bisulfite conversion efficiency of each individual molecule can be determined.
[0232] 2 – Promotes amplification by reducing the complementarity of the two strands paired by hairpins.
[0233] Following step (i) of this embodiment, two double-stranded DNA molecules are obtained from each DNA molecule containing a Y-aptamer, and each double-stranded DNA molecule is formed from the original DNA strand of the paired DNA molecule and its synthetic complementary strand (which may be physically linked by a hairpin molecule through one of their ends, or they may contain a barcode sequence, or both, as described above).
[0234] The pairing between the two strands of the original double-stranded DNA molecule allows for the tracking of both strands of each double-stranded DNA fragment originally used.
[0235] Therefore, each Y-aptamer can contain a unique combinatorial barcode, which allows for sample identification, multiplexing, and quantitative analysis. In a preferred embodiment, Y-aptamers are provided in the form of an aptamer library, wherein each member of the library is distinguishable from other members by a combinatorial sequence located within a double-stranded region formed by the 3' region of the first DNA strand and the 5' region of the second DNA strand of the aptamer. Thus, the present invention provides Y-hairpin aptamers containing at least one barcode sequence, particularly for use in the methods of the present invention.
[0236] When a library of Y-aptamers with combined sequences is used in step (i) of the method of the present invention, each DNA molecule containing a Y-aptamer obtained after step (i) will have two different combined sequences, each located at one of the Y-aptamers attached to each end of the DNA molecule. The two identifiers bind to individual molecules in the starting sample, thus giving the ability to distinguish constructs. The identifiers allow for the identification of specific constructs containing the identifiers and their progeny, because after step (b) of this embodiment, the original sense and antisense strands of the double-stranded DNA molecule remain in different molecules, but each of these molecules will contain both combined sequences. It is assumed that any amplification products of the original individual molecules carrying the identifiers are pedigree identical. The combined barcodes also allow for the quantification of the percentage of individual sequences within the sample and can be used to monitor bias and error control during the amplification step.
[0237] Optionally, the Y-aptamer is incorporated with a base labeled with a second member of the binding pair, which allows for the recovery of the original DNA template after the elongation or amplification step. This provides the advantages of identifying the sample used as a DNA template, preserving it during the process, and recovering, storing, and performing multiple amplifications under different conditions and sequencing without depleting the sample.
[0238] The construct obtained after step (i) according to this embodiment forms a double-stranded DNA library, which can be used for sequencing or in other conventional molecular biology techniques.
[0239] Optionally, the Y-adaptor may contain a “cleavage site,” as described above. The addition of a “cleavage site” adapts the final elements of the library to the needs of different sequencing platforms. While this adaptation can be achieved through the specific design of the Y-adaptor (by introducing sequences compatible with platform reagents, such as sequencing primers), the cleavage site allows for modularity, to add barcodes or aptamers for multiplexing (mixing samples from different sources) or to meet the needs of any platform used for large-scale sequencing (or also to eliminate potentially unnecessary nucleotides). A “cleavage site” is a specific sequence that allows the presence of a known target at the edge of multiple paired, adaptor-modified DNA molecules (a library of paired, adaptor-modified DNA molecules obtained in step (i), or a library of paired and transformed, adaptor-modified DNA molecules obtained in step (iii) (and optionally in step (iv))). The “cleavage site” can be ligated to multiple double-stranded DNA molecules before or after the step of ligating the Y-adaptor. As described above, the “cleavage site” may already be included in the Y-adaptor. This allows for the cutting of all fragments and the correct ligation of adaptors (thus, sequences of adaptors that are no longer needed can be removed to improve sequencing efficiency).
[0240] Subsequently, step (ii) of the method of the present invention is performed, namely, converting the (unmethylated) cytosine (preferably converted to uracil) present in the plurality of paired, adaptor-modified DNA molecules obtained in step (i) of this embodiment of the method of the present invention into a plurality of paired, adaptor-modified DNA molecules, as previously described.
[0241] Therefore, multiple paired DNA molecules containing aptamers are treated with a reagent that allows the conversion of unmethylated cytosine into a base (preferably uracil) that is detectably different from cytosine in terms of hybridization properties, and wherein the primers used in step (iii) (and optionally step (iv)) are complementary to at least a portion of the sequence derived from the double-stranded DNA molecule obtained by treating step (ii) with said reagent. This treatment imparts to the DNA molecule containing the aptamers that the complementarity between the original strand and the synthetic strand is partially or completely lost, thus promoting the annealing of the primers used in subsequent steps. Figure 5 and 6 A schematic diagram is shown, illustrating this point.
[0242] In a preferred embodiment, the reagent is a bisulfite that converts all unmethylated cytosine into uracil, which will be read as thymine in the molecule amplified in step (iii).
[0243] When the paired, aptamer-modified DNA molecule obtained in step (i) is treated with a reagent capable of converting unmethylated cytosine into bases that are detectably different from cytosine in terms of hybridization properties, the complementarity between the sense and antisense strands of the original double-stranded DNA molecule is partially or completely lost. This can promote the synthesis of the complementary strand.
[0244] The DNA molecules containing Y-aptamers obtained in step (ii) of the method of the present invention can be further processed prior to step (iii) under conditions suitable for the separation of the strands of DNA molecules containing Y-aptamers.
[0245] Suitable conditions for the separation of strands of DNA molecules containing Y-aptamers can be, but are not limited to, the following, wherein, for example, denaturation of the two strands is achieved by heating the molecules to 94-98°C for 20 seconds to 2 minutes, causing the hydrogen bonds between complementary bases to break and producing single-stranded DNA molecules. The separation of strands can also be achieved without heating the molecules using isothermal techniques, for example, by using a strand-displacement DNA polymerase, such as, but not limited to, a large fragment of Phi29 DNA polymerase or Bacillus stearothermophilus DNA polymerase. Subsequently, step (iii) of the method of the present invention is performed, namely, providing paired and transformed complementary strands of a plurality of adaptor-modified DNA molecules.
[0246] Optionally, the construct (step (iv)) can be amplified to increase the amount of material used in the following steps. Preferably, the double-stranded DNA molecule obtained in step (iii) is amplified using primers whose sequences are complementary to at least a portion of the double-stranded DNA molecule obtained in step (iii).
[0247] Preferably, primers can be used in the first amplification step to amplify the double-stranded DNA molecule obtained in step (iii), wherein the sequence of the primers is complementary to at least a portion of the complementary sequence of the 5' region of the first DNA strand of the aptamer molecule. Preferably, the primers used in step (iii) (and optionally step (iv)) are complementary to at least a portion of the sequence derived from the double-stranded DNA molecule obtained by treating it with reagents in step (ii). This embodiment in Figure 5 and 6 The data shows that the full sequence of a population of double-stranded DNA molecules used to generate double-stranded DNA libraries can be amplified.
[0248] Alternatively, primers can be used in the first amplification step to amplify the double-stranded DNA molecules obtained in step (iii), wherein the sequence of the primers is at least a portion complementary to the complementary sequence of the population of double-stranded DNA molecules used to generate the double-stranded DNA library. Preferably, the primers used in step (iii) (and optionally step (iv)) are at least a portion complementary to the sequence derived from the double-stranded DNA molecules obtained in step (ii) by treatment with reagents. This embodiment in Figure 6 It is displayed in the middle.
[0249] Subsequent amplification steps can also be performed using primer pairs. This invention covers any combination of primers. For example, the first primer may be complementary to at least a portion of the complementary sequence of the 5' region of the first DNA strand of the aptamer molecule, and the second primer may be complementary to the 3' region of the molecule obtained after the first amplification step. Figure 6 ).
[0250] Figure 12 The arrangement of different elements in the double-stranded DNA library obtained through this embodiment is shown. It is noteworthy that aliquots of the same sample can be amplified under different conditions in a manner that allows any bias (TA or CG bias) to be assessed and accounted for during the analysis phase.
[0251] Optionally, molecules obtained after steps (iii) and / or (iv) of the method of the present invention can be recovered from the reaction mixture. Thus, molecules obtained in steps (iii) and / or (iv) can be recovered from the reaction mixture, preferably using the first member of the binding pair, wherein the primers used in steps (iii) and / or (iv) are modified with the second member of said binding pair.
[0252] Optionally, the composite sequence may contain modified cytosine, which is resistant to treatment with a reagent that allows the conversion of unmethylated cytosine into bases that are detectably different from cytosine in terms of hybridization properties. Alternatively / additionally, the composite sequence may contain unmodified cytosine (which is intolerant to treatment with a reagent that allows the conversion of unmethylated cytosine into bases that are detectably different from cytosine in terms of hybridization properties).
[0253] The terms “double-stranded DNA molecular population,” “end,” “compatible,” “ligation,” “template,” “primer,” “complement,” “aptamer library,” “combination sequence” or “combination barcode,” “reaction mixture,” “binding pair,” “first member of binding pair,” “second member of binding pair,” and “bases that are detectable to a different degree from cytosine in terms of hybridization characteristics” are defined in the context of the method of this invention.
[0254] Y-aptamers can be provided immobilized in a support. Preferably, the immobilization is performed by binding a nucleotide (if present) to the 5' end of the first DNA strand of the Y-aptamer or to the hairpin loop of the second DNA strand. The primers used in step (iii) can also be attached to the support. Preferably, the binding of the aptamer and / or primers to the support can be covalent.
[0255] The terms “fixation”, “support” and “covalent bond” have already been defined above.
[0256] For example, multiple paired, adaptor-modified DNA molecules obtained in step (i) of the method of the present invention are obtained as follows:
[0257] (a) Contacting a population of double-stranded DNA molecules with a DNA Y-aptamer, said aptamer comprising a first DNA strand and a second DNA strand.
[0258] The 3' region of the first DNA strand and the 5' region of the second DNA strand form a double-stranded region through sequence complementarity, and the ends of the double-stranded region are compatible with the ends of the double-stranded DNA molecule.
[0259] The contact is performed under conditions suitable for the ligation of the Y-aptamer to both ends of a double-stranded DNA molecule, thereby obtaining multiple DNA molecules containing the Y-aptamer.
[0260] (b) Each strand of the DNA molecule containing the Y-aptamer is contacted with an extension primer under conditions suitable for hybridization of the extension primer with the second strand of the Y-aptamer, the extension primer containing a 3' region complementary to the second DNA strand of the Y-aptamer molecule, and generating overhanging ends after hybridization with the second DNA strand of the Y-aptamer molecule.
[0261] (c) The molecule produced in step (b) is brought into contact with the hairpin aptamer under conditions suitable for the connection between the hairpin aptamer and the molecule produced in step (b), the hairpin aptamer comprising a hairpin loop region and a cantilevered end, the cantilevered end being compatible with the cantilevered end in the molecule produced in step (b).
[0262] (d) Each strand of the DNA molecule obtained in step (c) is converted into a double-stranded DNA molecule by polymerase extension from the extension primers used in step (b).
[0263] The steps (c) of connecting to the hair clip adapter and (d) of extending the hair clip adapter can be performed in any order or simultaneously.
[0264] The terms "Y-aptamer" and "Y-adaptor" are used interchangeably and, similarly as above, refer to an aptamer formed by two DNA strands, wherein the 3' region of the first DNA strand and the 5' region of the second DNA strand form a double-stranded region through sequence complementarity, and wherein the ends of said double-stranded region are compatible with the ends of a double-stranded DNA molecule. In this case, the Y-aptamer does not contain a hairpin loop, i.e., the 3' region of the second DNA strand of the Y-aptamer. No A hairpin is formed by the hybridization between the first and second segments within the 3' region.
[0265] Preferably, the double-stranded DNA molecule used in step (a) is a genomic DNA fragment. Preferably, the double-stranded DNA molecule used in step (a) is end-repaired before step (a), and preferably further includes a step of adding dA tails to the DNA molecule after the end-repair step.
[0266] The contact step (a) is performed under conditions suitable for the ligation of the Y-aptamer to both ends of the double-stranded DNA molecule.
[0267] The result of the steps is a plurality of DNA molecules containing Y-aptamers, which are double-stranded DNA molecules having one Y-aptamer attached to each end of the molecule.
[0268] Step 2(b) involves contacting each strand of the DNA molecule containing the Y-aptamer with an elongating primer.
[0269] As used herein, the term "extension primer" refers to a primer used in a next step of the method for extension, the primer containing a 3' region complementary to the second strand of the Y-aptamer molecule, thereby generating a dangling end. The terms "primer" and "dangling end" are defined above.
[0270] Step 3 (c) involves bringing the molecules generated in step (b) into contact with the hairpin aptamer under conditions suitable for the connection between the hairpin aptamer and the molecules generated in step (b).
[0271] As used herein, the term "hairpin aptamer" refers to a double-stranded structure formed from a single-stranded nucleic acid that folds itself to form a double-stranded region maintained by base pairing between complementary base sequences on the same strand, a hairpin loop region formed by unpaired bases, and a dangling end that is compatible with the dangling ends in the molecule produced in step (b).
[0272] Figure 7 An embodiment of the method of the present invention is shown, wherein the hairpin adapter and the extension primer are provided separately.
[0273] Alternatively, by providing hairpin aptamers and elongation primers in the form of a complex... Single step The steps (b) and (c) are performed to contact each strand of the DNA molecule containing the Y-aptamer with an extension primer and to contact the molecule produced in step (b) with a hairpin aptamer.
[0274] As used in this article, the term "complex" refers to a unique molecule formed by a hairpin aptamer and an elongated primer. Figure 8 and 9 Different embodiments of the method according to the invention are shown, wherein the elongation primer and hairpin aptamer are provided in the form of a complex with different conformations. In this case, after the Y-aptamer is connected, the elongation primer contained in the complex is annealed, and a connection is performed between the 3' end of the second strand of the Y-aptamer and the 5' end of the hairpin aptamer.
[0275] Steps (c) for connecting the hairpin adapter and (d) for extending the hairpin adapter can be performed in any order or simultaneously. In one embodiment, step (c) is performed before step (d). In another embodiment, step (d) is performed before step (c) (i.e., the extension is performed before connecting the hairpin adapter to the building block). In yet another embodiment, steps (c) and (d) are performed simultaneously.
[0276] Preferably, the DNA molecule containing the Y-aptamer obtained in step (a) or step (c) is placed under conditions suitable for separating the strands of the DNA molecule containing the Y-aptamer.
[0277] The constructs obtained after step (d) form a double-stranded DNA library and can be used for sequencing or other conventional molecular biology techniques.
[0278] Y-aptamers are provided in the form of aptamer libraries to pair multiple double-stranded DNA fragments (original strand and its complementary original strand), wherein each member of the library is distinguishable from other members by a combinatorial sequence (also known as a barcode sequence) located in a double-stranded region formed by the 3' region of the first DNA strand and the 5' region of the second DNA strand of the aptamer.
[0279] Optionally, the Y-aptamer, extension primer, and / or hairpin aptamer are incorporated with bases labeled with a second member of the binding pair, which allows for the recovery of the original DNA template after the extension or amplification step.
[0280] Optionally, the molecules obtained after step (d) can be recovered from the reaction mixture. Thus, the molecules obtained in step (d) can be recovered from the reaction mixture, preferably using the first member of the binding pair, wherein the primer used in step (e) is modified with the second member of said binding pair.
[0281] Once multiple paired adaptor-modified DNA molecules have been provided (step (i) of the method of the present invention), the paired DNA molecules containing aptamers are treated with a reagent that allows unmethylated cytosine to be converted to a base (preferably uracil) that is detectably different from cytosine in terms of hybridization properties, and wherein the primers used in steps (i) and (b) above are complementary to at least a portion of the sequence derived from the double-stranded DNA molecules treated with said reagent (step (ii) of the method of the present invention).
[0282] Preferably, the reagent is a bisulfite.
[0283] Optionally, to prevent specific portions of the combined sequence (barcode sequence) from changing after treatment with the reagent, the combined sequence may contain modified cytosine that is resistant to treatment with the reagent, which allows unmethylated cytosine to be converted into bases that are detectably different from cytosine in terms of hybridization properties.
[0284] For example, hairpin aptamers and / or Y-aptamers can be provided fixed in a support, as previously described. Preferably, the fixation is performed by binding the nucleotides of the hairpin loop of the hairpin aptamer and / or the 5' end of the first DNA strand of the Y-aptamer to the support.
[0285] In another embodiment, the primers used in steps (i) and (d) as described above with respect to this embodiment, and optionally in steps (iii) and (iv) of the method of the invention (for optionally amplifying partially transformed paired double-stranded DNA molecules), are also attached to the support. Preferably, the binding between the aptamers and / or primers and the support is covalent.
[0286] Third embodiment of the method of the present invention
[0287] In a third embodiment of the method of the present invention, a plurality of paired, adaptor-modified DNA molecules of step (i) of the method of the present invention can be obtained by the following:
[0288] (a) A plurality of double-stranded DNA molecules are fragmented under conditions suitable for generating a plurality of double-stranded DNA molecular fragments with overhanging ends, wherein each end of each fragment is bound to a hemiaptamer molecule, the hemiaptamer molecule comprising a first DNA strand and optionally a second DNA strand, wherein the second strand forms a double-stranded region with the first strand via complementarity with the central region of the first strand, and wherein the hemiaptamer molecule binds to the double-stranded DNA molecular fragment between the 3' end of the first strand of the hemiaptamer and the overhanging end of the double-stranded DNA molecular fragment.
[0289] (b) Adding an alternative second strand or replacing the second DNA strand of a hemi-aptamer molecule with an alternative second strand, wherein the 5' region of the alternative second strand is complementary to the 3' region of the first strand of the hemi-aptamer molecule, and wherein the alternative second strand contains a region not complementary to the first strand of the hemi-aptamer molecule, thereby producing a plurality of Y-aptamer-modified DNA molecules.
[0290] (c) Optionally, fill the gap between the 5' end of the alternative second strand of the Y-aptamer and the 3' end of each DNA fragment.
[0291] (d) Each strand of the DNA molecule containing the Y-aptamer is contacted with an extension primer under conditions suitable for hybridization of the extension primer with the alternative second strand of the Y-aptamer, the extension primer comprising a 3' region complementary to the alternative second DNA strand of the Y-aptamer molecule and a 5' region not hybridizing with the alternative second DNA strand of the Y-aptamer molecule.
[0292] (e) Pairing the molecules produced in step (d). Preferably, pairing can be achieved by contacting the molecules produced in step (d) with hairpin molecules under conditions suitable for the connection of the hairpin aptamer and the molecules produced in step (d), the hairpin molecules comprising a hairpin loop region and an end compatible with the end of the molecules produced in step (d); optionally, pairing can be achieved by the presence of a barcode sequence (in the absence of hairpin molecules), the barcode sequence being added after step (d) or already present in the Y-adaptor and / or extension primer;
[0293] (f) Each strand of the DNA molecule obtained in step (e) is converted into a double-stranded DNA molecule by polymerase extension from the extension primers used in step (d), thereby obtaining multiple paired aptamer-modified DNA molecules.
[0294] The steps (e) of connecting to the hair clip adapter and the extension step (f) can be performed in any order or simultaneously.
[0295] This third embodiment of the invention is suitable for different fragmentation systems.
[0296] In the first step, the method for generating a double-stranded DNA library according to a third embodiment of the present invention involves fragmenting a plurality of double-stranded DNA molecules under conditions suitable for generating a plurality of double-stranded DNA molecule fragments with overhanging ends, wherein each end of each fragment is bound to a semi-aptamer molecule between the 3' end of the first strand of the semi-aptamer and the overhanging end of the fragment of the double-stranded DNA molecule.
[0297] The terms "hemapter" and "hemapter" are used interchangeably and refer to an incomplete aptamer formed by a first DNA strand and optionally a second DNA strand, wherein the second strand forms a double-stranded region with the first strand via complementarity with the central region of the first strand. The hemiapter of this embodiment does not contain a hairpin loop. For example, the hemiapter does not contain a second DNA strand. For example, the hemiapter contains both a first DNA strand and a second DNA strand.
[0298] Preferably, the double-stranded DNA molecule used in step (a) is a genomic DNA fragment.
[0299] The first step of this embodiment can be performed, for example, by in vitro transposition, which includes fragmentation of the double-stranded DNA molecule and ligation of the hemiaptamer to the double-stranded DNA molecule, wherein the transposable element is introduced from the donor DNA (hemiaptamer molecule) into the target DNA (population of double-stranded DNA molecules).
[0300] Preferably, the fragmentation step (a) is carried out by a method comprising contacting a population of double-stranded DNA molecules with a transposase dimer loaded with double-stranded aptamer molecules, wherein the aptamer molecules comprise a double-stranded region comprising a Tn5 inverted repeat and a 5' overhang of one of the strands, wherein optionally, cytosine nucleotides in the double-stranded region that do not form part of the Tn5 inverted repeat and cytosine nucleotides in the single-stranded region are methylated, and wherein the contact is carried out under conditions suitable for DNA fragmentation and attachment of hemiaptamer molecules to both ends of each DNA fragment.
[0301] As used herein, the term "transposase" refers to an enzyme (EC 3.1.-.-) that recognizes a specific DNA sequence, cleaves two double-stranded DNA molecules at four locations, and joins the strands. Transposases form complexes with nucleic acids capable of transposition (i.e., catalyzing the insertion of nucleic acids into the target DNA sequence).
[0302] As used herein, the term "transposase dimer" refers to a dimer of two chemically identical monomers containing some residues. Any transposase dimer (natural or mutant) from any species may be used in this invention. Of particular interest are, but not limited to, transposases Tn5, Tn3, Tn7 and their mutants, as well as retroviral integrase. In a preferred embodiment, the transposase dimer is the Tn5 transposase. The term "Tn5 transposase" refers to a member of the RNase protein superfamily, including retroviral integrase. The Tn5 transposase of this invention is found in *Escherichia coli* and is defined by sequence Q46731, version April 3, 2013, in the UniProt database. This invention also includes functionally equivalent variants of the said Tn5 transposase, including natural variants found in other species (e.g., *Shewanella*) and artificial variants obtained through molecular biology techniques (e.g., the mutant Tn5 transposase disclosed in US5965443).
[0303] As used in this article, the term "loading" refers to the binding of the transposase dimer to a double-stranded DNA fragment.
[0304] In the context of this embodiment, the term "double-stranded aptamer molecule" refers to an aptamer molecule containing a double-stranded region, wherein the double-stranded region includes a Tn5 reverse repeat and a 5' overhang of one of the strands.
[0305] As used herein, the term "Tn5 inverted repeat" refers to a transposable element. It is generally assumed that the length of the Tn5 inverted repeat is 18 or 19 bases and that it is an inverted repeat relative to another Tn5 inverted repeat (Johnson RC and Reznikoff WS 1983. Nature, 304:280). Tn5 inverted repeat sequences are well known in the art.
[0306] The cytosine nucleotides in the double-stranded region that do not form part of the Tn5 inverted repeat and in the single-stranded region may be methylated or unmethylated. In one specific embodiment, the cytosine nucleotides are methylated.
[0307] As used in the context of this embodiment, the phrase "conditions suitable for DNA fragmentation and attachment of hemiaptamer molecules to both ends of each DNA fragment" refers to the time, temperature, and buffer composition conditions suitable for the proper functioning of the loaded transposase dimerase. These conditions are well known to those skilled in the art. Exemplary conditions may be those disclosed in Adey A. and Shendure J. 2012. Genome Research, 22:1139-1143 and Adey A. et al. 2010. Genome Biology, 11:R119.
[0308] The kit suitable for the first step (a) of this embodiment is, for example, Nextera. TM DNA Sample Preparation Kit (Illumina).
[0309] In vitro transposition can be implemented using tagging (“Ultra-low-input, tagging-based-whole-genome bisulfite sequencing” Adey A. and Shendure J. 2012. Genome Research, 22:1139-1143; “Rapid, low-input, low-bias construction of shotgunfragment libraries by high-density in vitro transposition” Adey A. et al. 2010. Genome Biology, 11:R119) or modifications of the tagging method disclosed in patent US5965443. Other methods that can be used in this invention include, but are not limited to, those disclosed in EP2527438A1, US2003143740A, US7160682B, and WO9925817A.
[0310] The result of step (a) is a plurality of DNA molecules containing hemiaptamers, which are double-stranded DNA molecules having one hemiaptamer attached to each end of the molecule.
[0311] The second step (b) of this embodiment involves adding an alternative second strand or replacing the second DNA strand of the semi-aptamer molecule with an alternative second strand to obtain a plurality of DNA molecules containing Y-aptamers.
[0312] As used herein, the term "displacement" means that the second DNA strand of a hemiaptamer is replaced by an alternative second strand. Displacement occurs under conditions where the original second DNA strand of the hemiaptamer loses affinity relative to the alternative second strand, by controlling the temperature, strand concentration, and strand melting temperature during the reaction. In short, by raising the temperature to, for example, 50°C over a period of 2 minutes in a suitable buffer, the original second DNA strand loses hybridization with the first strand of the hemiaptamer. The mixture is then cooled, and the alternative second strand displaces the original second DNA strand. Exemplary conditions for displacing the second DNA strand of a hemiaptamer are disclosed in Adey A. and Shendure J. 2012. Genome Research, 22:1139-1143.
[0313] As used herein, the term "alternative second strand" refers to a strand having a 5' region complementary to the 3' region of the first strand of the hemimeric molecule and containing regions not complementary to the first strand of the hemimeric molecule. In one embodiment, the 5' region of the alternative second strand complementary to the 3' region of the first strand of the hemimeric molecule has no gap between the 5' end of the alternative second strand and the 3' end of the DNA fragment. In another embodiment, a gap exists between the 5' end of the alternative second strand and the 3' end of the DNA fragment.
[0314] In the context of this embodiment, the term "Y-aptamer" refers to an aptamer formed by two DNA strands, wherein the 3' region and / or central region of the first DNA strand and the 5' region of the alternative second strand form a double-stranded region through sequence complementarity, and wherein the 5' region of the first strand and the 3' region of the alternative second strand are not complementary.
[0315] In some cases, the 5' end of the alternative second strand of the Y-aptamer and the 3' end of each DNA fragment are not connected because gaps may exist between the ends.
[0316] Optionally, this embodiment involves step (c) of filling the gap between the 5' end of the alternative second strand of the Y-aptamer and the 3' end of each DNA fragment.
[0317] As used herein, the term “gap” refers to an interruption in one of the two DNA strands due to the loss of one or more nucleotides.
[0318] As used in this article, the term "gap filling" refers to adding a missing nucleotide to the strand. DNA polymerase inserts the correct nucleotide into the gap and ligates it to the 3' end of the strand by recognizing which base corresponds to the gap on the complementary DNA strand. Because there is a nick at the gap, DNA ligase is used to join the DNA strands on either side of the nick.
[0319] When the first DNA strand of the hemiaptamer contains a combined sequence, step (c) of this embodiment generates a complementary copy of the combined sequence by gap filling.
[0320] Step (d) of this embodiment involves contacting each strand of the DNA molecule containing the Y-aptamer with the extension primer under conditions suitable for hybridization of the extension primer with the alternative second strand of the Y-aptamer.
[0321] As used in the context of this embodiment, the term "extension primer" refers to a primer used for extension in a subsequent step of the method, which comprises a 3' region complementary to the alternative second DNA strand of the Y-aptamer and a 5' region preferably not hybridizing to the alternative second DNA strand of the Y-aptamer.
[0322] Step (d) of this embodiment can produce a flat end or a cantilevered end, preferably a cantilevered end.
[0323] Step (e) of this embodiment involves the molecules produced in step (d) pairing, preferably by contacting the molecules produced in step (d) with the hairpin aptamers under conditions suitable for linking the hairpin aptamers to the molecules produced in step (d). Optionally, pairing can be achieved by the presence of a barcode sequence (in the absence of hairpin molecules), which is added after step (d) or is already present in the Y-aptamer and / or the extended primer.
[0324] In the context of this embodiment, the terms "hairpin aptamer," "hairpin sequence," and / or "hairpin molecule" refer to a double-stranded structure formed from a single-stranded nucleic acid, wherein the single-stranded nucleic acid folds itself to form a double-stranded region maintained by base pairing between complementary base sequences on the same strand, a hairpin loop region formed by unpaired bases, and a terminal end compatible with the terminal end of the molecule produced in step (d). The hairpin aptamer may contain blunt ends or overhanging ends, preferably overhanging ends.
[0325] In another embodiment, the steps (d) of contacting each strand of the DNA molecule containing the Y-aptamer with the extension primer and (e) of contacting the molecule produced in step (d) with the hairpin molecule are carried out in a single step by providing the hairpin aptamer and the extension primer in the form of a complex.
[0326] As used in this article, the term "complex" refers to a unique molecule formed by a hairpin aptamer and an elongated primer. Figure 10 An example according to this embodiment is shown, wherein an extension primer and a hairpin aptamer are provided in the form of a complex. In this case, the extension primer contained in the complex is annealed, and a connection is performed between the 3' end of the alternative second strand of the Y-aptamer and the 5' end of the hairpin aptamer.
[0327] In another embodiment, the steps of adding a second DNA strand or replacing a semi-aptamer molecule with a second DNA strand in a single step are performed by providing an alternative second strand, a hairpin aptamer, and an extension primer in the form of a complex (b), contacting each strand of the DNA molecule containing the Y-aptamer with the extension primer (d), and contacting the molecule produced in step (d) with the hairpin molecule (e).
[0328] As used in this article, the term "complex" refers to a unique molecule formed by an alternative second strand, a hairpin aptamer, and an elongated primer. Figure 11 An embodiment according to this invention is shown, wherein an alternative second strand, an extension primer, and a hairpin molecule are provided in the form of a complex. In this case, the alternative second strand contained in the complex is annealed with the first DNA strand of the semi-aptamer molecule. The gap-filling step (c) and the extension step (f) can be performed simultaneously with the addition of DNA polymerase and DNA ligase; or they can be performed separately during the reversible closure of the 3' end of the extension primer. Preferably, steps (b) to (f) are performed simultaneously.
[0329] In this embodiment, the steps (e) of connecting the hairpin molecule and the extension step (f) can be performed in any order or simultaneously.
[0330] In a preferred embodiment, the DNA molecule containing a Y-aptamer is placed in step (b) or, depending on the specific circumstances, in step (c) under conditions suitable for separating the strands of the molecule, or, depending on the specific circumstances, in step (e) or in step (f) as a paired DNA molecule containing an aptamer.
[0331] The construct obtained after step (f) of this embodiment forms a double-stranded DNA library and can be used for sequencing or other conventional molecular biology techniques.
[0332] Preferably, the semi-aptamers used in step (a) are provided in the form of a semi-aptamer library, each member of which is distinguishable from other members by a combination sequence in the 3' region of the first strand of the semi-aptamer. For example, the second strand of the semi-aptamer or the alternative second strand used in step (b) does not show any substantial overlap with the combination region. In another example, the second strand of the semi-aptamer or the alternative second strand used in step (b) shows substantial overlap with the combination region.
[0333] As used herein, the expression "not showing any substantial overlap" means that the second strand of the semi-aptamer or the alternative second strand used in step (b) does not extend and cover the combinatorial region. As used herein, the expression "showing substantial overlap" means that the second strand of the semi-aptamer or the alternative second strand used in step (b) extends and partially or completely covers the combinatorial region.
[0334] Optionally, the Y-aptamer, extension primer, and / or hairpin aptamer and / or barcode sequence (if present) are incorporated with bases labeled with a second member of the binding pair, which allows the original DNA template to be recovered after the extension or amplification step.
[0335] Optionally, the molecules obtained after step (f) of this embodiment can be recovered from the reaction mixture. Therefore, in one embodiment of this invention, the molecules obtained after step (f) are preferably recovered from the reaction mixture because the Y-aptamer, extension primer, and / or hairpin molecule, and / or (if so) barcode sequence are incorporated with bases labeled with a second member of the binding pair, which allows for the recovery of the original DNA template after the extension or amplification step.
[0336] Once a plurality of paired, adaptor-modified DNA molecules have been provided according to this embodiment (step (i) of the method of the present invention), the (unmethylated) cytosine present in both strands of the paired, adaptor-modified DNA molecules is converted, for example, into uracil in the paired, adaptor-modified DNA molecules (step (ii) of the method of the present invention).
[0337] Therefore, multiple paired aptamer-modified DNA molecules are treated with a reagent that allows the conversion of unmethylated cytosine into bases that are detectably different from cytosine in terms of hybridization properties, and wherein the primers used in step (g) are complementary to at least a portion of the sequence derived from the double-stranded DNA molecule obtained in step (f) by treating it with said reagent.
[0338] Preferably, the reagent is a bisulfite.
[0339] Optionally, the composite sequence (barcode sequence) (if present) may contain one or more modified cytosines resistant to treatment with a reagent that allows the conversion of unmethylated cytosines into bases that are detectably different from cytosines in terms of hybridization properties. Alternatively / additionally, the composite sequence may contain unmodified cytosines (which are intolerant to treatment with a reagent that allows the conversion of unmethylated cytosines into bases that are detectably different from cytosines in terms of hybridization properties).
[0340] Once multiple paired and transformed adaptor-modified DNA molecules have been generated (step (ii) of the method of the present invention), the complementary strand of the paired and transformed adaptor-modified DNA molecules is provided (step (iii) of the method of the present invention).
[0341] Preferably, step (iii) of the method of the present invention is performed using primers, wherein the sequence of the primers is complementary to at least a portion of the complementary sequence of the 5' region of the first DNA strand of the semi-aptamer molecule (preferably derived from the treatment of a double-stranded DNA molecule with a reagent (step (ii) of the method of the present invention)).
[0342] Preferably, in step (iii) of the method of the present invention, the primer sequence is complementary to at least a portion of the complementary sequence of a population of paired and transformed adaptor-modified double-stranded DNA molecules (preferably complementary to the sequence derived from the treatment of double-stranded DNA molecules with reagents (step (ii) of the method of the present invention, as described above)).
[0343] Subsequent amplification steps using primer pairs are also possible (step (iv) of the method of the present invention). The present invention covers any combination of primers. For example, in one specific embodiment, the primer of step (iii) is complementary to at least a portion of the complementary sequence of the 5' region of the first DNA strand of the semi-aptamer molecule, and the primer of step (iv) is complementary to the 3' region of the molecule obtained after the first amplification step.
[0344] Optionally, molecules obtained after steps (iii) and / or (iv) can be recovered from the reaction mixture. Therefore, the recovery of molecules obtained after steps (iii) and / or (iv) from the reaction mixture preferably uses the first member of the binding pair, wherein the primers used in steps (iii) and / or (iv) are modified with the second member of said binding pair.
[0345] In this embodiment, a first DNA strand of a hairpin aptamer and / or hemiaptamer may be provided immobilized in a support. Preferably, the immobilization is performed by binding the nucleotides of the hairpin loop of the hairpin aptamer and / or the 5' end of the first DNA strand of the hemiaptamer to the support. In this embodiment, the primers used in steps (iii) and / or (iv) may also be attached to the support. Preferably, the binding between the aptamer and / or primer and the support is covalent.
[0346] Other embodiments of the method of the present invention
[0347] 1. The method of the present invention, wherein the pairing of strands of multiple double-stranded DNA molecules is achieved by the presence of a barcode sequence.
[0348] For example, a plurality of paired, adaptor-modified DNA molecules of step (i) of the method of the present invention can be obtained by: ligating a population of double-stranded DNA molecules to a population of DNA aptamers, each aptamer comprising a first DNA strand and a second DNA strand, wherein the 3' region of the first DNA strand and the 5' region of the second DNA strand form a double-stranded region by sequence complementarity, and wherein the ends of the double-stranded regions are compatible with the ends of the double-stranded DNA molecules, wherein each aptamer of the population is distinguishable from other aptamers by a combination sequence located within the double-stranded region, the double-stranded region being formed between the 3' region of the first DNA strand and the 5' region of the second DNA strand, wherein the ligation is performed under conditions suitable for ligation of the aptamers to each end of the double-stranded DNA molecules, thereby obtaining a plurality of DNA molecules containing aptamers.
[0349] Thus, a double-stranded DNA library was obtained that is particularly useful for analyzing the methylation of samples, in which the original sense and antisense strands of the DNA molecule are not physically bound to each other by adapters, but can be identified by combinatorial markers (combined sequences or barcode sequences). Specifically, each of the sense and antisense strands is linked to two combinatorial sequences.
[0350] As used herein, the term “DNA library” refers to a collection of DNA fragments that have been ligated with aptamer molecules to identify and isolate DNA fragments of interest.
[0351] As used herein, the term "double-stranded DNA library" refers to a library containing two strands of a DNA molecule (i.e., a sense strand and an antisense strand), but wherein the strands are not physically linked by a linker. The sense and antisense strands of the original molecule are identified by combinatorial markers (combination sequences or barcode sequences), because the sense strand contains two unique combinatorial barcodes, which are also present in the antisense strand. Each original strand of the DNA molecule binds to a first strand with an aptamer having a unique combinatorial sequence at one end and to a second strand with a different aptamer having a different unique combinatorial sequence at the other end.
[0352] The terms "aptamer" and "adaptor" are used interchangeably and refer to oligonucleotides or nucleic acid fragments or segments that can be linked to nucleic acid molecules of interest.
[0353] As used herein, a “DNA aptamer” comprises a first DNA strand and a second DNA strand, wherein the 3' region of the first DNA strand and the 5' region of the second DNA strand form a double-stranded region through sequence complementarity, and wherein the ends of the double-stranded region are compatible with the ends of a plurality of double-stranded DNA molecules, and wherein each aptamer is distinguishable from other aptamers by a combination of sequences located within the double-stranded region, which is formed between the 3' region of the first DNA strand and the 5' region of the second DNA strand. A DNA aptamer may be formed from substantially complementary first and second DNA strands. Therefore, the 5' region of the first DNA strand and the 3' region of the second DNA strand may be complementary. A DNA aptamer may be a Y-aptamer. Therefore, a DNA aptamer may be a Y-aptamer, wherein the 3' region of the first DNA strand and the 5' region of the second DNA strand may form a double-stranded region through sequence complementarity, and wherein the 5' region of the first DNA strand and the 3' region of the second DNA strand may not be complementary. Figure 14 A schematic diagram illustrating this aspect of the method of the invention is shown. Optionally, the DNA aptamer may be an aptamer of the aptamer disclosed in the "Y-aptamer library of the present invention, methods for synthesizing them and kits" section of this specification. Any embodiments disclosed in that section are applicable to aptamers as used herein. Preferably, the 3' end of the second DNA strand in each aptamer is reversibly closed by a linker that connects the 5' end of the first DNA strand and the 3' end of the second DNA strand.
[0354] Preferably, the double-stranded DNA molecule used in step (i) is a genomic DNA fragment. Optionally, the double-stranded DNA molecule used in step (i) is end-repaired before step (i), preferably further including a step of adding dA tails to the DNA molecule after the end-repair step, as described above.
[0355] Optionally, multiple double-stranded DNA molecules are pretreated with aptamers in step (i) under conditions suitable for the ligation of aptamer molecules to DNA molecules, thereby introducing sticky ends into the DNA molecules.
[0356] The ligation step (i) is performed under conditions suitable for ligating aptamers to each end of a double-stranded DNA molecule to produce multiple DNA molecules containing aptamers.
[0357] The result of these steps is multiple paired DNA molecules containing aptamers. Each paired aptamer-modified DNA molecule will have two distinct combination sequences, each located at one of the aptamers attached to each end of the DNA molecule.
[0358] The constructs obtained after step (i) form the double-stranded libraries of this invention and can be used for sequencing or other conventional molecular biology techniques. The advantage of these libraries is that the combined sequences allow for cross-referencing of sequence information obtained from the initially combined sense and antisense strands to obtain more reliable results.
[0359] The DNA molecule containing the aptamer obtained in step (i) is treated with a reagent that allows unmethylated cytosine to be converted into a base that is detectably different from cytosine in terms of hybridization properties (step (ii)).
[0360] In a preferred embodiment, the reagent is a bisulfite that converts all unmethylated cytosine into uracil, which is read as thymine in the molecule amplified in step (iii).
[0361] Following the treatment in step (ii), the complementarity between the sense and antisense strands of the original double-stranded DNA molecule is partially or completely lost. This facilitates the annealing of primers used in subsequent steps.
[0362] The constructs obtained after step (ii) can also be used for sequencing or other conventional molecular biology techniques, particularly for analyzing sample methylation.
[0363] In the third step (step (iii)) of the method of the present invention, the complementary strand of the paired and transformed aptamer-modified DNA molecule obtained in step (ii) is provided. In this step, primers are used, the sequences of which are complementary to at least a portion of the second DNA strand of the aptamer (derived from the treatment of the double-stranded DNA molecule with the reagent (step (ii)) of the method of the present invention). This demonstrates... Figure 14 In addition, it allows for the acquisition of the complementary strand of the full sequence of a population of paired and transformed adaptor-modified double-stranded DNA molecules used to generate double-stranded DNA libraries.
[0364] After step (ii) has occurred, the primers used in step (iii) are capable of hybridizing with the aptamer-modified DNA molecule that has been treated with a reagent that allows for the conversion of unmethylated cytosine to bases that are detectably different from cytosine in terms of hybridization characteristics. For example, the primers may not hybridize with the aptamer molecule before the transformation. If the aptamer molecule contains unmethylated cytosine, the primers used in step (iii) contain adenine bases instead of guanine bases at the position where they pair with the unmethylated cytosine of the original aptamer molecule. The aptamer molecule may contain methylated or unmethylated cytosine. Optionally, to prevent specific portions of the sequence of the aptamer molecule, hairpin sequence, and / or barcode sequence from changing after treatment with the reagent, the sequence of the first aptamer molecule and preferably the combined sequence within the aptamer sequence may contain modified cytosine resistant to treatment with the reagent that allows for the conversion of unmethylated cytosine to bases that are detectably different from cytosine in terms of hybridization characteristics. Aptamers can contain unmethylated (unmodified) cytosine.
[0365] It is worth noting that aliquots of the same sample can be amplified under different conditions in a manner that allows any bias (TA or CG bias) to be assessed and accounted for during the analysis phase.
[0366] Optionally, the DNA aptamer is incorporated with a base labeled with a second member of the binding pair, which allows the original DNA template to be recovered after steps (i), (ii), (iii), or (iv).
[0367] Optionally, the DNA molecules obtained in step (i) or step (ii) of the method of the present invention may be recovered from the reaction mixture obtained after step (i), step (ii), or, depending on the specific circumstances, after step (iii) and / or after step (iv), preferably by using the first member of the binding pair, wherein the aptamer is modified with the second member of the binding pair, as described above.
[0368] Optionally, the DNA molecules obtained in step (iii) and / or after step (iv) are recovered from the reaction mixture, preferably by using the first member of the binding pair, wherein the primers used in step (iii) and / or after step (iv) are modified with the second member of the binding pair.
[0369] Combination barcodes allow for sample identification, multiplexing, pairing, and quantitative analysis, and can be used to monitor bias and control errors during amplification steps. The constructs obtained in this case have two distinct identifiers that bind to the sense strand, and the same two identifiers that bind to the antisense strand. Because these two combined sequences are identical for both the sense and antisense strands, the two strands can be tracked during the process. After the entire process, the complementary strands that were initially together will share the same two combined sequences. The two unique identifiers bind to each individual molecule in the starting sample, thus giving the ability to distinguish constructs. The unique identifiers allow for the identification of specific constructs containing the identifiers and their progeny, because after step (ii), the original sense and antisense strands of the double-stranded DNA molecule can be separated, but each of these strands will contain both combined sequences. Therefore, it is assumed that any amplification products of the original individual molecules carrying the two unique identifiers are phylogenetically identical.
[0370] This provides the advantages of being able to identify samples used as DNA templates, preserve them during the process, and recover, store them, and perform multiple amplifications under different conditions and sequencing without depleting the sample.
[0371] The terms “double-stranded DNA molecular population”, “combination sequence”, “end”, “compatible”, “ligation”, “template”, “primer”, “complement”, “amplification”, “binding pair”, “first member of binding pair”, “second member of binding pair”, “bases that are detectably different from cytosine in terms of hybridization characteristics”, and “modified cytosine” are defined in the context of the method of this invention.
[0372] All implementation methods and definitions used in the methods of this invention are applicable to this example.
[0373] 2. The method of the present invention, wherein the pairing of strands of multiple double-stranded DNA molecules is achieved by the presence of a barcode sequence.
[0374] For example, multiple paired, adaptor-modified DNA molecules of step (i) of the method of the present invention can be obtained by the following:
[0375] (a) A plurality of double-stranded DNA molecules are fragmented under conditions suitable for generating a plurality of double-stranded DNA molecule fragments with overhanging ends, wherein each end of each fragment is bound to a hemiaptamer molecule comprising a first DNA strand and optionally a second DNA strand, wherein each hemiaptamer is distinguishable from other hemiaptamers by a combination sequence located in a 3' region of the first DNA strand, wherein the second strand forms a double-stranded region with the first strand via complementarity with the central region of the first strand, and wherein the hemiaptamer molecule binds to a double-stranded DNA molecule fragment located between the 3' end of the first strand of the hemiaptamer and the overhanging end of the fragment of the double-stranded DNA molecule.
[0376] (b) Adding an alternative second strand or replacing the second DNA strand of a hemi-aptamer molecule with an alternative second strand, wherein the 5' region of the alternative second strand is complementary to the 3' region of the first strand of the hemi-aptamer molecule, and wherein the alternative second strand contains a region not complementary to the first strand of the hemi-aptamer molecule, thereby producing a plurality of Y-aptamer-modified DNA molecules.
[0377] (c) Optionally, fill the gap between the 5' end of the alternative second strand of the Y-aptamer and the 3' end of each DNA fragment.
[0378] In this case, the multiple paired adaptor-modified DNA molecules obtained after step (i) of the method of the present invention are suitable for different fragmentation systems.
[0379] In the first step, multiple double-stranded DNA molecules are fragmented under conditions suitable for generating a population of double-stranded DNA molecule fragments with overhanging ends, wherein each end of each fragment is bound to a hemiaptamer molecule located between the 3' end of the first strand of the hemiaptamer and the overhanging end of the double-stranded DNA molecule fragment.
[0380] As used herein, the term "hemaptamer" refers to an incomplete aptamer formed of a first DNA strand and optionally a second DNA strand, wherein the second strand forms a double-stranded region with the first strand via complementarity with the central region of the first strand, and wherein each hemiaptamer is distinguishable from other hemiaptamers by a combination of sequences located within the 3' region of the first DNA strand. In one embodiment, a hemiaptamer does not contain a second DNA strand. A hemiaptamer may contain both a first DNA strand and a second DNA strand.
[0381] Preferably, the second strand of the semi-aptamer does not show any substantial overlap with the combined sequence.
[0382] Preferably, the double-stranded DNA molecule used in step (i) is a genomic DNA fragment.
[0383] Optionally, the double-stranded DNA molecules used in step (i) are end-repaired before step (i), preferably further including a step of adding dA tails to the DNA molecules after the end-repair step.
[0384] Optionally, the population of double-stranded DNA molecules is pretreated with aptamers in step (i) under conditions suitable for the ligation of aptamer molecules to DNA molecules, thereby introducing sticky ends into the DNA molecules.
[0385] The first step, which includes fragmentation and ligation of hemiaptamers to double-stranded DNA molecules, can be performed, for example, by in vitro transposition, in which transposable elements are introduced from donor DNA (hemiaptamer molecules) into target DNA (a population of double-stranded DNA molecules).
[0386] Preferably, the fragmentation step (i) is carried out by a method comprising contacting a population of double-stranded DNA molecules with a transposase dimer loaded with double-stranded aptamer molecules, wherein the aptamer molecules comprise a double-stranded region comprising a Tn5 inverted repeat and a 5' overhang of one of the strands, wherein optionally, cytosine nucleotides in the double-stranded region that do not form part of the Tn5 inverted repeat and cytosine nucleotides in the single-stranded region are methylated, and wherein the contact is carried out under conditions suitable for DNA fragmentation and attachment of hemiaptamer molecules to both ends of each DNA fragment.
[0387] As used herein, the term "double-stranded aptamer molecule" refers to an aptamer molecule containing a double-stranded region comprising a Tn5 reverse repeat and a 5' overhang of one of the strands.
[0388] The result of the steps is a plurality of DNA molecules containing hemiaptamers, which are double-stranded DNA molecules having one hemiaptamer attached to each end of the molecule. Each DNA molecule containing hemiaptamers will have two distinct combined sequences, each of which is located in the 3' region of the first DNA strand of the hemiaptamer attached to each end of the DNA molecule.
[0389] Step 2(b) involves adding an alternative second strand or replacing the second DNA strand of a semi-aptamer molecule with an alternative second strand to obtain multiple DNA molecules containing Y-aptamers.
[0390] As used herein, the term "alternative second strand" refers to a strand having a 5' region complementary to the 3' region of the first strand of the hemimeric molecule and containing regions not complementary to the first strand of the hemimeric molecule. Optionally, the 5' region of the alternative second strand complementary to the 3' region of the first strand of the hemimeric molecule has no gap between the 5' end of the alternative second strand of the hemimeric molecule and the 3' end of the DNA fragment. In another embodiment, a gap exists between the 5' end of the alternative second strand of the hemimeric molecule and the 3' end of the DNA fragment.
[0391] Preferably, the alternative second strand does not show any substantial overlap with the combinatorial region of the first DNA strand of the semi-aptamer.
[0392] As used herein, the term "Y-aptamer" refers to an aptamer formed by two DNA strands, wherein the 3' region and / or central region of the first DNA strand and the 5' region of the alternative second strand form a double-stranded region through sequence complementarity, and wherein the 5' region of the first strand and the 3' region of the alternative second strand are not complementary.
[0393] In some implementations, the 5' end of the alternative second strand of the Y-aptamer and the 3' end of each DNA fragment are not connected because gaps can exist between the ends.
[0394] The constructs obtained after step (b) form the double-stranded libraries of this invention and can be used for sequencing or other conventional molecular biology techniques. The advantage of these libraries is that the combined sequences allow for cross-referencing of sequence information obtained from the initially combined sense and antisense strands to obtain more reliable results.
[0395] Optionally, step (c) involves filling the gap between the 5' end of the alternative second strand of the Y-aptamer and the 3' end of each DNA fragment.
[0396] This step generates a complementary copy of the combined sequence of the first DNA strand of the semi-aptamer by filling in the gaps.
[0397] Subsequently, the resulting multiple paired, adaptor-modified DNA molecules continue to proceed with steps (ii) and (iii) (and optionally, step (iv)) of the method of the present invention, as described above.
[0398] Optionally, the Y-aptamer incorporates a base labeled with a second member of the binding pair, which allows for the recovery of the original DNA template. The Y-aptamer may include methylated cytosine and / or unmethylated cytosine. The expressions “double-stranded DNA molecular population,” “combination sequence,” “end,” “compatible,” “ligation,” “template,” “primer,” “complementary,” “amplification,” “binding pair,” “first member of the binding pair,” “second member of the binding pair,” and “bases that are detectably different from cytosine in terms of hybridization characteristics” and “modified cytosine” have been defined above.
[0399] The terms “transposase,” “transposase dimer,” “loading,” “Tn5 reverse repeat,” “conditions suitable for DNA fragmentation and attachment of semi-aptamer molecules to both ends of each DNA fragment,” “displacement,” “gap,” “gap filling,” and “no substantial overlap” have been defined above.
[0400] The terms “DNA library” and “double-stranded DNA library” have been defined above.
[0401] Sequencing steps (The method of the present invention) Step (v) )
[0402] The paired DNA molecules (also known as double-stranded DNA libraries or DNA libraries) generated in steps (ii), (iii) and / or (iv) of the method of the present invention are suitable for sequencing technology (preferably sequencing of the paired DNA molecules generated in steps (iii) and / or (iv).
[0403] The design of the DNA library generated in the method of the present invention allows for better monitoring of any biases that may arise during previous steps and for detecting sequencing and transformation errors than currently used methods.
[0404] When amplification occurs, some fragments may be selectively amplified for various reasons. This undesirable effect is a major problem for quantification purposes, which are crucial in many sequencing applications, particularly for analyzing the methylation status of DNA (because each allele in each cell can have a different methylation status, and even samples can have heterogeneous compositions, making quantification and bias control essential for most applications).
[0405] The double-stranded DNA library generated in the method of this invention has the advantage of simultaneously reading both strands of each double-stranded DNA molecule during the sequencing process. This dual reading improves the confidence of the method because it allows for the detection and correction of systematic potential sequencing errors arising in each individual sequence read.
[0406] Currently, sequencers have an error rate that requires assumptions. Most of these errors cannot be described and remain hidden in the final results. This negatively impacts subsequent processing and analysis of the results. The method of this invention provides up to four sources of information for each nucleotide (given the upper and lower strands of dsDNA and, depending on the specific case, their respective synthetic complementary strands), thereby allowing verification of each nucleotide read, since all reads must be consistent. Therefore, the method of this invention allows for the detection and even correction of errors in sequencing (both for primary sequencing and for cytosine methylation analysis).
[0407] Sequencing errors are detected only when they occur in one of the strands. If this error cannot be confused with asymmetric methylation and the nucleotide detection has high confidence and is distinguishable by a reference genome sequence, then it can be correctable.
[0408] Genetic variants (i.e., mutations and SNPs) should not be confused with sequencing errors because the change must occur in both strands that are considered actual gene variants. In fact, the dual detection of the method of this invention validates the detection of changes. Furthermore, because primary sequence and methylation information are evaluated simultaneously, thymine from unmethylated cytosine can be distinguished from those from mutations or SNPs.
[0409] Additionally, when the aptamer contains a composite sequence (barcode sequence), it is possible to track amplification bias and count the unique double-stranded DNA molecules initially present in the sample. This can also be quantified during the sequencing method of this invention.
[0410] Double-stranded DNA libraries can be used in any conventional sequencing method, including next-generation sequencing (NGS). The method for generating double-stranded DNA libraries according to the present invention can be integrated into current and upcoming DNA sequencing pipelines, i.e., NGS technology, etc. NGS of the library can be performed using most available platforms. Paired-end sequencing can also be used, but it is not required. Paired-end sequencing will allow the study of longer fragments. Even when sequencing does not cover the entire molecule, information based on the pairing of the two strands can be obtained, provided that the sequenced region contains complementary regions and barcodes (where necessary) (see, for example...). Figure 13 ).
[0411] Alternatively, locus-specific sequencing can be performed.
[0412] The libraries obtained by the method according to the invention can be used for sequencing directly from the reaction mixture, wherein they have already been obtained or can be purified prior to the sequencing process.
[0413] In one aspect, the present invention relates to a method for determining the sequence of a population of double-stranded DNA molecules, comprising generating a library from the population of double-stranded DNA molecules using the method of the present invention in any embodiment thereof, and sequencing the DNA molecules obtained in steps ((ii)), (iii) of the method of the present invention, or, depending on the specific circumstances, step (iv) (step (v) of the method of the present invention).
[0414] The term "sequencing," or expressions such as "sequencing determination" or "sequencing," refers to the determination of information related to the nucleotide base sequence of nucleic acids, particularly involving the determination and ordering of multiple consecutive nucleotides within nucleic acids. This information may include partial or complete sequence information identifying or determining the nucleic acid. The information refers to the primary sequence of a double-stranded DNA library, epigenetic modifications (e.g., methylation or hydroxymethylation), or both the primary sequence and epigenetic modifications of a double-stranded DNA library. Sequence information can be determined with varying degrees of statistical reliability or confidence. As explained above, the method of this invention allows for high-confidence sequencing of double-stranded DNA molecules.
[0415] The method of this invention can be further used to sequence the primary sequence of a double-stranded DNA library.
[0416] Primary sequence determination includes detecting mutations or gene variants, such as polymorphisms (SNPs, etc.).
[0417] Preferably, the method of the present invention allows for the simultaneous determination of both the primary sequence and cytosine methylation within the same read. By analyzing the sequencing output, each read will contain information on the primary sequence (including mutations and SNPs) and methylation sequences of both strands, as well as information on the combined sequences contained in the aptamers.
[0418] Identification of methylated cytosine
[0419] The method of this invention allows for the identification of methylated cytosine in a population of double-stranded DNA molecules.
[0420] The double-stranded DNA molecules (also known as double-stranded DNA libraries) produced by the method of the present invention (step (iii) or step (iv)) are suitable for the identification of methylated cytosine, as described above.
[0421] If cytosine is present in one of the chains and guanine is present at the corresponding position in the opposite chain, the presence of methylated cytosine at the given position is determined; or if uracil or thymine is present in one of the chains and guanine is present at the corresponding position in the opposite chain, the presence of unmethylated cytosine at the given position is determined.
[0422] As used in this article, the term "strand" refers to each strand of a double-stranded DNA molecule. If pairing of the two strands occurs in the presence of a hairpin molecule, the two strands of the DNA molecule are linked together to form a unique molecule. If pairing occurs in the absence of a hairpin molecule, the strands of the DNA molecule reside in separate molecules; that is, they are not physically linked together, but can be identified by combinatorial sequencing.
[0423] As used herein, the term "opposite strand" in reference to the first strand can refer to the strand complementary to the first strand before treatment with a reagent that allows the conversion of unmethylated cytosine into bases that are detectably different from cytosine in terms of hybridization properties. For example, the antisense strand of a DNA molecule is the opposite strand of the sense strand. The term "opposite strand" can be broader than the term "complementary strand" because the complementarity between strands can be partially or completely lost after treatment with a reagent.
[0424] As used herein, the term "corresponding position" refers to the same position in the opposite strand (i.e., the nucleotide position that pairs with a given nucleotide position in the opposite strand, but the pairing need not be complementary).
[0425] The terms “double-stranded DNA molecular population”, “library”, “aptamer-modified DNA molecule”, “reagent”, “bases that are detectable to a different degree from cytosine in terms of hybridization characteristics”, “primer”, and “specificity” have been defined above.
[0426] The term "sequencing" has been defined in the context of the sequencing method of this invention. Therefore, the sequencing step (v) of the method of this invention (for identifying methylated cytosine) allows for the simultaneous acquisition of primary sequence and methylation information.
[0427] Preferably, the reagent that allows the conversion of unmethylated cytosine into a base that differs from cytosine in detectable order in terms of hybridization properties is a bisulfite. Preferably, the base that differs from cytosine in detectable order is thymine or uracil, more preferably uracil.
[0428] All unmethylated cytosine in the molecules obtained in step (i) of the method of the present invention are converted to uracil by treatment with a reagent (e.g., bisulfite) that allows the conversion of unmethylated cytosine to a base that is detectably different from cytosine in terms of hybridization properties. When the treated molecules are carried out in step (iii) of the method of the present invention, the synthesized complementary DNA molecule will have adenine at the position where uracil is present in the template molecule. The molecule (optionally amplified in step (iv) of the method) will have thymine at the position where uracil is present in the first molecule (provided that thymine, rather than uracil, is used for amplification). Thus, at the end of the process, unmethylated cytosine will be read as either uracil or thymine.
[0429] After elongation or PCR amplification, uracil is amplified to thymine, while the 5-methylcytosine residue remains cytosine, thus allowing differentiation between methylated and unmethylated cytosine in the original double-stranded DNA molecule. The presence of cytosine at a given position in the original double-stranded DNA molecule can be inferred from the presence of guanine at the corresponding position in the opposite strand during molecular sequencing of the double-stranded DNA library. Similarly, the presence of uracil or thymine at a given position in the original double-stranded DNA molecule can be inferred from the presence of guanine at the corresponding position in the opposite strand during molecular sequencing of the double-stranded DNA library.
[0430] When using the paired double-stranded DNA molecules (also known as DNA libraries) obtained in step (iii) or (iv) of the method of the present invention, each read has a reagent-transformed sequence of both strands of the double-stranded DNA molecule. If pairing is achieved by using barcodes, the initially complementary strand can be deduced from the combined sequence of each strand (barcode sequence or combined barcode).
[0431] Using this information, the original sequences of the two strands (before reagent transformation) can be deduced. Therefore, for each read, information on the primary sequences (for localization to a reference genome and assessment of polymorphism) and methylation status of the two strands of the original double-stranded DNA molecule is obtained. Combined labeling of the molecule allows for the assessment of the reagent-transformed sequences of the two original strands of the double-stranded DNA molecule.
[0432] Software that processes the output sequences performs a process that allows for the inference of the original double-stranded DNA molecules' sequences, obtaining the sequence of each molecule in the double-stranded DNA library before and after reagent treatment. Subsequently, the software cross-references information from each strand and its true complementary strand, taking into account that the true complementary strand is the one initially bound to the first strand. Information can be directly obtained from the library obtained by the method of this invention, where the two strands are physically linked. This information can also be obtained indirectly if the original strands are paired through sequence combination.
[0433] The developed software is capable of: (i) integrating primary sequence and methylation information, (ii) detecting mutations, SNPs, and CNVs (loss and gain), and (iii) detecting sequencing errors and biases.
[0434] Because it measures the methylation information of both strands of the original double-stranded DNA molecule, this allows for the assessment of hemimethylation or methylation symmetry in each strand.
[0435] The terms “hemimethylation” and “asymmetric methylation” are used interchangeably and refer to double-stranded DNA, such as sequences in CpG, where only one of the two strands is methylated.
[0436] y-aptamer libraries, methods for synthesizing them, and kits
[0437] The present invention also provides a method and kit for synthesizing Y-aptamers used in the method of the present invention. The present invention also provides a kit containing a Y-aptamer library.
[0438] The present invention also relates to a Y-aptamer library, wherein each Y-aptamer comprises a first DNA strand and a second DNA strand, wherein the 3' region of the first DNA strand and the 5' region of the second DNA strand form a double-stranded region through sequence complementarity, wherein the 3' region of the second DNA strand can form a hairpin loop through hybridization between a first segment and a second segment of the 3' region, the first segment being located at the 3' end of the 3' region, and the second segment being located near the region of the second DNA strand forming the double-stranded region with the 3' region of the first DNA strand, and wherein each member of the library is distinguishable from other members by one or more (preferably one) combination sequences located within the double-stranded region formed between the 3' region of the first DNA strand and the 5' region of the second DNA strand, and / or by one or more combination sequences located within the single-stranded region of the Y-aptamer. Preferably, the Y-aptamer comprises a barcode sequence in the double-stranded region of the Y-aptamer and a barcode sequence in the single-stranded region of the Y-aptamer. Optionally, the Y-aptamer further comprises a "cleavage site" as described above. Preferably, the "cleavage site" is located in the single-stranded region of the Y-aptamer.
[0439] In a preferred embodiment of the Y-aptamer library, the combined sequence contains one or more modified cytosines resistant to treatment with a reagent that allows the conversion of unmethylated cytosines into bases that are detectably different from cytosines in terms of hybridization properties.
[0440] In another embodiment of the Y-aptamer library, the 3' end of the second polynucleotide in each aptamer is reversibly closed.
[0441] As used herein, the term "reversible blocking" refers to, but is not limited to, reversible modification with a reversible blocking group that prevents polynucleotide initiation of synthesis. Because the modification is reversible, the blocking group can be removed in a further step.
[0442] As used herein, the term “reversible clogging group” refers to a group that replaces the 3'OH- group at the 3' end of a polynucleotide to prevent polymerase elongation. Exemplary reversible clogging groups are, but are not limited to, 3'-p; 3'-amino; 3'-reverse 3-3'-deoxynucleotide (idN), inosine, 3'-O ethers, such as 3'-O-allyl (Intelligent Bio-Systems), 3'-O-methoxymethyl, 3'-O-nitrobenzyl, and 3'-O-azidomethyl (Illumina / Solexa); 3'-aminoalkoxy and 3'-O-amino. Exemplary reversible clogging groups are also those disclosed in Gardner AF et al. 2012. Nucleic Acids Research, 40(15):7404-7415, and are referred to as Lightning Terminators. TM (Lasergen, Inc.). In the context of this invention, all dideoxynucleotides (ddCMP, ddAMP, ddTMP, ddGMP) can also be considered as "reversible blocking groups".
[0443] As used herein, the term "reversible blocking group" also includes a linker that connects the 5' end of the first DNA strand and the 3' end of the second DNA strand and is removable. The linker can be, but is not limited to, a typical nucleotide, a modified nucleotide, or a sequence of typical or modified nucleotides. When the linker is removed, the first and second DNA strands of the aptamer are separated. The aptamer can be used to ligate to a population of double-stranded DNA molecules in the method of the present invention, before or after the linker has been removed. The linker can be attached to a support.
[0444] Therefore, in another embodiment of the Y-aptamer library, the 3' end of the second polynucleotide in each aptamer is reversibly blocked by a linker that connects the 5' end of the first DNA strand and the 3' end of the second DNA strand.
[0445] In another embodiment of the Y-aptamer library, the terminal segment of the double-stranded region formed between the 3' region of the first DNA strand and the 5' region of the second DNA strand in each aptamer contains a target site for a restriction endonuclease.
[0446] The term "restriction endonuclease" refers to an enzyme that cuts DNA at or near a specific recognition nucleotide sequence (the sequence of oligonucleotides recognized by the restriction enzyme) called a restriction site or target site. Exemplary restriction endonucleases are well known in the art. Restriction endonucleases include, but are not limited to, type I, type II, type II, type III, and type IV enzymes. The REBASE database provides a comprehensive database of information on restriction endonucleases (Roberts RJ et al. 2010. Nucleic Acids Research, 38: D234-D236).
[0447] As used in this article, the term "target site" refers to a nucleotide sequence that is specifically recognized by a restriction endonuclease.
[0448] In another aspect, the present invention relates to a method for generating DNA Y-aptamers, the method comprising the steps of:
[0449] (i) Contacting a first single-stranded polynucleotide with a second single-stranded polynucleotide, wherein at least a portion of the 3' region of the first single-stranded polynucleotide is complementary to the 5' region of the second polynucleotide.
[0450] The 3' region of the second single-stranded polynucleotide forms a hairpin loop through the hybridization of a first segment and a second segment of the 3' region. The first segment is located at the 3' end of the 3' region of the second polynucleotide, and the second segment is located near the 5' region of the second polynucleotide. The 3' end of the second polynucleotide is reversibly closed.
[0451] The contact is carried out under conditions suitable for hybridization of complementary regions within the 3' region of the first single-stranded polynucleotide and the 5' region of the second single-stranded polynucleotide, thereby producing a double-stranded DNA molecule.
[0452] (ii) Extending the 3' end of the first polynucleotide to generate a sequence complementary to the 5' region of the second polynucleotide within the first single-stranded polynucleotide, and
[0453] (iii) Optionally, unblock the 3' end of the second polynucleotide.
[0454] Step (i) is performed under conditions suitable for hybridization of complementary regions within the 3' region of the first single-stranded polynucleotide and the 5' region of a member of the second single-stranded polynucleotide group. The conditions suitable for hybridization are defined in the context of the first method of this invention.
[0455] Step (ii) involves extension. The conditions suitable for extension are well known to those skilled in the art.
[0456] Step (iii) involves unblocking the 3' end of the second polynucleotide. The term "unblocking" refers to the removal of the blocking group from the 3' end of the polynucleotide and the restoration of the 3'-OH group. Suitable conditions for unblocking are those that neither break the duplex nor damage the DNA, and depend on the specific blocking group to be removed. For example, 3'-O-allyl groups can be cleaved by transition metal catalysis, 3'-O-methoxymethyl groups by acid cleavage, 3'-O-nitrobenzyl groups by photo-cleavage, and 3'-oazidomethylene groups by phosphine cleavage. When the reversible blocking group is a dideoxynucleotide, the entire dideoxynucleotide is removed from the 3' end of the second polynucleotide without breaking the duplex (i.e., the pairing between complementary regions within the 3' region of the first single-stranded polynucleotide and the 5' region of the second single-stranded polynucleotide is maintained). The term "unblocking" also refers to the removal of the linker connecting the 5' end of the first single-stranded polynucleotide and the 3' end of the second single-stranded polynucleotide.
[0457] In a preferred embodiment of the method for generating DNA Y-aptamers, the extension carried out in step (ii) is performed in the presence of a modified cytosine treated with a reagent that allows the conversion of unmethylated cytosine into bases that are detectably different from cytosine in terms of hybridization properties.
[0458] In another embodiment, the first single-stranded polynucleotide and / or the second single-stranded polynucleotide used in step (i) also contain modified cytosine.
[0459] In another preferred embodiment, the method for generating DNA Y-aptamers according to the third method of the invention further comprises treating the Y-aptamers with a reagent that allows the conversion of unmethylated cytosine into bases that are detectably different from cytosine in terms of hybridization properties.
[0460] In another preferred embodiment, at least one position within the 3' end of the first polynucleotide is guanine, and wherein the position within the 5' region of the second polynucleotide that hybridizes with said one or more positions within the 3' end of the first polynucleotide is cytosine or methylcytosine.
[0461] In another preferred embodiment, the 5' end of the second single-stranded polynucleotide contains a sequence that is copied onto the complementary strand when the complementary strand is formed during the elongation step (ii) to form the target site of the restriction endonuclease.
[0462] The second single-stranded polynucleotide contains a sequence that may or may not be a palindromic sequence. A palindromic sequence is a nucleic acid sequence that is identical whether read from 5' to 3' on one strand or from 5' to 3' on the complementary strand.
[0463] The phrase "copied onto the complementary strand" means that the synthesized complementary strand contains a complementary and antiparallel sequence of a second single-stranded polynucleotide.
[0464] The Y-adaptor may contain a composite barcode that allows tracking of the product of the method of the present invention.
[0465] A second single-stranded polynucleotide is provided in the form of a nucleotide library, wherein each member of the library is distinguishable from other members by a combinatorial sequence (also known as a barcode sequence or combinatorial barcode) located within the 5' region of the polynucleotide, and wherein the combinatorial sequence is located upstream of a region showing sequence complementarity with the first single-stranded polynucleotide, thereby producing a DNA Y-aptamer molecular library. In a more preferred embodiment, the combinatorial sequence contains one or more cytosines resistant to treatment with a reagent that allows the conversion of unmethylated cytosines into bases that are detectably different from cytosine in terms of hybridization properties. Figure 15 An embodiment of a method for generating DNA Y-aptamers containing combined sequences according to a third method of the present invention is shown.
[0466] A composite sequence is formed from a number of degenerate nucleotides (each degenerate nucleotide is a mixture of two or more nucleotides). If X is the number of different degenerate nucleotides in the composite sequence, and Y is the length of the nucleotides in the composite sequence, then the number of different aptamers that can be obtained is X. Y For example, if the combined sequence is 4 nucleotides long and can be A, T, or G, then the number of different aptamers is 3. 4 =81, and the number of combinations of two aptamers in a DNA molecule containing a Y-aptamer is 81. 2 =6561. If the combined sequence is 5 nucleotides long and can be A, T or G, then the number of different aptamers is 243, and the number of combinations of two aptamers in a DNA molecule containing a Y-aptamer is 59049.
[0467] As used in this article, the term "upstream" refers to the region facing the 5' end of the chain.
[0468] In another embodiment, a first single-stranded polynucleotide and / or a second single-stranded polynucleotide are provided immobilized in a support, wherein the immobilization is performed by binding a nucleotide at the 5' end of the first single-stranded polynucleotide or a hairpin loop of the second single-stranded polynucleotide to the support. Preferably, the binding is covalent.
[0469] In another embodiment, the first and second single-stranded polynucleotides are linked via a linker at the 5' end of the first single-stranded polynucleotide and the 3' end of the second single-stranded polynucleotide, and the linker is fixed in a support. The binding of the linker to the support facilitates the elongation step. The binding between the linker and the support can be broken after the synthesis of the Y-aptamer has been completed to release the aptamer. The aptamer linked to the support via the linker can also be used to link molecules.
[0470] The present invention also provides a kit containing polynucleotides for obtaining Y-aptamers of the method of the present invention.
[0471] In another aspect, the present invention relates to a kit comprising:
[0472] (i) A first single-stranded polynucleotide containing the 5' and 3' regions
[0473] (ii) A second polynucleotide comprising a 5' region and a 3' region, wherein the 3' region forms a hairpin loop through hybridization between a first segment and a second segment within the 3' region, the first segment being located at the 3' end of the 3' region, and the second segment being located near the 5' region, and wherein the 3' end of the second polynucleotide is reversibly closed.
[0474] The 3' region of the first single-stranded polynucleotide is complementary to at least a portion of the 5' region of the second polynucleotide.
[0475] As used herein, the term "reagent kit" refers to a combination of two or more elements or components, including other types of biochemical reagents, containers, packaging, such as packaging intended for commercial-scale use, electronic hardware components, etc.
[0476] Suitable kits contain various reagents used according to the invention in suitable containers and packaging materials, including tubes, vials, and shrink-wrapped and blow-molded packages. Additionally, the kits of the invention may contain instructions regarding the simultaneous, sequential, or separate use of the different components of the kit. These instructions may be in the form of printed material or in the form of an electronic support capable of storing the instructions so that they can be read by an object, such as electronic storage media (disks, magnetic tapes, etc.), optical media (CD-ROMs, DVDs), etc. Alternatively / or, the media may contain an Internet address providing the instructions.
[0477] In a preferred embodiment of the kit of the present invention, the second polynucleotide is provided in the form of a polynucleotide library, wherein each member is distinguishable from other members by a combinatorial sequence located within the 5' region of the second polynucleotide and upstream of a region showing sequence complementarity with the first single-stranded polynucleotide. In a more preferred embodiment, the combinatorial sequence contains one or more modified cytosines resistant to treatment with a reagent that allows the conversion of unmethylated cytosines into bases that are detectably different from cytosines in terms of hybridization properties.
[0478] In another embodiment, the first single-stranded polynucleotide and / or the second polynucleotide also contain modified cytosine.
[0479] In another embodiment, the 5' region of the second polynucleotide contains a sequence that generates a target site for a restriction endonuclease when converted into a double-stranded region.
[0480] Optionally, the kit may contain one or more additional components.
[0481] In another embodiment, the kit further comprises one or more components selected from:
[0482] (i) DNA polymerase,
[0483] (ii) One or more nucleotides selected from A, G, C and T,
[0484] (iii) One or more modified cytosines resistant to treatment with a reagent that allows the conversion of unmethylated cytosines into bases that are detectably different from cytosine in terms of hybridization properties.
[0485] (iv) A reagent capable of removing the blocking group from the 3' end of the second polynucleotide.
[0486] (v) A reagent that allows the conversion of unmethylated cytosine into bases that are detectably different from cytosine in terms of hybridization properties, and
[0487] (vi) A target site-specific restriction endonuclease, wherein the target site is formed by a sequence within the 5' end of a second polynucleotide.
[0488] As used herein, the term "DNA polymerase" refers to an enzyme that synthesizes DNA strands de novo by using a nucleic acid strand as a template and adding only free nucleotides to the 3' hydroxyl terminus of the newly formed strand. This results in the new strand elongating in the 5'-3' direction. DNA polymerases can be naturally occurring DNA polymerases or variants of natural enzymes having the activities described above.
[0489] The nucleotides provided in the kit may be modified nucleotides. Examples of modified nucleotides are modified cytosines, such as methylcytosine and hydroxymethylcytosine.
[0490] As used herein, the expression "a reagent capable of removing the blocking group from the 3' end of a second polynucleotide" refers to a reagent that unblocks the 3' end of a second polynucleotide. The reagent depends on the specific blocking group used. Suitable reagents (e.g., acids, phosphine, etc.) have been disclosed above.
[0491] Target-specific restriction endonucleases are restriction endonucleases that can specifically recognize target sequences and cleave at or near the target sequence.
[0492] In a preferred embodiment, one or more modified cytosines are selected from methylcytosine, hydroxymethylcytosine, and combinations thereof.
[0493] The present invention also provides a kit comprising a Y-aptamer library of the method of the present invention and other components.
[0494] In another aspect, the present invention relates to a kit comprising:
[0495] (i) A Y-aptamer library, wherein each Y-aptamer comprises a first DNA strand and a second DNA strand, wherein the 3' region of the first DNA strand and the 5' region of the second DNA strand form a double-stranded region through sequence complementarity, wherein the 3' region of the second DNA strand forms a hairpin loop through hybridization between a first segment and a second segment within the 3' region, the first segment being located at the 3' end of the 3' region, and the second segment being located near the second DNA strand region forming the double-stranded region with the 3' region of the first DNA strand, and wherein each member of the library is distinguishable from other members by a combination of sequences within the double-stranded region formed between the 3' region of the first DNA strand and the 5' region of the second DNA strand.
[0496] (ii) one or more components, said components being selected from:
[0497] a) DNA polymerase,
[0498] b) One or more nucleotides selected from A, G, C, and T.
[0499] c) One or more modified cytosines resistant to treatment with a reagent that allows the conversion of unmethylated cytosine into bases that are detectably different from cytosine in terms of hybridization properties.
[0500] d) A reagent capable of removing the blocking group from the 3' end of the second polynucleotide.
[0501] e) A reagent that allows the conversion of unmethylated cytosine into bases with a detectable degree different from cytosine in terms of hybridization properties, and
[0502] f) A target-site specific restriction endonuclease, wherein the target site is formed by a sequence within the 5' end of a second polynucleotide.
[0503] All the specific embodiments previously disclosed regarding Y-aptamer libraries and other components of kits for obtaining Y-aptamers are applicable to kits incorporating the methods of the present invention for Y-aptamer libraries.
[0504] Other terms and expressions have already been defined.
[0505] A library containing double-stranded DNA aptamers with combined sequences and a method for obtaining said library.
[0506] The present invention also provides a method for synthesizing any library of double-stranded DNA aptamers having a combined sequence (“barcode sequence” or “combined barcode”), which can be used in the method of the present invention or in any other method requiring a combined barcode.
[0507] To synthesize the aptamer, a precursor consisting of two partially complementary oligonucleotides can be used, one of which carries the combinatorial region. This combinatorial region is single-stranded DNA in precursor form. After incubation of the precursor form with a suitable enzyme, the incomplete oligonucleotide is completed, and the combinatorial region becomes double-stranded DNA.
[0508] In one aspect, the present invention relates to a method for obtaining a double-stranded DNA aptamer library, wherein each aptamer comprises a first DNA strand and a second DNA strand, and wherein each aptamer is distinguishable from other aptamers by a combination of sequences located within a double-stranded region formed between the 3' region of the first DNA strand and the 5' region of the second DNA strand, the method comprising the following steps:
[0509] (i) Providing a population of single-stranded DNA molecules comprising a constant region and a combinatorial region, wherein the single-stranded DNA molecules are distinguishable from other molecules by a sequence in the combinatorial region, wherein the constant region is located at the 3' end relative to the combinatorial region, and wherein the 3' end is reversibly closed.
[0510] (ii) By using the single-stranded DNA molecule from step (i) as a template and using an extension primer that fully or partially hybridizes with the constant region of the single-stranded DNA molecule to generate double-stranded DNA, the combinatorial region is copied onto the newly generated strand, thereby generating a double-stranded combinatorial sequence.
[0511] As used in this article, the term "constant region" refers to a region in a single-stranded DNA molecule in which the sequence is identical for every member of a population of single-stranded DNA molecules.
[0512] As used in the context of this aspect of the invention, the term "reversible blocking" refers to, but is not limited to, reversible modification with a reversible blocking agent that prevents polynucleotide-initiated synthesis. Because the modification is reversible, the blocking group can be removed in a further step. Exemplary reversible blocking groups have been disclosed in the context of "Y-aptamer libraries of the present invention, methods for synthesizing them, and kits."
[0513] The term "reversible closure" also includes a linker that connects the 5' end of the first DNA strand to the 3' end of the second DNA strand and can be removed.
[0514] The term "reversible closure" also encompasses immobilizing a single-stranded DNA molecule within a support by binding its 3' end to the support. Because the closure is reversible, the connection between the 3' end of the single-stranded DNA molecule and the support can be disrupted in a further step.
[0515] As used herein, the expression "complete or partial hybridization with a constant region of a single-stranded DNA molecule" means extending the primer's overall or partial non-covalent binding to form a stable double-stranded polynucleotide with the constant region of a single-stranded DNA molecule. The expressions "hybridization" and "hybridization conditions" have already been defined in the context of the first method of this invention.
[0516] When 100% of the extended primer hybridizes to the constant region of the single-stranded DNA molecule, the primer completely hybridizes to the region. When less than 100% of the extended primer hybridizes to the constant region of the single-stranded DNA molecule, the primer partially hybridizes to the region. Preferably, at least 0.1%, at least 0.5%, at least 1%, at least 2%, at least 5%, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, and at least 95% of the extended primer does not hybridize to the constant region of the single-stranded DNA molecule.
[0517] As used herein, the term "combinatorial region" refers to a variable region in a single-stranded DNA molecule, where the sequence is different for each member of a population of single-stranded DNA molecules. The term combinatorial sequence has been previously defined.
[0518] As used in this article, the phrase "copying the combinatorial region onto the newly generated strand" means that the newly generated strand contains a complementary and antiparallel sequence to the combinatorial region of a single-stranded DNA molecule.
[0519] In a preferred embodiment, the method further includes removing the blocking group from the 3' end of the single-stranded DNA molecule.
[0520] In another embodiment, the extended primer includes a 5' region of a protrusion that does not hybridize with the constant region of a single-stranded DNA molecule.
[0521] In another embodiment, the constant region of the single-stranded DNA molecule includes a 3' region of a protrusion that does not hybridize with an elongation primer.
[0522] In another embodiment, a hairpin loop is formed in the constant region of a single-stranded DNA molecule through hybridization between a first segment and a second segment within the constant region.
[0523] Combinatorial sequences can be generated separately and then added to another aptamer to produce combinatorial Y-aptamers. Therefore, library aptamers can be used as modules that can be attached to other incomplete aptamers to obtain more complex aptamers. For example, combinatorial aptamers of a library formed by two complementary chains can be joined with non-combinatorial Y-aptamers to obtain a set of Y-aptamers with combinatorial regions.
[0524] In another embodiment, the method further includes linking aptamers of the library to a second DNA molecule having a double-stranded region whose ends are compatible with the ends of the aptamer molecule. Preferably, the second DNA molecule contains overhanging regions in the 5' region of the first strand and / or the 3' region of the second strand, the overhanging regions not hybridizing with each other. In a more preferred embodiment, the 3' overhanging region in the second strand forms a hairpin loop through hybridization between a first segment and a second segment within the region.
[0525] In another embodiment, each single-stranded DNA molecule of step (i) is provided immobilized in a support. In a more preferred embodiment, the immobilization is performed by binding the 5' end of the single-stranded DNA molecule to the support, preferably by covalent binding. In another embodiment, the immobilization is performed by binding the 3' end of the single-stranded DNA molecule to the support, preferably by covalent binding. In another preferred embodiment, if a hairpin loop is formed by hybridization between a first segment and a second segment within the constant region of the single-stranded DNA molecule, the immobilization is performed by binding the nucleotides of the hairpin loop of the single-stranded DNA molecule to the support, preferably by covalent binding.
[0526] In another embodiment, the first and second DNA strands are connected via a linker between the 5' end of the first DNA strand and the 3' end of the second DNA strand, and the linker is fixed in a support. The binding of the linker to the support facilitates the elongation step. The binding between the linker and the support can be broken after aptamer synthesis has been completed to release the aptamer. Aptamers connected to the support via linkers can also be used for molecule linking.
[0527] The present invention also relates to a double-stranded DNA aptamer molecular library obtained by the method described above.
[0528] In another aspect, the present invention relates to a library of double-stranded DNA aptamer molecules, wherein each DNA aptamer molecule comprises a constant region and a variable region, wherein each double-stranded DNA aptamer comprises a first DNA strand and a second DNA strand, and wherein each aptamer is distinguishable from other aptamers by a combination of sequences in a variable region within a double-stranded region formed between a 3' region of the first DNA strand and a 5' region of the second DNA strand.
[0529] In a preferred embodiment, the 3' end of at least one chain is reversibly closed.
[0530] In another embodiment, one or two chains include overhanging regions that do not hybridize with opposite chains. Preferably, a constant region of one of the chains forms a hairpin loop through hybridization between a first segment and a second segment within the constant region.
[0531] Figure 16 Different embodiments of the aspects described herein are shown, wherein a hairpin-containing Y-adaptor is obtained according to the third method of the invention, and Y-adaptors, half-adaptors and plain adaptors according to the fourth, sixth and seventh methods of the invention are obtained.
[0532] Although the foregoing invention has been described in considerable detail for clarity and understanding purposes, those skilled in the art will understand from reading this disclosure that various changes in form and detail may be made without departing from the true scope of the invention and the appended claims.
[0533] The invention is described below with reference to the following embodiments, which should be considered merely illustrative and in no way limit the scope of the invention. Example
[0534] Example 1: The procedure involves physically pairing two original strands and selectively filtering, converting to bisulfite, amplifying, and sequencing only the desired intermediate product.
[0535] Aptamer preparation. Hairpin aptamers were prepared from 10 μM oligonucleotide c15_Hairp01-5'P (SEQ ID NO:1). Double-stranded aptamers were prepared from 20 μM oligonucleotide c14_Hang04 5'P (SEQ ID NO:2) and 20 μM oligonucleotide c14_BIO05noB (SEQ ID NO:3). Aptamers were hybridized by initial denaturation in a thermal cycler followed by stepwise cooling (95°2'; 80°2'; 65°10'; 37°10'; 25°5'; 4°-').
[0536] Ligation process: Using 3 μM hairpin aptamers and 3 μM double-stranded aptamers, 4 pmol dsDNA fragments (30 meters) were mixed to a final concentration of 0.2 μM in a 20 μL final volume with T4 DNA ligase and T4 buffer, and incubated at 23 °C for 15 minutes. The reaction product was purified using a G50 column.
[0537] Ligation product capture. The correct reaction product is captured by a biotinylated oligonucleotide complementary to the hairpin aptamer: 3 μL of 20 μM c14_BIO04-5'B (SEQ ID NO:4) is added to 10 μL of the ligation reaction in the presence of SSC. The reaction tube is gently mixed and incubated in a thermal cycler (90°2'; 65°5'; 60°5'; 55°5'; 25°-'). 30 μL of M-270 beads is resuspended, and BW1x is prepared to a final volume of 90 μL following the manufacturer's instructions. The M-270 beads are added to the capture reaction and incubated at ambient temperature for 15 min. The reaction is washed, and the ligation product is released from the M-270 beads. The beads are resuspended in SCC1x and incubated at 95°C in a heating block for 2 min. The supernatant containing the ligation product of interest is recovered.
[0538] Bisulfite. Treat 50 μL of the ligation product capture with sodium bisulfite according to standard procedures or the manufacturer's instructions. Elute the final product in 30 μL of elution buffer.
[0539] Amplification. Using polymerases such as Zymotag or TurboPfu, 1 μL of the product from the previous step was amplified with primers c14_amp02F (SEQ ID NO:5) and c14_amp02R (SEQ ID NO:6). For a final reaction volume of 30 μL, 1 μL of 20 μM each primer was used. Twenty PCR cycles were performed (95° 2'; 62° 30"; 72° 1'). The product was purified using a G50 column and evaluated by PAGE electrophoresis.
[0540] Sequencing. The products obtained from the previous step were processed according to the Ion Torrent pipeline and visualized using IntegrativeGenomics Viewer.
[0541] Example 2: The procedure involves generating a complementary strand, physically pairing it with a template, and converting, amplifying, and sequencing the resulting intermediate product using bisulfite.
[0542] Aptamer preparation. Y-aptamers were prepared from 20 μM oligonucleotides c15_YA4 (SEQ ID NO:7) and c15_Hairp06 (SEQ ID NO:8). The aptamers were hybridized by initial denaturation in a thermal cycler followed by stepwise cooling (95°2'; 80°2'; 65°10'; 37°10'; 25°5'; 4°-').
[0543] Ligation procedure: Using 2.5 μM Y-aptamer, in the presence of T4 DNA ligase and T4 buffer, 4 pmol dsDNA fragments (30 meters) were mixed to a final concentration of 0.2 μM in a final volume of 20 μL and incubated at 23 °C for 15 minutes. The reaction product was purified using a G50 column.
[0544] Ligation product extension. The ligation product was extended by adding polymerase to the reaction with dNTPs in an optimal buffer solution at a reaction volume of 30 μL. The reaction tubes were gently mixed and incubated in a thermal cycler (25°2'; 72°10'; 4°-').
[0545] Bisulfite. Treat 20 μL of the ligation product capture with sodium bisulfite according to standard procedures or the manufacturer's instructions. Elute the final product in 20 μL of elution buffer.
[0546] Amplification. Using polymerases such as Zymotag or TurboPfu, 1 μL of the product from the previous step was amplified with primers c14_amp02F (SEQ ID NO:5) and c14_amp02R (SEQ ID NO:6). For a final reaction volume of 30 μL, 1 μL of 20 μM each primer was used. Twenty PCR cycles were performed (95° 2'; 62° 30"; 72° 1'). The product was purified using a G50 column and evaluated by PAGE electrophoresis.
[0547] Sequencing. The products obtained from the previous step were processed according to the Ion Torrent pipeline and visualized using IntegrativeGenomics Viewer.
[0548] Table I. Sequences of oligonucleotides used in embodiments of the present invention
[0549]
[0550]
[0551] pho: phosphate ester; bio: biotin
[0552] Item of this invention
[0553] The present invention provides the following items.
[0554] [1]. A method for generating a double-stranded DNA library from a population of double-stranded DNA molecules, the method comprising the following steps:
[0555] (i) Contact the population of double-stranded DNA molecules with the first aptamer molecule and the second aptamer molecule.
[0556] The first aptamer molecule is a double-stranded DNA molecule with an end that is compatible with the end of the double-stranded DNA.
[0557] The second aptamer molecule is a hairpin aptamer, which includes a hairpin loop region and a double-stranded region, wherein the double-stranded region contains ends compatible with the ends of the double-stranded DNA.
[0558] The contacting step is performed under conditions suitable for the ligation of the first aptamer molecule and / or the second aptamer molecule to the DNA molecule to produce a plurality of aptamer-modified DNA molecules.
[0559] (ii) From the population of aptamer-modified DNA molecules obtained in step (i), recover those molecules that contain the second aptamer molecule at one or both ends of the aptamer-modified DNA molecules.
[0560] (iii) Using the aptamer-modified DNA molecule obtained in step (ii) as a template and primers to synthesize a DNA strand under conditions allowing for strand synthesis, wherein the sequence of the primers is at least a portion complementary to the sequence of the first aptamer molecule, and
[0561] (iv) Optionally, the double-stranded DNA molecule obtained in step (iii) is amplified using primers whose sequences are complementary to at least a portion of the first aptamer region.
[0562] [2]. According to the method of [1], the recovery step (ii) is carried out using a polynucleotide, the polynucleotide comprising a sequence complementary to at least a portion of the sequence of the second aptamer and a purification tag.
[0563] [3]. The method according to [1] or [2], wherein the DNA molecules obtained in step (ii) are recovered from the reaction mixture obtained after step (iii) or, depending on the case, after step (iv).
[0564] [4]. The method according to [3], wherein the recovery from the reaction mixture is carried out using a first member of the binding pair, wherein the first aptamer and / or the second aptamer is modified with a second member of the binding pair.
[0565] [5]. A method for generating a double-stranded DNA library from a population of double-stranded DNA molecules, the method comprising the following steps:
[0566] (i) Contact the population of double-stranded DNA molecules with the first aptamer molecule and the second aptamer molecule.
[0567] The first aptamer molecule is a double-stranded DNA molecule with an end that is compatible with the end of the double-stranded DNA.
[0568] The second aptamer molecule is a hairpin aptamer comprising a hairpin loop region and a double-stranded region, wherein the double-stranded region contains ends compatible with the ends of the double-stranded DNA molecule.
[0569] The first aptamer molecule, or the second aptamer molecule, or both the first aptamer molecule and the second aptamer molecule, are provided in a manner immobilized in a support, wherein the immobilization is achieved by binding a nucleotide to the end of one of the chains of the first aptamer molecule or a hairpin loop of the second aptamer molecule to the support, and
[0570] The contacting step is performed under conditions suitable for the ligation of the first and / or second aptamer molecules with the DNA molecule to produce a plurality of aptamer-modified DNA molecules.
[0571] (ii) Using the aptamer-modified DNA molecule obtained in step (i) as a template and synthesizing a DNA strand using primers, wherein the sequence of the primers is at least a portion complementary to the sequence of the first aptamer molecule, and
[0572] (iii) Optionally, the double-stranded DNA molecule obtained in step (ii) is amplified using primers, wherein the sequence of the primers is complementary to at least a portion of the first aptamer region.
[0573] [6]. The method according to [1] to [5], wherein the double-stranded DNA molecule used in step (i) is a genomic DNA fragment.
[0574] [7]. The method according to [1] to [6], wherein the double-stranded DNA molecule used in step (i) is end-repaired prior to step (i).
[0575] [8]. According to the method of [7], the method further includes a step of adding dA tail to the DNA molecule after the end repair step.
[0576] [9]. According to the method described in [1] to [8], the first aptamer molecule and / or the second aptamer molecule are provided in the form of a first aptamer molecule library and a second aptamer molecule library, respectively, wherein each member in the library is distinguishable from other members by a combination sequence within the aptamer sequence.
[0577]
[10] . According to the method of [1] to [9], wherein the population of aptamer-modified DNA molecules is treated with a reagent before step (iii) if the library has been obtained by the method of [1] or before step (ii) if the library has been obtained by the method of [5], the reagent allowing the conversion of unmethylated cytosine to bases that are detectably different from cytosine in terms of hybridization properties, and wherein the primers used in steps (iii) and (iv) of [1] or, depending on the case, steps (ii) and (iii) of [5] are specific to the first aptamer molecules that have been treated with the reagent.
[0578]
[11] . According to the method of
[10] , the combined sequence within the aptamer sequence contains a modified cytosine resistant to treatment with a reagent that allows unmethylated cytosine to be converted into bases that are detectably different from cytosine in terms of hybridization properties.
[0579]
[12] . The method according to [1] to
[11] , wherein the DNA molecules obtained in step (iii) or, depending on the case, in step (iv) when the library has been obtained by the method according to [1] or in step (ii) when the library has been obtained by the method according to [5] or, depending on the case, in step (iii) are recovered from the reaction mixture.
[0580]
[13] . The method according to
[12] , wherein the recovery from the reaction mixture is carried out using the first member of the binding pair, wherein the primer used in step (iii) of [1] or, depending on the case, in step (iv) or in step (ii) of [5] or, depending on the case, in step (iii) is modified with the second member of the binding pair.
[0581]
[14] . According to the method of [1] to
[13] , the population of double-stranded DNA molecules is treated with aptamers before step (i) under conditions suitable for the ligation of the aptamer molecules to the DNA molecules, thereby introducing sticky ends into the DNA molecules.
[0582]
[15] . A method for determining the sequence of a population of double-stranded DNA molecules, the method comprising generating a library from the population of double-stranded DNA molecules using the method according to [1] to
[14] , and sequencing the DNA molecules obtained in step (iii) or, depending on the case, in step (iv) if the library has already been obtained by the method according to [1], or in step (ii) or, depending on the case, in step (iii) if the library has already been obtained by the method according to [5].
[0583]
[16] . A method for identifying methylated cytosine in a population of double-stranded DNA molecules, the method comprising the following steps:
[0584] (i) A library is generated from the population of double-stranded DNA molecules using the method according to [1] through
[14] , wherein, if the library has already been obtained by the method according to [1], the population of aptamer-modified DNA molecules is treated with a reagent before step (iii) or, if the library has already been obtained by the method according to [5], before step (ii), the reagent allowing the conversion of unmethylated cytosine to bases that are detectably different from cytosine in terms of hybridization properties, and wherein, if the library has already been obtained by the method according to [1], the primers used in steps (iii) and (iv) or, if the library has already been obtained by the method according to [5], are specific to the first aptamer molecules treated with the reagent, and
[0585] (ii) Sequencing the DNA molecules obtained in step (ii) or step (iii) or, as appropriate, step (iv) if the library has been obtained according to the method of [1], or in step (i) or step (ii) or, as appropriate, step (iii) if the library has been obtained according to the method of [5], wherein the DNA molecules have been treated with a reagent prior to step (iii) if the library has been obtained according to the method of [1], or prior to step (ii) if the library has been obtained according to the method of [5], the reagent allowing unmethylated cytosine to be converted to bases that are detectably different from cytosine in terms of hybridization properties.
[0586] If cytosine is present in one of the chains and guanine is present at the corresponding position in the opposite chain, the presence of methylated cytosine at the given position is determined; or if uracil or thymine is present in one of the chains and guanine is present at the corresponding position in the opposite chain, the presence of unmethylated cytosine at the given position is determined.
[0587]
[17] . A method for generating a double-stranded DNA library from a population of double-stranded DNA molecules, comprising the following steps:
[0588] (i) Contacting a population of double-stranded DNA molecules with a DNA Y-aptamer, the aptamer comprising a first DNA strand and a second DNA strand.
[0589] The 3' region of the first DNA strand and the 5' region of the second DNA strand form a double-stranded region through sequence complementarity.
[0590] The ends of the double-stranded region, formed by the 3' region of the first DNA strand of the Y-aptamer and the 5' region of the second DNA strand of the Y-aptamer, are compatible with the ends of the double-stranded DNA molecule.
[0591] The second DNA strand of the Y-aptamer forms a hairpin loop in its 3' region through hybridization between a first segment and a second segment within the 3' region. The first segment is located at the 3' end of the 3' region of the second DNA strand, and the second segment is located near the 5' region of the second DNA strand.
[0592] The contact is performed under conditions suitable for the ligation of the Y-aptamer to both ends of the double-stranded DNA molecule, thereby obtaining a plurality of DNA molecules containing the Y-aptamer.
[0593] (ii) Using each strand of the DNA molecule obtained in step (i) as a template, each strand of the DNA molecule obtained in step (i) is converted into a double-stranded DNA molecule by polymerase extension from the 3' end of the second DNA strand in the Y-aptamer molecule, and
[0594] (iii) Optionally, at least the double-stranded DNA molecule obtained in step (ii) is amplified using primers whose sequences are complementary to at least a portion of the double-stranded DNA molecule obtained in step (ii).
[0595]
[18] . The method according to
[17] , wherein the DNA molecule containing the Y-aptamer is obtained in step (i) under conditions suitable for separating the strands of the DNA molecule containing the Y-aptamer.
[0596]
[19] . A method for generating a double-stranded DNA library from a population of double-stranded DNA molecules, comprising the following steps:
[0597] (i) Contacting a population of double-stranded DNA molecules with a DNA Y-aptamer, the aptamer comprising a first DNA strand and a second DNA strand.
[0598] The 3' region of the first DNA strand and the 5' region of the second DNA strand form a double-stranded region through sequence complementarity, and the ends of the double-stranded region are compatible with the ends of the double-stranded DNA molecule.
[0599] The contact is performed under conditions suitable for the ligation of the Y-aptamer to both ends of the double-stranded DNA molecule, thereby obtaining a plurality of DNA molecules containing the Y-aptamer.
[0600] (ii) Each strand of the DNA molecule containing the Y-aptamer is contacted with an extension primer under conditions suitable for hybridization of the extension primer with the second strand of the Y-aptamer, the extension primer containing a 3' region complementary to the second DNA strand of the Y-aptamer molecule, and generating overhanging ends after hybridization with the second DNA strand of the Y-aptamer molecule.
[0601] (iii) The molecule produced in step (ii) is brought into contact with the hairpin aptamer under conditions suitable for the connection between the hairpin aptamer and the molecule produced in step (ii), the hairpin aptamer comprising a hairpin loop region and a cantilever end, the cantilever end being compatible with the cantilever end in the molecule produced in step (ii).
[0602] (iv) Each strand of the DNA molecule obtained in step (iii) is converted into a double-stranded DNA molecule by polymerase extension from the extension primers used in step (ii), and
[0603] (v) Optionally, the double-stranded DNA molecule obtained in step (iv) is amplified using at least the primers described below, wherein the sequences of the primers are complementary to at least a portion of the double-stranded DNA molecule obtained in step (iv).
[0604] The steps (iii) of connecting to the hair clip adapter and (iv) of extending the hair clip adapter can be performed in any order or simultaneously.
[0605]
[20] . The method according to
[19] , wherein the DNA molecule containing the Y-aptamer obtained in step (i) or step (iii) is placed under conditions suitable for separating the strands of the DNA molecule containing the Y-aptamer.
[0606]
[21] . The method according to
[19] or
[20] , wherein step (ii) of contacting each strand of the DNA molecule containing the Y-aptamer with the extension primer and step (iii) of contacting the molecule produced in step (ii) with the hairpin aptamer are performed in a single step by providing the hairpin aptamer and the extension primer in the form of a complex.
[0607]
[22] . The method according to
[17] to
[21] , wherein the double-stranded DNA molecule used in step (i) is a genomic DNA fragment.
[0608]
[23] . The method according to
[17] to
[22] , wherein the double-stranded DNA molecule used in step (i) is end-repaired prior to step (i).
[0609]
[24] . According to the method of
[23] , the method further includes a step of adding dA tails to the DNA molecule after the end repair step.
[0610]
[25] . The method according to
[17] to
[24] , wherein the Y-aptamer is provided in the form of an aptamer library, wherein each member of the library is distinguishable from other members by a combination sequence located in a double-stranded region formed by the 3' region of the first DNA strand and the 5' region of the second DNA strand of the aptamer.
[0611]
[26] . According to the method of
[25] , the combined sequence contains modified cytosine resistant to treatment with a reagent that allows unmethylated cytosine to be converted into bases that are detectably different from cytosine in terms of hybridization properties.
[0612]
[27] . According to the method of
[17] to
[26] , wherein the DNA molecule containing aptamers is treated with a reagent before step (iii) if the library has been obtained by the method of
[17] or before step (v) if the library has been obtained by the method of
[19] , the reagent allowing the conversion of unmethylated cytosine to bases that are detectably different from cytosine in terms of hybridization properties, and wherein the primer used in step (iii) of
[17] or, depending on the case, step (v) of
[19] is complementary to at least a portion of the sequence derived from the double-stranded DNA molecule obtained in step (ii) of
[17] or, depending on the case, in step (iv) of
[19] .
[0613]
[28] . The method according to
[17] to
[27] , wherein the hairpin aptamer and / or the Y-aptamer are provided fixed in a support, wherein the fixation is performed by binding the nucleotides of the hairpin loop of the hairpin aptamer and / or the nucleotides of the hairpin loop of the second DNA strand of the Y-aptamer and / or the 5' end of the first DNA strand of the Y-aptamer to the support.
[0614]
[29] . A method for determining the sequence of a population of double-stranded DNA molecules, the method comprising generating a library from the population of double-stranded DNA molecules using the method according to
[17] through
[28] , and sequencing the DNA molecules obtained in step (ii) or, depending on the case, in step (iii) if the library has been obtained by the method according to
[17] , or, depending on the case, in step (iv) if the library has been obtained by the method according to
[19] , or, depending on the case, in step (v).
[0615]
[30] . A method for identifying methylated cytosine in a population of double-stranded DNA molecules, the method comprising the following steps:
[0616] (i) A library is generated from the population of double-stranded DNA molecules using the method according to
[17] through
[28] , wherein, if the library has been obtained by the method according to
[17] prior to step (iii) or if the library has been obtained by the method according to
[19] prior to step (v), the population of aptamer-modified DNA molecules is treated with a reagent that allows the conversion of unmethylated cytosine to bases that are detectably different from cytosine in terms of hybridization properties, and wherein, if the library has been obtained by the method according to
[17] in step (iii) or if the library has been obtained by the method according to
[19] in step (v), the primers used are specific to the sequences of the double-stranded DNA molecules obtained in step (ii) of
[17] or step (iv) of
[19] after treatment with the reagent, and
[0617] (ii) Sequencing the DNA molecules obtained in step (ii) or, depending on the case, in step (iii) if the library has been obtained according to the method of
[17] , or in step (iv) or, depending on the case, in step (v) if the library has been obtained according to the method of
[19] , wherein the DNA molecules have been treated with a reagent prior to step (iii) if the library has been obtained according to the method of
[17] , or prior to step (v) if the library has been obtained according to the method of
[19] , the reagent allowing unmethylated cytosine to be converted into bases that are detectably different from cytosine in terms of hybridization properties.
[0618] If cytosine is present in one of the chains and guanine is present at the corresponding position in the opposite chain, the presence of methylated cytosine at the given position is determined; or if uracil or thymine is present in one of the chains and guanine is present at the corresponding position in the opposite chain, the presence of unmethylated cytosine at the given position is determined.
[0619]
[31] . A method for generating a double-stranded DNA library from a population of double-stranded DNA molecules, the method comprising the following steps:
[0620] (i) Fragmenting the double-stranded DNA molecules into a population under conditions suitable for generating a population of fragments of double-stranded DNA molecules with overhanging ends, wherein each end of each fragment binds to a hemiaptamer molecule comprising a first DNA strand and optionally a second DNA strand, wherein the second strand forms a double-stranded region with the first strand via complementarity to the central region of the first strand, and wherein the hemiaptamer molecule binds to a fragment of the double-stranded DNA molecule between the 3' end of the first strand of the hemiaptamer and the overhanging end of the fragment of the double-stranded DNA molecule.
[0621] (ii) Adding an alternative second strand or replacing the second DNA strand of the hemi-aptamer molecule with an alternative second strand, wherein the 5' region of the alternative second strand is complementary to the 3' region of the first strand of the hemi-aptamer molecule, and wherein the alternative second strand contains a region not complementary to the first strand of the hemi-aptamer molecule, thereby producing a plurality of DNA molecules containing Y-aptamers.
[0622] (iii) Optionally, fill the gap between the 5' end of the alternative second strand of the Y-aptamer and the 3' end of each DNA fragment.
[0623] (iv) Each strand of the DNA molecule containing the Y-aptamer is contacted with an extension primer under conditions suitable for hybridization of the extension primer with the alternative second strand of the Y-aptamer, the extension primer comprising a 3' region complementary to the alternative second DNA strand of the Y-aptamer molecule and a 5' region not hybridizing with the alternative second DNA strand of the Y-aptamer molecule.
[0624] (v) The molecule produced in step (iv) is brought into contact with the hairpin aptamer under conditions suitable for connecting the hairpin aptamer to the molecule produced in step (iv), the hairpin aptamer comprising a hairpin loop region and an end, the end being compatible with the end of the molecule produced in step (iv).
[0625] (vi) Each strand of the DNA molecule obtained in step (v) is converted into a double-stranded DNA molecule by polymerase extension using the extension primers used in step (iv), and
[0626] (vii) Optionally, the double-stranded DNA molecule obtained in step (vi) is amplified using at least the primers described below, wherein the sequence of the primers is complementary to at least a portion of the double-stranded DNA molecule obtained in step (vi).
[0627] The steps (v) of connecting to the hair clip adapter and the extension step (vi) can be performed in any order or simultaneously.
[0628]
[32] . The method according to
[31] , wherein the fragmentation step (i) is performed by means of contacting a population of double-stranded DNA molecules with a transposase dimer loaded with double-stranded aptamer molecules, wherein the aptamer molecules comprise a double-stranded region comprising a Tn5 inverted repeat and a 5' overhang of one of the strands, wherein optionally, cytosine nucleotides in the double-stranded region that do not form part of the Tn5 inverted repeat and cytosine nucleotides in the single-stranded region are methylated, and wherein the contact is performed under conditions suitable for DNA fragmentation and suitable for the attachment of hemiaptamer molecules to both ends of each DNA fragment.
[0629]
[33] . According to the method of
[31] or
[32] , the DNA molecule containing a Y-aptamer obtained in step (ii) or, depending on the case, in step (iii), or the DNA molecule containing a hairpin aptamer obtained in step (v) or, depending on the case, in step (vi), is placed under conditions suitable for separating the strands of the DNA molecule containing the Y-aptamer.
[0630]
[34] . The method according to
[31] to
[33] , wherein step (iv) of contacting each strand of the DNA molecule containing the Y-aptamer with the extension primer and step (v) of contacting the molecule produced in step (iv) with the hairpin aptamer are carried out in a single step by providing the hairpin aptamer and the extension primer in the form of a complex.
[0631]
[35] . The method according to
[31] to
[33] , wherein the steps of adding a second DNA strand or replacing a semi-aptamer molecule with a second DNA strand in a single step are performed by providing an alternative second strand, a hairpin aptamer and an extension primer in the form of a complex, (ii), contacting each strand of the DNA molecule containing the Y-aptamer with the extension primer, (iv), and contacting the molecule produced in step (iv) with the hairpin aptamer, (v).
[0632]
[36] . The method according to
[35] , wherein steps (ii) to (vi) are performed simultaneously.
[0633]
[37] . The method according to
[31] to
[36] , wherein the semi-aptamers used in step (i) are provided in the form of a semi-aptamer library, each member of which is distinguishable from other members by a combination sequence in the 3' region of the first chain of the semi-aptamer.
[0634]
[38] . According to the method of
[37] , the second chain of the semi-aptamer or the alternative second chain used in step (ii) does not show any substantial overlap with the combination region.
[0635]
[39] . The method according to
[37] or
[38] , wherein the combined sequence contains one or more modified cytosines resistant to treatment with a reagent that allows unmethylated cytosines to be converted into bases that are detectably different from cytosines in terms of hybridization properties.
[0636]
[40] . The method according to
[31] to
[39] , wherein the DNA molecule containing the aptamer is treated with a reagent that allows unmethylated cytosine to be converted to a base that is detectably different from cytosine in terms of hybridization properties before step (vii), and wherein the primer used in step (vii) is complementary to at least a portion of the sequence derived from the double-stranded DNA molecule obtained in step (vi) by treating it with the reagent.
[0637]
[41] . According to the method described in
[31] to
[40] , the molecules obtained in step (vi) or, depending on the specific circumstances, in step (vii) are recovered from the reaction mixture.
[0638]
[42] . A method for determining the sequence of a population of double-stranded DNA molecules, the method comprising generating a library from the population of double-stranded DNA molecules using the method according to
[31] to
[41] , and sequencing the DNA molecules obtained in step (vi) or, depending on the specific circumstances, in step (vii).
[0639]
[43] . A method for identifying methylated cytosine in a population of double-stranded DNA molecules, the method comprising the following steps:
[0640] (i) A library is generated from the population of double-stranded DNA molecules using the method according to
[31] to
[41] , wherein prior to step (vii), the population of aptamer-modified DNA molecules is treated with a reagent that allows unmethylated cytosine to be converted to bases that are detectably different from cytosine in terms of hybridization characteristics, and wherein the primers used in step (vii) are specific to the sequences of the double-stranded DNA molecules obtained in step (vi) after treatment with said reagent, and
[0641] (ii) Sequencing the DNA molecules obtained in step (vi) or, depending on the specific circumstances, in step (vii), said DNA molecules having been treated prior to step (vii) with a reagent that allows unmethylated cytosine to be converted to bases that are detectably different from cytosine in terms of hybridization properties.
[0642] If cytosine is present in one of the chains and guanine is present at the corresponding position in the opposite chain, the presence of methylated cytosine at the given position is determined; or if uracil or thymine is present in one of the chains and guanine is present at the corresponding position in the opposite chain, the presence of unmethylated cytosine at the given position is determined.
[0643]
[44] . A method for generating a double-stranded DNA library from a population of double-stranded DNA molecules, the method comprising the following steps:
[0644] (i) Contacting a population of double-stranded DNA molecules with a population of DNA aptamers, each aptamer containing a first DNA strand and a second DNA strand.
[0645] The 3' region of the first DNA strand and the 5' region of the second DNA strand form a double-stranded region through sequence complementarity, and the ends of the double-stranded region are compatible with the ends of the double-stranded DNA molecule.
[0646] Each aptamer in the population is distinguishable from other aptamers by a combinatorial sequence located within a double-stranded region formed between the 3' region of the first DNA strand and the 5' region of the second DNA strand.
[0647] The contact is performed under conditions suitable for the ligation of the aptamer to both ends of a double-stranded DNA molecule, thereby obtaining multiple DNA molecules containing the aptamer.
[0648] (ii) Optionally, the aptamer-containing DNA molecule obtained in step (i) is treated with a reagent that allows unmethylated cytosine to be converted into bases that are detectably different from cytosine in terms of hybridization properties, and
[0649] (iii) Optionally, at least the primers described below are used to amplify the DNA molecule containing aptamers obtained in step (i) or, depending on the specific circumstances, step (ii), wherein the sequence of the primers is complementary to at least a portion of the DNA molecule containing aptamers obtained in step (i) or step (ii).
[0650]
[45] . According to the method of
[44] , the 5' region of the first DNA strand and the 3' region of the second DNA strand are complementary.
[0651]
[46] . According to the method of
[44] , wherein the aptamer is a Y-aptamer, wherein the 5' region of the first DNA strand and the 3' region of the second DNA strand are not complementary.
[0652]
[47] . A method for generating a double-stranded DNA library from a population of double-stranded DNA molecules, comprising the following steps:
[0653] (i) A population of double-stranded DNA molecules is fragmented under conditions suitable for generating a population of fragments of double-stranded DNA molecules with overhanging ends, wherein each end of each fragment is bound to a hemiaptamer molecule comprising a first DNA strand and optionally a second DNA strand, wherein each hemiaptamer is distinguishable from other hemiaptamers by a combination sequence located within a 3' region of the first DNA strand, wherein the second strand forms a double-stranded region with the first strand via complementarity with the central region of the first strand, and wherein the hemiaptamer molecule binds to a double-stranded DNA molecule fragment located between the 3' end of the first strand of the hemiaptamer and the overhanging end of the double-stranded DNA molecule fragment.
[0654] (ii) Adding an alternative second strand or replacing the second DNA strand of the hemi-aptamer molecule with an alternative second strand, wherein the 5' region of the alternative second strand is complementary to the 3' region of the first strand of the hemi-aptamer molecule, and wherein the alternative second strand contains a region not complementary to the first strand of the hemi-aptamer molecule, thereby producing a plurality of DNA molecules containing Y-aptamers.
[0655] (iii) Optionally, fill the gap between the 5' end of the alternative second strand of the Y-aptamer and the 3' end of each DNA fragment.
[0656] (iv) Optionally, the DNA molecule containing the Y-aptamer obtained in step (iii) is treated with a reagent that allows unmethylated cytosine to be converted into bases that are detectably different from cytosine in terms of hybridization properties, and
[0657] (v) Optionally, at least the primers described below are used to amplify the DNA molecule containing the Y-aptamer obtained in step (ii), step (iii), or, depending on the specific circumstances, step (iv), wherein the sequence of the primers is complementary to at least a portion of the DNA molecule containing the Y-aptamer obtained in step (ii), step (iii), or, depending on the specific circumstances, step (iv).
[0658]
[48] . The method according to
[47] is carried out by means of (i), the method comprising contacting a population of double-stranded DNA molecules with a transposase dimer loaded with double-stranded aptamer molecules, wherein the aptamer molecules comprise a double-stranded region comprising a Tn5 inverted repeat and a 5' overhang of one of the strands, wherein optionally, cytosine nucleotides in the double-stranded region that do not form part of the Tn5 inverted repeat and cytosine nucleotides in the single-stranded region are methylated, and wherein the contact is carried out under conditions suitable for DNA fragmentation and suitable for the attachment of hemiaptamer molecules to both ends of each DNA fragment.
[0659]
[49] . According to the method of
[47] or
[48] , the second chain of the semi-aptamer or the alternative second chain used in step (ii) does not show any substantial overlap with the combined region.
[0660]
[50] . The method according to
[44] to
[49] , wherein the double-stranded DNA molecule used in step (i) is a genomic DNA fragment.
[0661]
[51] . The method according to
[44] to
[50] , wherein the double-stranded DNA molecule used in step (i) is end-repaired prior to step (i).
[0662]
[52] . According to the method of
[51] , the method further includes a step of adding dA tails to the DNA molecule after the end repair step.
[0663]
[53] . According to the method of
[44] to
[52] , wherein the population of double-stranded DNA molecules is pretreated with aptamers in step (i) under conditions suitable for the ligation of aptamers to the DNA molecules, thereby introducing sticky ends into the DNA molecules.
[0664]
[54] . According to the method of
[44] to
[53] , the combined sequence within the aptamer sequence contains one or more modified cytosines that are resistant to treatment with a reagent that allows the conversion of unmethylated cytosines into bases that are detectably different from cytosines in terms of hybridization properties.
[0665]
[55] . The method according to
[44] to
[54] , wherein the DNA molecules obtained in step (i) or step (ii) of
[44] or, depending on the case, step (iii) are recovered from the reaction mixture obtained in step (iii), step (iv) of
[47] or, depending on the case, step (v). The DNA molecules obtained in step (i) or step (ii) in cases where the library has been obtained by the method according to
[44] or in step (iii) or step (iv) in cases where the library has been obtained by the method according to
[47] .
[0666]
[56] . The method according to
[55] , wherein the recovery from the reaction mixture is carried out using a first member of the binding pair, wherein the aptamer is modified with a second member of the binding pair.
[0667]
[57] . The method according to
[44] to
[56] , wherein the DNA molecules obtained in step (iii) in cases where the library has been obtained by the method according to
[44] or the DNA molecules obtained in step (v) in cases where the library has been obtained by the method according to
[47] are recovered from the reaction mixture.
[0668]
[58] . The method according to
[57] , wherein the recovery from the reaction mixture is carried out using the first member of the binding pair, wherein the primer used in step (iii) of
[44] or step (v) of
[47] is modified with the second member of the binding pair.
[0669]
[59] . A method for determining the sequence of a population of double-stranded DNA molecules, the method comprising generating a library from the population of double-stranded DNA molecules using the method according to
[44] through
[58] , and sequencing the DNA molecules obtained in step (i) or step (ii) or, depending on the case, in step (iii) if the library has been obtained by the method according to
[44] , or in step (iv) if the library has been obtained by the method according to
[47] , or, depending on the case, in step (v).
[0670]
[60] . A method for identifying methylated cytosine in a population of double-stranded DNA molecules, the method comprising:
[0671] (i) A library is generated from the population of double-stranded DNA molecules using the method according to
[44] through
[58] , wherein, if the library has already been obtained by the method according to
[44] , the population of aptamer-modified DNA molecules is treated with a reagent before step (iii) or, if the library has already been obtained by the method according to
[47] , before step (v), the reagent allowing the conversion of unmethylated cytosine to bases that are detectably different from cytosine in terms of hybridization characteristics, and wherein, if the library has already been obtained by the method according to
[44] , the primers used in step (iii) or, if the library has already been obtained by the method according to
[47] , in step (v) are specific to the sequence of the DNA molecules obtained in step (ii) of
[44] or step (iv) of
[47] , and
[0672] (ii) sequencing of DNA molecules obtained in step (ii) or, depending on the case, in step (iii) if the library has already been obtained by the method according to
[44] , or sequencing of DNA molecules obtained in step (iv) or, depending on the case, in step (v) if the library has already been obtained by the method according to
[47] .
[0673] If cytosine is present in one of the chains and guanine is present at the corresponding position in the opposite chain, the presence of methylated cytosine at the given position is determined; or if uracil or thymine is present in one of the chains and guanine is present at the corresponding position in the opposite chain, the presence of unmethylated cytosine at the given position is determined.
[0674]
[61] . A DNA library that can be obtained by the method according to [1] to
[14] ,
[17] to
[28] ,
[31] to
[41] or
[44] to
[58] .
[0675]
[62] . A Y-aptamer library
[0676] Each Y-aptamer contains a first DNA strand and a second DNA strand, wherein the 3' region of the first DNA strand and the 5' region of the second DNA strand form a double-stranded region through sequence complementarity.
[0677] The 3' region of the second DNA strand forms a hairpin loop through hybridization between a first segment and a second segment of the 3' region. The first segment is located at the 3' end of the 3' region, and the second segment is located near the region of the second DNA strand that forms a double-stranded region with the 3' region of the first DNA strand.
[0678] Each member of the library can be distinguished from other members by a combination of sequences located in the double-stranded region formed between the 3' region of the first DNA strand and the 5' region of the second DNA strand.
[0679]
[63] . According to the library described in
[62] , the combined sequence contains one or more modified cytosines that are resistant to treatment with a reagent that allows the conversion of unmethylated cytosines into bases that are detectably different from cytosines in terms of hybridization properties.
[0680]
[64] . According to the library described in
[62] or
[63] , the 3' end of the second polynucleotide in each aptamer is reversibly closed.
[0681]
[65] . According to the libraries of
[62] to
[64] , the terminal segment of the double-stranded region formed between the 3' region of the first DNA strand and the 5' region of the second DNA strand in each aptamer contains the target site of a restriction endonuclease.
[0682]
[66] . A method for generating DNA Y-aptamers, the method comprising the following steps:
[0683] (i) Contacting a first single-stranded polynucleotide with a second single-stranded polynucleotide, wherein at least a portion of the 3' region of the first single-stranded polynucleotide is complementary to the 5' region of the second polynucleotide.
[0684] The 3' region of the second single-stranded polynucleotide forms a hairpin loop through the hybridization of a first segment and a second segment of the 3' region. The first segment is located at the 3' end of the 3' region of the second polynucleotide, and the second segment is located near the 5' region of the second polynucleotide. The 3' end of the second polynucleotide is reversibly closed.
[0685] The contact is carried out under conditions suitable for hybridization of complementary regions within the 3' region of the first single-stranded polynucleotide and the 5' region of the second single-stranded polynucleotide, thereby producing a double-stranded DNA molecule.
[0686] (ii) Extending the 3' end of the first polynucleotide to generate a sequence complementary to the 5' region of the second polynucleotide within the first single-stranded polynucleotide, and
[0687] (iii) Optionally, unblock the 3' end of the second polynucleotide.
[0688]
[67] . The method according to
[66] , wherein the extension carried out in step (ii) is carried out in the presence of a modified cytosine that is resistant to treatment with a reagent that allows the unmethylated cytosine to be converted into a base that is detectably different from cytosine in terms of hybridization properties.
[0689]
[68] . According to the method of
[66] or
[67] , at least one position within the 3' end of the first polynucleotide is guanine, and the position within the 5' region of the second polynucleotide that hybridizes with the one or more positions within the 3' end of the first polynucleotide is cytosine or methylcytosine.
[0690]
[69] . According to the method of
[66] to
[68] , the 5' end of the second single-stranded polynucleotide contains a sequence that is copied onto the complementary strand when the complementary strand is formed during the elongation step (ii) to form a target site for the restriction endonuclease.
[0691]
[70] . The method according to
[66] to
[69] , wherein the second single-stranded polynucleotide is provided in the form of a polynucleotide library, wherein each member of the library is distinguishable from other members by a combination sequence located in the 5' region of the polynucleotide, and wherein the combination sequence is located upstream relative to a region that shows sequence complementarity with the first single-stranded polynucleotide, thereby producing a DNA Y-aptamer molecular library.
[0692]
[71] . According to the method of
[70] , the combined sequence contains one or more modified cytosines that are resistant to treatment with a reagent that allows unmethylated cytosines to be converted into bases that are detectably different from cytosines in terms of hybridization properties.
[0693]
[72] . The method according to
[66] to
[71] , wherein the first single-stranded polynucleotide and / or the second single-stranded polynucleotide are provided immobilized in a support, wherein the immobilization is performed by binding the nucleotide at the 5' end of the first single-stranded polynucleotide or the hairpin loop of the second single-stranded polynucleotide to the support.
[0694]
[73] . A kit comprising:
[0695] (i) A first single-stranded polynucleotide containing the 5' and 3' regions
[0696] (ii) A second polynucleotide comprising a 5' region and a 3' region, wherein the 3' region forms a hairpin loop through hybridization between a first segment and a second segment within the 3' region, the first segment being located at the 3' end of the 3' region, and the second segment being located near the 5' region, and wherein the 3' end of the second polynucleotide is reversibly closed.
[0697] The 3' region of the first single-stranded polynucleotide is complementary to at least a portion of the 5' region of the second polynucleotide.
[0698]
[74] . The kit according to
[73] , wherein the second polynucleotide is provided in the form of a polynucleotide library, wherein each member is distinguishable from other members by a combinatorial sequence located in the 5' region of the second polynucleotide and upstream relative to a region that shows sequence complementarity with the first single-stranded polynucleotide.
[0699]
[75] . According to the kit described in
[74] , the combined sequence contains one or more modified cytosines that are resistant to treatment with a reagent that allows unmethylated cytosines to be converted into bases that are detectably different from cytosines in terms of hybridization properties.
[0700]
[76] . The kit according to
[73] to
[75] , wherein the 5' region of the second polynucleotide contains a sequence that generates a target site for restriction endonuclease when converted into a double-stranded region.
[0701]
[77] . The kit according to
[73] to
[76] further comprises one or more components selected from:
[0702] (i) DNA polymerase,
[0703] (ii) One or more nucleotides selected from A, G, C and T,
[0704] (iii) One or more modified cytosines resistant to treatment with a reagent that allows the conversion of unmethylated cytosine into bases that are detectably different from cytosine in terms of hybridization properties.
[0705] (iv) A reagent capable of removing the blocking group from the 3' end of the second polynucleotide.
[0706] (v) A reagent that allows the conversion of unmethylated cytosine into bases that are detectably different from cytosine in terms of hybridization properties, and
[0707] (vi) A target site-specific restriction endonuclease, wherein the target site is formed by a sequence within the 5' end of the second polynucleotide.
[0708]
[78] . A kit comprising:
[0709] (i) the Y-adapter sub-library as described in
[62] to
[65] ; and
[0710] (ii) one or more components, said components being selected from:
[0711] i. DNA polymerase,
[0712] ii. One or more nucleotides selected from A, G, C, and T,
[0713] iii. One or more modified cytosines resistant to treatment with a reagent, said reagent allowing the conversion of unmethylated cytosine into bases that are detectably different from cytosine in terms of hybridization properties.
[0714] iv. A reagent capable of removing the blocking group from the 3' end of the second polynucleotide.
[0715] v. A reagent that allows the conversion of unmethylated cytosine into bases that are detectably different from cytosine in terms of hybridization properties, and
[0716] vi. A target-site specific restriction endonuclease, wherein the target site is formed by a sequence within the 5' end of the second polynucleotide.
[0717]
[79] . According to the kit of
[77] or
[78] , the one or more modified cytosines are selected from: methylcytosine, hydroxymethylcytosine and combinations thereof.
[0718]
[80] . A method for obtaining a double-stranded DNA aptamer library, wherein each aptamer comprises a first DNA strand and a second DNA strand, and wherein each aptamer is distinguishable from other aptamers by a combination sequence located within a double-stranded region formed between the 3' region of the first DNA strand and the 5' region of the second DNA strand, the method comprising the following steps:
[0719] (i) Providing a population of single-stranded DNA molecules comprising a constant region and a combinatorial region, wherein the single-stranded DNA molecules are distinguishable from other molecules by a sequence in the combinatorial region, wherein the constant region is located at the 3' end relative to the combinatorial region, and wherein the 3' end is reversibly closed, and
[0720] (ii) By using the single-stranded DNA molecule of step (i) as a template and using an extension primer that fully or partially hybridizes with the constant region of the single-stranded DNA molecule to generate double-stranded DNA, the combined region is copied onto the newly generated strand, thereby generating a double-stranded combined sequence.
[0721]
[81] . According to the method of
[80] , the method further includes removing the blocking group from the 3' end of the single-stranded DNA molecule.
[0722]
[82] . The method according to
[80] or
[81] , wherein the elongation primer comprises a 5' region of a protrusion that does not hybridize with the constant region of the single-stranded DNA molecule.
[0723]
[83] . According to the method of
[80] to
[82] , the constant region of the single-stranded DNA molecule forms a hairpin loop through hybridization between a first segment and a second segment within the constant region.
[0724]
[84] . The method according to
[80] or
[81] further includes linking an aptamer of the library to a second DNA molecule having a double-stranded region whose ends are compatible with the ends of the aptamer molecule.
[0725]
[85] . According to the method of
[84] , wherein the second DNA molecule is contained in a 5' region of the first strand and / or a 3' region of the second strand, the suspended regions not hybridizing with each other.
[0726]
[86] . According to the method of
[85] , the 3' overhang region in the second chain forms a hairpin loop through hybridization between the first segment and the second segment in the region.
[0727]
[87] . The method according to
[80] to
[82] or
[84] to
[86] , wherein each single-stranded DNA molecule provided in step (i) is immobilized in a support, wherein the immobilization is performed by binding the 5' end of the single-stranded DNA molecule to the support.
[0728]
[88] . The method according to
[83] , wherein each single-stranded DNA molecule provided in step (i) is immobilized in a support, wherein the immobilization is performed by binding nucleotides of the hairpin loop of the single-stranded DNA molecule to the support.
[0729]
[89] . A library of double-stranded DNA aptamer molecules, wherein each DNA aptamer molecule comprises a constant region and a variable region, wherein each double-stranded DNA aptamer comprises a first DNA strand and a second DNA strand, and wherein each aptamer is distinguishable from other aptamers by a combination of sequences in a variable region within a double-stranded region formed between a 3' region of the first DNA strand and a 5' region of the second DNA strand.
[0730]
[90] . According to the library described in
[89] , at least one of the chains has a reversibly closed 3' end.
[0731]
[91] . According to the library described in
[89] or
[90] , one or two chains contain a dangling region that does not hybridize with the opposite chain.
[0732]
[92] . According to the library described in
[91] , the constant region of one of the chains forms a hairpin loop through hybridization between a first segment and a second segment within the constant region.
Claims
1. A method for identifying methylated cytosine in a population of double-stranded DNA molecules, the method comprising the following steps: (i) To connect a double-stranded DNA integrator to at least one end of the strands of a plurality of double-stranded DNA molecules and to pair the strands of the plurality of double-stranded DNA molecules to provide a plurality of paired integrator-modified DNA molecules; (ii) Converting any unmethylated cytosine in the paired, adaptor-modified DNA molecule into uracil in the paired, adaptor-modified DNA molecule; (iii) Using nucleotides A, G, C and T and primers to provide complementary strands of paired and transformed adaptor-modified DNA molecules, wherein the sequences of the primers are complementary to at least a portion of the double-stranded DNA adaptor to provide partially transformed paired double-stranded molecules; (iv) Optionally, the partially converted paired double-stranded DNA molecules obtained in step (iii) are amplified to provide amplified paired double-stranded DNA molecules; (v) Sequencing the paired DNA molecules obtained in step (iii) or step (iv), If cytosine is present in one strand of the paired double-stranded DNA molecule obtained in step (iii) or step (iv) and guanine is present at the corresponding position in the other strand of the paired double-stranded DNA molecule, then the presence of methylated cytosine at the given position is determined; and / or if uracil or thymine is present in one strand of the paired double-stranded DNA molecule obtained in step (iii) or step (iv) and guanine is present at the corresponding position in the other strand of the paired double-stranded DNA molecule, then the presence of unmethylated cytosine at the given position is determined. The plurality of paired, adaptor-modified DNA molecules obtained in step (i) are obtained as follows: (a) To attach a DNA Y-aparasome to each end of a plurality of double-stranded DNA molecules, the aparasome comprising a first DNA strand and a second DNA strand, The 3' region of the first DNA strand and the 5' region of the second DNA strand form a double-stranded region through sequence complementarity. The ends of the double-stranded region formed by the 3' region of the first DNA strand and the 5' region of the second DNA strand of the DNA Y-aptamer are compatible with the ends of the double-stranded DNA molecule. (b) For each strand of the DNA molecule obtained in step (a), a complementary strand is synthesized by polymerase extension from the 3' end of the second DNA strand in the DNA Y-aptamer molecule using each strand of the DNA molecule obtained in step (a) as a template, thereby pairing each strand of the DNA molecule obtained in step (a) with its synthesized complementary strand to provide a plurality of paired adaptor-modified DNA molecules. In step (a), the plurality of double-stranded DNA molecules are genomic DNA fragments. The strands of the genomic DNA fragments are paired to provide multiple paired genomic DNA fragments. The pairing of the genomic DNA fragment strands is performed using barcode sequences. Step (b) uses nucleotides A, G, C, and T.
2. The method according to claim 1, wherein in step (i), the pairing can be performed before or after the connection, or simultaneously with the connection.
3. The method according to one or more of the preceding claims, wherein at least a portion of the double-stranded DNA adaptor has a sequence common to all the double-stranded DNA adaptors used in step (i).
4. The method according to one or more of the preceding claims, wherein prior to step (ii), the plurality of paired adaptor-modified DNA molecules are separated to generate a library of paired adaptor-modified DNA molecules.
5. The method according to one or more of the preceding claims, wherein the conversion of unmethylated cytosine to uracil in the paired DNA molecules is performed using bisulfite.
6. The method according to one or more of claims 1-5, wherein the 3' region of the second DNA strand of the DNA Y-aptamer forms a hairpin loop through hybridization between a first segment and a second segment within the 3' region, the first segment being located at the 3' end of the 3' region of the second DNA strand, and the second segment being located near the 5' region of the second DNA strand.
7. The method according to one or more of claims 1-6, wherein the DNA Y aptamer has a first barcode sequence in the double-stranded region of the DNA Y aptamer and / or a second barcode sequence in the 3' region of the second DNA strand.
8. The method according to one or more of claims 1 to 7, wherein the DNA Y aptamer has a restriction site in the 5' region of the first DNA strand of the DNA Y aptamer.
9. A method for identifying methylated cytosine in a population of double-stranded DNA molecules, the method comprising the following steps: (i) To connect a double-stranded DNA integrator to at least one end of the strands of a plurality of double-stranded DNA molecules and to pair the strands of the plurality of double-stranded DNA molecules to provide a plurality of paired integrator-modified DNA molecules; (ii) Converting any unmethylated cytosine in the paired, adaptor-modified DNA molecule into uracil in the paired, adaptor-modified DNA molecule; (iii) Using nucleotides A, G, C and T and primers to provide complementary strands of paired and transformed adaptor-modified DNA molecules, wherein the sequences of the primers are complementary to at least a portion of the double-stranded DNA adaptor to provide partially transformed paired double-stranded molecules; (iv) Optionally, the partially converted paired double-stranded DNA molecules obtained in step (iii) are amplified to provide amplified paired double-stranded DNA molecules; (v) Sequencing the paired DNA molecules obtained in step (iii) or step (iv), If cytosine is present in one strand of the paired double-stranded DNA molecule obtained in step (iii) or step (iv) and guanine is present at the corresponding position in the other strand of the paired double-stranded DNA molecule, then the presence of methylated cytosine at the given position is determined; and / or if uracil or thymine is present in one strand of the paired double-stranded DNA molecule obtained in step (iii) or step (iv) and guanine is present at the corresponding position in the other strand of the paired double-stranded DNA molecule, then the presence of unmethylated cytosine at the given position is determined. The plurality of paired, adaptor-modified DNA molecules obtained in step (i) are obtained as follows: (a) Contacting the population of double-stranded DNA molecules with a DNA Y-aptamer, the aptamer comprising a first DNA strand and a second DNA strand, The 3' region of the first DNA strand and the 5' region of the second DNA strand form a double-stranded region through sequence complementarity, and the ends of the double-stranded region are compatible with the ends of the double-stranded DNA molecule. The contact is performed under conditions suitable for the ligation of the DNA Y-aptamer to both ends of the double-stranded DNA molecule, thereby obtaining a plurality of DNA molecules containing the DNA Y-aptamer. (b) Each strand of the DNA molecule containing the DNA Y-aptamer is contacted with an extension primer under conditions suitable for hybridization of the extension primer with the second DNA strand of the DNA Y-aptamer, the extension primer containing a 3' region complementary to the second DNA strand of the DNA Y-aptamer molecule, and generating overhanging ends after hybridization with the second DNA strand of the DNA Y-aptamer molecule. (c) The molecule produced in step (b) is brought into contact with a hairpin aptamer under conditions suitable for connecting the hairpin aptamer to the molecule produced in step (b), the hairpin aptamer comprising a hairpin loop region and a cantilever end, the cantilever end being compatible with the cantilever end in the molecule produced in step (b). (d) Each strand of the DNA molecule obtained in step (b) or (c) is converted into a double-stranded DNA molecule by polymerase extension using the extension primers used in step (b). The steps (c) and (d) of connecting the hair clip adapter can be performed in any order or simultaneously. In step (i), the plurality of double-stranded DNA molecules are genomic DNA fragments to provide a plurality of paired, adaptor-modified genomic DNA fragments. The pairing described in step (i) is performed using a barcode sequence.
10. A method for identifying methylated cytosine in a population of double-stranded DNA molecules, the method comprising the following steps: (i) To connect a double-stranded DNA integrator to at least one end of the strands of a plurality of double-stranded DNA molecules and to pair the strands of the plurality of double-stranded DNA molecules to provide a plurality of paired integrator-modified DNA molecules; (ii) Converting any unmethylated cytosine in the paired, adaptor-modified DNA molecule into uracil in the paired, adaptor-modified DNA molecule; (iii) Using nucleotides A, G, C and T and primers to provide complementary strands of paired and transformed adaptor-modified DNA molecules, wherein the sequences of the primers are complementary to at least a portion of the double-stranded DNA adaptor to provide partially transformed paired double-stranded molecules; (iv) Optionally, the partially converted paired double-stranded DNA molecules obtained in step (iii) are amplified to provide amplified paired double-stranded DNA molecules; (v) Sequencing the paired DNA molecules obtained in step (iii) or step (iv), If cytosine is present in one strand of the paired double-stranded DNA molecule obtained in step (iii) or step (iv) and guanine is present at the corresponding position in the other strand of the paired double-stranded DNA molecule, then the presence of methylated cytosine at the given position is determined; and / or if uracil or thymine is present in one strand of the paired double-stranded DNA molecule obtained in step (iii) or step (iv) and guanine is present at the corresponding position in the other strand of the paired double-stranded DNA molecule, then the presence of unmethylated cytosine at the given position is determined. The plurality of paired, adaptor-modified DNA molecules obtained in step (i) are obtained as follows: (a) Contacting the population of double-stranded DNA molecules with a DNA Y-aptamer, the aptamer comprising a first DNA strand and a second DNA strand, The 3' region of the first DNA strand and the 5' region of the second DNA strand form a double-stranded region through sequence complementarity, and the ends of the double-stranded region are compatible with the ends of the double-stranded DNA molecule. The contact is performed under conditions suitable for the ligation of the DNA Y-aptamer to both ends of the double-stranded DNA molecule, thereby obtaining a plurality of DNA molecules containing the DNA Y-aptamer. (b) Contacting each strand of the DNA molecule containing the DNA Y-aptamer with a complex of an extension primer and a hairpin aptamer under conditions suitable for hybridization of the extension primer with the second DNA strand of the DNA Y-aptamer, wherein the extension primer contains a 3' region complementary to the second DNA strand of the DNA Y-aptamer molecule and generates a dangling end after hybridization with the second DNA strand of the DNA Y-aptamer molecule, and wherein the hairpin aptamer comprises a hairpin loop region and a dangling end, the dangling end being compatible with the dangling end formed after hybridization of the extension primer and the second DNA strand of the DNA Y-aptamer. (c) Each strand of the DNA molecule obtained in step (b) is converted into a double-stranded DNA molecule by polymerase extension using the extension primers used in step (b). In step (i), the plurality of double-stranded DNA molecules are genomic DNA fragments to provide a plurality of paired, adaptor-modified genomic DNA fragments. The pairing described in step (i) is performed using a barcode sequence.
11. A DNA Y-aptamer for use in the method of any one of claims 1 to 10, wherein the DNA Y-aptamer comprises a first DNA strand and a second DNA strand. The 3' region of the first DNA strand and the 5' region of the second DNA strand form a double-stranded region through sequence complementarity. The ends of the double-stranded region formed by the 3' region of the first DNA strand and the 5' region of the second DNA strand of the DNA Y-aptamer are compatible with the ends of the double-stranded DNA molecule. The double-stranded region of the DNA Y-aptamer contains one or more barcode sequences, and The 3' region of the second DNA strand of the DNA Y-aptamer forms a hairpin loop through hybridization between a first segment and a second segment within the 3' region, wherein the first segment is located at the 3' end of the 3' region of the second DNA strand, and the second segment is located near the 5' region of the second DNA strand, and / or The DNA Y-aptamer contains at least one barcode sequence within the single-stranded region of the DNA Y-aptamer.
12. The DNA Y-aptamer of claim 11, wherein the DNA Y-aptamer has a restriction site in the 5' region of the first DNA strand of the DNA Y-aptamer.
13. The DNA Y-aptamer according to any one of claims 11 to 12, wherein the DNA Y-aptamer comprises at least one barcode sequence in a single-stranded region of the DNA Y-aptamer, and wherein the 3' region of the second DNA strand of the DNA Y-aptamer forms a hairpin loop through hybridization between a first segment and a second segment within the 3' region, the first segment being located at the 3' end of the 3' region of the second DNA strand, and the second segment being located near the 5' region of the second DNA strand.
14. A DNA Y-aptamer library comprising a DNA Y-aptamer as defined in any one of claims 11 to 13, wherein each member of the library is distinguishable from other members by a combinatorial sequence located within the double-stranded region formed by the 3' region of the first DNA strand of the aptamer and the 5' region of the second DNA strand.
15. A kit comprising a DNA Y-aptamer library as defined in claim 14.
16. A reagent kit comprising: (i) A DNA Y-aptamer library, wherein the aptamers comprise a first DNA strand and a second DNA strand, wherein, The 3' region of the first DNA strand and the 5' region of the second DNA strand form a double-stranded region through sequence complementarity, wherein the ends of the double-stranded region are compatible with the ends of the double-stranded DNA molecule. (ii) Multiple extension primers, wherein each extension primer contains a 3' region complementary to the second DNA strand of the DNA Y-aptamer molecule as defined in (i), and produces overhanging ends upon hybridization with the second DNA strand of the DNA Y-aptamer molecule; and (iii) Multiple hairpin aptamers, wherein each hairpin aptamer includes a hairpin loop region and a drooping end, the drooping end being compatible with the drooping end formed after hybridization of the second DNA strand of the extension primer defined in (ii) and the DNA Y-aptamer defined in (i); The extended primer in (ii) and the hairpin aptamer in (iii) can be provided in the form of a complex; The DNA Y-aptamer of (i), the elongation primer of (ii), and the hairpin aptamer of (iii) are adapted to obtain a DNA Y-aptamer library for use in the method of any one of claims 1 to 10.
Citation Information
Patent Citations
Improved method for bisulfite treatment
EP1394172A1
Methods and compositions for DNA fragmentation and tagging by transposases
EP2527438A1
Processes for transposase mediated integration into mammalian cells
US20030143740A1
System for in vitro transposition
US5965443A
Nucleic acid transfer vector for the introduction of nucleic acid into the DNA of a cell
US7160682B2