DNA sequencing method

By employing GEUS molecular technology and multi-primer sequencing methods, the problem of high error rates in nucleic acid sequencing has been solved, achieving low-error-rate and high-quality determination of nucleic acid molecular sequences, especially accurate identification of cytosine methylation, which is suitable for high-precision analysis of genomic DNA and RNA molecules.

CN121443752APending Publication Date: 2026-01-30ANILIN CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202480040077.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-04-13
Filing Date
2024-04-15
Publication Date
2026-01-30

AI Technical Summary

Technical Problem

Existing nucleic acid sequencing technologies have a high error rate, especially when detecting low-frequency gene variations, which are often confounded by many factors and make it difficult to accurately identify cytosine nucleic acid methylation, thus affecting molecular diagnosis and monitoring of cancer.

Method used

Using GEUS molecular technology, nucleic acid molecules in the 5' and 3' regions are covalently linked, and multiple primers are used to sequence them, ensuring the consistency of reads for each nucleotide and providing multiple sources of information to verify and correct errors in sequencing.

Benefits of technology

Significantly reduces sequencing error rate, improves sequencing quality and process efficiency, and enables faster and lower-cost determination of nucleic acid molecular sequences, especially accurate identification of cytosine methylation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The present invention relates to a method for determining the sequence of a nucleic acid molecule. Specifically, the present invention provides a method comprising: i. Providing a nucleic acid molecule comprising a 5 '-region and a 3'-region wherein the 5 '-region and the 3'-region are covalently linked by a nucleotide sequence that can bind to a primer wherein the 5 '-region and the 3'-region are covalently linked by a nucleotide sequence that can bind to the primer, the base identity in one of the 5'region or the 3 'region and the base identity in the other region independently provide information about the base identity in the corresponding locus in the original nucleic acid molecule wherein the molecule further comprises: a linker located at the 5'end of the molecule; a linker located at the 3'end of the molecule; ii. Sequencing the molecule provided in step (i) using at least two different primers wherein the at least two different primers bind to at least three, preferably at least four, different regions of the nucleic acid molecule provided in (i), wherein: 1. At least one of the primers is capable of binding at least partially to at least a portion of the linker at the 5'end of said molecule for sequencing at least a portion of the 5 'region of the nucleic acid molecule provided in (i); 2. At least one of the primers is capable of at least partially binding to a nucleotide sequence region covalently linking the 5'region and the 3 'region of the nucleic acid molecule provided in (i) to sequence the 3' region of the nucleic acid molecule provided in (i); 3. At least one of the primers is capable of binding at least partially to at least a portion of the linker at the 3'end of the molecule to sequence at least a portion of the 3 'region of the nucleic acid molecule provided in (i); and / or 4. At least one of the primers is capable of at least partially binding to a region covalently linked to the 5'region and the 3 'region of the nucleic acid molecule provided in a to sequence the 5' region of the nucleic acid molecule provided in (i).
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present invention relates to a method of determining the sequence of a nucleic acid molecule, comprising the identification of methylated cytosines (e.g. cytosine nucleic acid methylation (5mCs) in any sequence context (e.g. CpG, CHG, CHH, wherein H = A, C or T). The present invention also relates to computer programs and kits associated with the method of the present invention. BACKGROUND

[0002] The analysis of the primary structure of nucleic acids (such as DNA and RNA), including epigenetic modifications (such as DNA or RNA methylation), can be processed by using different techniques commonly referred to as "sequencing".

[0003] Single-end (SE) sequencing refers to sequencing a nucleic acid from only one end of the insert; this is the earliest and simplest method for high-throughput sequencing. A simple modification to the standard SE library preparation process allows the forward and reverse template strands of each cluster to be read in a single paired-end run.

[0004] In addition to generating twice as many reads as single-end reads using the same amount of nucleic acid (input) at the same library preparation time and effort, aligning sequences in read pairs enables more accurate read alignment and enables detection of inversions and indel variants, which is not possible with single-end read data (Nakazato T, Ohta T, Bono H. Experimental design-based functional mining and characterization of high-throughput sequencing data in the sequence read archive. PLoS One. 2013;8(10):e77910).

[0005] Paired-end (PE) sequencing enables sequencing of both ends of a nucleic acid template. Because both end reads contain long-range positional information, highly accurate read alignment is possible. If the sample is prepared, because the distance between each pair of reads is known, alignment algorithms can use this information to more accurately map reads to repetitive regions. This results in better read alignment, especially in repetitive regions of the genome that are difficult to sequence.

[0006] If one of the reads in a PE maps to a different region of the genome than its mate, but its mate does not map to a different region of the genome, the first read is likely to be in close proximity to its mate (following the insert size distribution known from library preparation). Even if the edge alignment probability of each read is the same, some pair alignments are more likely than others, and thus, this approach works equally well even if both reads are multiply mapped.

[0007] Not only does PE provide twice the amount of DNA sequencing to find breakpoints of structural variations, but it also allows to estimate whether there is a breakpoint between two paired reads if they are further apart than expected, or located on different chromosomes. With SE reads, a breakpoint of a structural variation can only be detected if a read overlaps with the breakpoint. The principle is similar to chromosome conformation capture and its ligation sites, or RNA containing splice sites.

[0008] In summary, the advantages of paired-end reads over single-end reads are: Higher accuracy: able to detect errors and gaps in the sequencing data. Any error in one read can be corrected with the other read.

[0009] Longer read length: because the sequence of the original fragment can be inferred using the combination of both read lengths.

[0010] Better genome assembly: because it provides information on the orientation and distance between fragments, which helps to identify insertions and deletions in the genome.

[0011] Improved detection of structural variations: because it collects information on the location and size of these variations (insertions, deletions, and inversions).

[0012] Next Generation Sequencing (NGS) technologies are quite accurate, but not perfect, and errors are introduced at various steps of the whole workflow, for example, see Ma, X., Shao, Y., Tian, L. et al., Analysis of error profiles in deep next-generation sequencing data. Genome Biol 20, 50 (2019) in Figure 1 : Sample handling: in sample collection, storage, nucleic acid isolation; Library preparation: through fragmentation, ligation, adapter contamination, index jumping, amplification-PCR errors and biases, and capture biases; Sequencing itself: errors in base calling through inaccurate detection of emitted fluorescent signals or through misincorporated reversible terminators; Bioinformatics data analysis: error alignment at read mapping, repeat sequence identification, incorrect variant calling and insertion / deletion (INDEL) assembly errors.

[0013] The raw NGS error rate is on average 10 -3 but varies depending on the input sample type, technology, instrument, sequence context, position of the base in the read and type of nucleotide substitution.

[0014] For example, C>T / G>A and A>G / T>C substitutions account for about 70% of the total number of mutated nucleotides, see Zhang Z, Gerstein M. Patterns of nucleotide substitution, insertion and deletion in the human genome inferred from pseudogenes. Nucleic Acids Res. 2003 Sep 15;31(18):5338-48. Figure 1 .

[0015] When methylation is considered, error sources are increased by using the deamination process, by creating C>T / G>A ambiguities (WGBS) or by inappropriate or failed conversion rates.

[0016] These errors are a key confounder for the use of NGS to detect low-frequency genetic variants, which are crucial for cancer molecular diagnostics, therapy and monitoring.

[0017] A new sequencing technology called Genomic and Epigenomic Unified Sequencing (GEUS) (as described in WO 2015 / 104302) can take into account methylation and greatly reduce the error rate by querying the same position of the original sequence from different contexts of the two related strands.

[0018] In the GEUS technology (as described in WO 2015 / 104302), the 5 real bases of the original molecule, for example A, T, G, unmethylated C (C) and methylated C (M), are inferred from the following 2-letter matching code:

[0019] Any other combination that does not match the 2-letter code can only be the result of an error in one of the steps of the process and will thus be detected and removed (for example, N can be assigned to ignore that particular position, or a decision tree based on the base quality (BQ) of each nucleotide can be used to temporarily assign an ambiguity to a certain base or to two bases, to be finally resolved (inferred base recalibration step) when the read is mapped to the reference sequence and the reference nucleotides are known).

[0020]

[0021] With this high-precision approach, most types of nucleotide substitutions, and in particular short INDELs, can be detected and removed, thus greatly reducing the error rate.

[0022] For example, to mistake a T for a G, two errors are required to produce another matching 2-letter code at some positions where the two reads correspond to the same base in the original template, thus this type of error is greatly reduced (by about 10,000 times) under optimal conditions. This requires a T to C (T>C) error in read 1, plus a T to A (T>A) error in read 2, to produce a false positive in one of the inferred reads.

[0023]

[0024] If only one of the two errors occurs, the error will produce a non-matching 2-letter code, which will be detected and removed.

[0025]

[0026] Another example:

[0027] If two errors occur, but both are not in the direction that would produce a matching 2-letter code, the code will also be detected and removed. For example:

[0028] However, due to the ambiguity of the 2-letter matching code, some types of substitution errors (which are precisely the most common changes in DNA, such as G>A / C>T and T>C / A>G) need to be further reduced.

[0029] Furthermore, as described in WO 2015 / 104302, in the GEUS technology, the molecules that are sequenced (e.g. referred to as "GEUS molecules") have two regions that independently provide information about the base identity at the corresponding loci in the original nucleic acid molecule, e.g. see the molecules described in claim 1. However, when such molecules are subjected to conventional PE sequencing, only single-end sequencing results are actually obtained. This is due to the fact that the two regions containing the relevant information are present on the same molecule. Thus, only one end of each region of the GEUS molecule can be read, while the information from the other end provided by traditional paired-end sequencing is lost, e.g. see Figure 7 A.

[0030] Therefore, there is a need for further sequencing methods to reduce the error rate when sequencing nucleic acid molecules, such as genomic DNA or RNA molecules, in particular nucleic acid molecules as described in step i of the present application. SUMMARY

[0031] The present application meets the above needs and provides a new sequencing method that allows for very low error rates. In particular, with the method described herein, for each molecule, a total of more than two (e.g. three, or preferably four) reads covering the insert to be sequenced in parts or completely can be obtained. The method described herein allows for determining the start position and / or the end position of each insert to be sequenced, independent of its size, structural rearrangements or contained base insertions or deletions (INDELs). The method of the present application provides higher quality information and / or lower error rates compared to the sequencing methods of the prior art. Due to the improved quality, the efficiency of the process is also improved, resulting in a faster and less costly process.

[0032] Currently, NGS workflows introduce errors that need to be considered. As described above, each step of the traditional NGS workflow introduces errors that can be attributed to sample handling, library preparation, enrichment PCR, sequencing, mapping, duplication and variant calling. Most of these errors cannot be identified and are hidden in the final results. This negatively impacts the subsequent results processing and analysis. The method of the present application provides more than two (e.g. three, preferably up to four) sources of information for each nucleic acid molecule (up to eight if the original Watson & Crick double stranded molecule is considered, see below), which allows for verifying the read results for each nucleotide, as all read results must agree. Thus, the method of the present application allows for detecting and even correcting errors in the sequence determination (including primary sequence determination and modified nucleotide analysis, such as cytosine methylation analysis).

[0033] The method of the present application is suitable for sequencing nucleic acid molecules comprising a 5' region and a 3' region, wherein the 5' region and the 3' region are covalently linked by a nucleotide sequence that can bind to a primer, wherein the base identity of one of the 5' region or the 3' region independently of the base identity in the other region provides information about the base identity in the corresponding locus in the original nucleic acid molecule, and wherein the molecule further comprises: one linker located at the 5' end of the molecule; one linker located at the 3' end of the molecule.

[0034] In the present specification, such a molecule is referred to as "a molecule as defined in step i of the invention" or "a molecule according to the invention" or the like. A particular example is the so-called "GEUS molecule" described in WO 2015 / 104302.

[0035] In particular, the present invention provides a method, the method comprising: i. providing a nucleic acid molecule comprising a 5' region and a 3' region, wherein the 5' region and the 3' region are covalently linked by a known nucleotide sequence that can bind to a primer; wherein the base identity of one of the 5' region or the 3' region independently of the base identity in the other region provides information about the base identity in the corresponding locus in the original nucleic acid molecule, wherein the molecule further comprises: one linker located at the 5' end of the molecule; one linker located at the 3' end of the molecule; ii. sequencing the molecule provided in step (i) using at least two different primers, preferably at least three different primers, more preferably four different primers, wherein the at least two different primers bind to at least three, preferably at least four different regions in the nucleic acid molecule provided in step (i), wherein: 1. at least one of the primers is capable of at least partially binding to at least a part of the linker at the 5' end of the molecule to sequence at least a part of the 5' region of the nucleic acid molecule provided in (i); 2. at least one of the primers is capable of at least partially binding to a region of the nucleotide sequence covalently linking the 5' region and the 3' region of the nucleic acid molecule provided in (i) to sequence at least a part of the 3' region of the nucleic acid molecule provided in (i); 3. at least one of the primers is capable of at least partially binding to at least a part of the linker at the 3' end of the molecule to sequence at least a part of the 3' region of the nucleic acid molecule provided in (i); and / or 4. At least one of the primers is capable of at least partially binding to a region of the nucleotide sequence of the 5' region and the 3' region of the nucleic acid molecule provided in (a) to sequence at least a portion of the 5' region of the nucleic acid molecule provided in (i).

[0036] In one embodiment, step ii comprises sequencing the molecule provided in step (i) using at least two different primers, preferably at least three different primers, wherein the at least two different primers bind to at least three different regions in the nucleic acid molecule provided in step (i), wherein: 1. At least one of the primers (e.g. a first primer) is capable of at least partially binding (hybridising) to at least a portion of the adaptor at the 5' end of the nucleic acid molecule provided in (i) to sequence at least a portion of the 5' region of the nucleic acid molecule provided in (i); 2. At least one of the primers (e.g. a second primer) is capable of at least partially binding to at least a portion of the adaptor at the 3' end of the molecule to sequence at least a portion of the 3' region of the nucleic acid molecule provided in (i); and 3. At least one of the primers (e.g. a third primer) is capable of at least partially binding (hybridising) to: 3.1. a region of the nucleic acid molecule covalently linked to the 5' region and the 3' region of the nucleic acid molecule provided in (i) to sequence at least a portion of the 3' region of the nucleic acid molecule provided in (i); or 3.2 a region of the nucleic acid molecule covalently linked to the 5' region and the 3' region of the nucleic acid molecule provided in (a) to sequence at least a portion of the 5' region of the nucleic acid molecule provided in (i).

[0037] In a preferred embodiment, the present application provides a method comprising: i. providing a nucleic acid molecule comprising a 5' region and a 3' region, wherein the 5' region and the 3' region are covalently linked by a known nucleotide sequence that can bind to a primer; wherein the base identity of one of the 5' region or the 3' region independently of the base identity in the other region provides information about the base identity in the corresponding locus in the original nucleic acid molecule, wherein the molecule further comprises: - one adaptor located at the 5' end of the molecule; - one adaptor located at the 3' end of the molecule; ii. Sequencing the molecule provided in step (i) using at least two different primers, such as at least three different primers, preferably four different primers, wherein the at least two different primers, preferably four different primers, bind to at least four, preferably four different regions of the nucleic acid molecule provided in step (i), wherein: 1. At least one of the primers is capable of binding (hybridizing) at least partially to at least a portion of the 5' end linker of the nucleic acid molecule provided in (i) to sequence at least a portion of the 5' region of the nucleic acid molecule provided in (i); 2. At least one of the primers is capable of binding (hybridizing) at least partially to regions of the nucleotide sequences of the 5' and 3' regions of the nucleic acid molecule provided in (i) to sequence at least a portion of the 3' region of the nucleic acid molecule provided in (i); 3. At least one of the primers is capable of binding (hybridizing) at least partially to at least a portion of the linker at the 3' end of the nucleic acid molecule provided in (i) to sequence at least a portion of the 3' region of the nucleic acid molecule provided in (i); 4. At least one of the primers is capable of binding (hybridizing) at least partially to regions of the nucleotide sequences of the 5' and 3' regions of the nucleic acid molecule provided in (i) to sequence at least a portion of the 5' region of the nucleic acid molecule provided in (i).

[0038] Therefore, using the method of the present invention, in any embodiment thereof, the identity (e.g., A, G, C, T, U, M, or any modification thereof) and / or base quality (BQ, i.e., the probability that the specified identity is true or false) of the true base at a specific position in the original nucleic acid molecule can be determined. Therefore, the method of the present invention preferably includes, in any embodiment thereof, the further step of determining, based on the information provided in step (ii), the identity (e.g., A, G, C, T, U, M, or any modification thereof) and / or base quality (BQ, i.e., the probability that the specified identity is true or false) of the true base at a specific position in the nucleic acid molecule provided in step i.

[0039] The present invention also provides a computer program containing instructions that, when executed by a computer, are able to determine the identity (e.g., A, G, C, T, U, M or any modification thereof) and / or base quality (BQ, i.e. the probability that the specified identity is true or false) of a specific base at a particular position in the original nucleic acid molecule based on the information provided in step (ii) of the method of the present invention.

[0040] The present application also provides a kit comprising at least two different primers, preferably at least three different primers, more preferably four different primers, wherein the at least two different primers, preferably three different primers, more preferably four different primers, are capable of at least partially binding to at least three, preferably at least four, different regions of the nucleic acid molecule provided in step (i) of the method of the present application, wherein: 1. at least one of the primers is capable of at least partially binding to at least a portion of the linker at the 5' end of the molecule to sequence at least a portion of the 5' region of the nucleic acid molecule provided in (i); 2. at least one of the primers is capable of at least partially binding to a region of the nucleotide sequence covalently linking the 5' region and the 3' region of the nucleic acid molecule provided in (i) to sequence at least a portion of the 3' region of the nucleic acid molecule provided in (i); 3. at least one of the primers is capable of at least partially binding to at least a portion of the linker at the 3' end of the molecule to sequence at least a portion of the 3' region of the nucleic acid molecule provided in (i); and / or 4. at least one of the primers is capable of at least partially binding to a region of the nucleotide sequence covalently linking the 5' region and the 3' region of the nucleic acid molecule provided in (i) to sequence at least a portion of the 5' region of the nucleic acid molecule provided in (i).

[0041] The present application also provides a kit comprising at least two different primers, preferably at least three different primers, wherein the at least two different primers, preferably three different primers, are capable of at least partially binding to at least three different regions of the nucleic acid molecule provided in step (i) of the method of the present application, wherein: 1. at least one of the primers (e.g. a first primer) is capable of at least partially binding (hybridizing) to at least a portion of the linker at the 5' end of the nucleic acid molecule provided in (i) to sequence at least a portion of the 5' region of the nucleic acid molecule provided in (i); 2. at least one of the primers (e.g. a second primer) is capable of at least partially binding to at least a portion of the linker at the 3' end of the molecule to sequence at least a portion of the 3' region of the nucleic acid molecule provided in (i); 3. at least one of the primers (e.g. a third primer) is capable of at least partially binding (hybridizing) to: 3.1. a region of the nucleic acid molecule covalently linking the 5' region and the 3' region of the nucleic acid molecule provided in (i) to sequence at least a portion of the 3' region of the nucleic acid molecule provided in (i); or 3.2 the region of the nucleic acid molecule covalently linking the 5' region and the 3' region of the nucleic acid molecule provided in a. to sequence at least a portion of the 5' region of the nucleic acid molecule provided in (i).

[0042] In a preferred embodiment, the present application provides a kit comprising at least two different primers, preferably four different primers, wherein the at least two different primers, preferably four different primers, are capable of at least partially binding to at least four different regions of the nucleic acid molecule provided in step (i) of the method of the present application, wherein: 1. at least one of the primers is capable of at least partially binding (hybridizing) to at least a portion of the linker at the 5' end of the nucleic acid molecule provided in (i) to sequence at least a portion of the 5' region of the nucleic acid molecule provided in (i); 2. at least one of the primers is capable of at least partially binding (hybridizing) to the region of the nucleotide sequence covalently linking the 5' region and the 3' region of the nucleic acid molecule provided in (i) to sequence at least a portion of the 3' region of the nucleic acid molecule provided in (i); 3. at least one of the primers is capable of at least partially binding (hybridizing) to at least a portion of the linker at the 3' end of the nucleic acid molecule provided in (i) to sequence at least a portion of the 3' region of the nucleic acid molecule provided in (i); 4. at least one of the primers is capable of at least partially binding (hybridizing) to the region of the nucleotide sequence covalently linking the 5' region and the 3' region of the nucleic acid molecule provided in (i) to sequence at least a portion of the 5' region of the nucleic acid molecule provided in (i).

[0043] Preferably, any of the above primers at least partially hybridize (bind) to a specific position (region) in the nucleic acid molecule provided in step (i) of the method of the present application when synthesized. BRIEF DESCRIPTION OF DRAWINGS

[0044] Figure 1According to the schematic of the nucleic acid molecule defined in step i of the present application, for example, a GEUS molecule (as described in WO2015 / 104302) is used as template. The 5' region is represented by the nucleotides A-T-T-G-A-A-C-G-C-T from left to right in a gradient line (grey scale). The darker side of the gradient represents the start of the original molecule (5'), the lighter side represents the end of the molecule (3'). The 3' region is represented by the nucleotides A-G-T-G-T-T-T-G-A-T from left to right. Similar to before, the darker side of the gradient represents the start of the original molecule, the lighter side represents the end. The linkers at the 5' and 3' end of the molecule are represented by the grey solid lines. The nucleotide sequence covalently linking the 5' and 3' region is represented in black, it links the 5' and 3' region of the molecule. The optional unique molecular barcodes (UMI) are represented by the white lines. At least two different primers are represented by the thin black arrows, indicating the direction of the sequencing synthesis. The white thick arrows comprising the sequence correspond to the reads synthesized in the 5' to 3' direction after primer hybridization. The reads can have the same or different length. R1 represents read 1, R2 represents read 2, R3 represents read 3, R4 represents read 4.

[0045] Figure 2 In case of full overlap of all reads, the four reads of the nucleic acid molecule shown in Figure 1 are analyzed. The deduced reads ("GEUS deduced" in the figure) are based on the information of 4 sequencing bases. The reads can have the same or different length (number of cycles of the NGS instrument). The DNA template corresponds to the Watson strand of the original double stranded DNA (dsDNA) sequence, see Figure 1 .

[0046] Figure 3 In case of no overlap of the reads, the four reads of the nucleic acid molecule shown in Figure 1 are analyzed. In this case, the deduced reads are based on the information of 2 sequencing bases at each end. The DNA template corresponds to the Watson strand of the original double stranded DNA sequence.

[0047] Figure 4 In case both strands (Watson and Crick strand) of the original double stranded DNA sequence have been sequenced, the four reads of the nucleic acid molecule shown in Figure 1The four reads of the nucleic acid molecule shown were analyzed. In this case, the inferred sequence (reference genome format) is based on information from up to eight sequencing bases from a double-stranded insert that has been processed by two GEUS molecules paired with barcodes (as described in WO2015 / 104302). Note that because the two strands of the double-stranded DNA molecule are antiparallel and complementary, the last base of the Cricket chain inferred position can indicate the first base of the inferred reference position on the GEUS molecule, as shown by the gray gradient.

[0048] Figure 5 A schematic diagram of the method for obtaining the nucleic acid sequence provided in step (i) of the present invention is shown, for example, in WO2015 / 104302. Similar to the previous example, the darker side of the gradient represents the start end of the original molecule, and the lighter side represents the end end of the original molecule.

[0049] Figure 6 A schematic diagram is shown illustrating the sequencing process inside an NGS machine (e.g., an Illumina MiSeq NGS machine) in steps (i)(A) ​​and (ii)(B) of the method of the present invention. (A) is a diagram of a unit in a flow cell where cluster amplification will occur, to which a nucleic acid molecule according to the present invention is attached, in this example a DNAGEUS molecule (as described, for example, in WO2015 / 104302). From bottom to top, the molecule ends contain an NGS linker (Illumina P7 in this example), with a first sample index (for multiplexing samples in the same lane), then an external unique molecular index (UMI), followed by the 5' region of the nucleic acid molecule, then an internal UMI, then a known sequence (i.e., the nucleotide sequence that the primer can bind to, which may be a hairpin if it is a GEUS molecule), a synthesized internal UMI, a 3' region of the nucleic acid molecule (in this example, the synthesized 3' region of the molecule), a synthesized external UMI, and finally an Illumina P5 NGS sequencing linker with a second sample index that prevents index skipping.

[0050] In (B), the sequencing of the first three reads is described step-by-step after each step of primer hybridization (hybridization of primer 1 and sequencing of read 1, hybridization of primer 2 and sequencing of read 2, and hybridization of the sample index primer and the first index read). Then, complementary molecules are synthesized and amplified via cluster synthesis, and the final three additional reads are generated step-by-step in the same manner: primer hybridization and synthesis of read 3, primer hybridization and synthesis of read 4, and primer hybridization and synthesis of the second index sample. The order of read synthesis may vary depending on the instrument, therefore the NGS instrument protocol needs to be adjusted accordingly.

[0051] Figure 7 A) Perform paired-end (PE) sequencing (actually single-end (SE) sequencing) on ​​the molecule provided in step i of the method of the present invention. B) Perform paired-end (PE) sequencing on the molecule provided in step i of the method of the present invention. Detailed Implementation

[0052] This invention relates to a method for determining (e.g., identifying) the sequence of a nucleic acid molecule, comprising determining (e.g., identifying) variations in a nucleic acid molecule (e.g., a fragment of genomic DNA) obtained from a subject, wherein the nucleic acid molecule includes a 5' region and a 3' region, wherein the 5' region and the 3' region are covalently linked by a nucleotide sequence that can bind to a primer, wherein the base identity of one of the 5' region or the 3' region and the base identity in the other region each independently provide information about the base identity at a corresponding locus in the original nucleic acid molecule, wherein the molecule further comprises: - A linker located at the 5' end of the molecule; and - A connector located at the 3' end of the molecule.

[0053] Sequencing nucleic acid molecules can include identifying the bases present at specific loci of the nucleic acid molecule (such as adenine (A), cytosine (C), thymine (T), guanine (G), uracil (U) and their modifications, such as methylcytosine (5mC, 5hmC)).

[0054] The method of this invention, in any of the described embodiments, ensures sequence fidelity and improves sequencing quality, particularly achieving a very low error rate. Using the method described herein, each molecule can obtain a total of four reads (partially or completely covering the insert to be sequenced). Therefore, the information provided by the method of this invention has higher quality and a lower error rate compared to conventional synthesis methods.

[0055] Furthermore, because sequencing is more accurate, less sequencing depth and less material are required to obtain reliable reads.

[0056] The method of the present invention In a first aspect, the present invention provides a method, such as a nucleic acid sequencing method, comprising two steps (i) and (ii).

[0057] Step (i) includes providing Nucleic acid molecule (Hereinafter referred to as "the nucleic acid molecule of the present invention", "the molecule defined in step i of the present invention", "the molecule according to the present invention" or similar names), the nucleic acid molecule comprises a 5' region and a 3' region. In a preferred embodiment, the nucleic acid molecule is a DNA molecule, but it can also be any nucleic acid molecule, such as RNA. The nucleic acid molecule can be provided in the form of multiple nucleic acid molecules, as described in detail below.

[0058] The nucleic acid molecule provided in step (i) of the present application comprises a 5' region and a 3' region. In preferred embodiments, the 5' and / or 3' region of the nucleic acid molecule is obtained from an organism, such as a human or non-human animal, or a plant, a bacterium, a fungus, a yeast and / or a virus, i.e. the 5' and / or 3' region of the nucleic acid molecule is preferably a fragment of genomic DNA (e.g. nuclear DNA, mitochondrial DNA and chloroplast DNA). The 5' and / or 3' region of the nucleic acid molecule of the present application can also be a synthetic nucleic acid, such as synthetic DNA. Thus, in preferred embodiments, the 5' region and / or 3' region of the nucleic acid molecule provided in step (i) is a fragment of genomic DNA. For example, the 5' region and the 3' region of the nucleic acid molecule provided in step (i) can be a fragment of genomic DNA. For example, the 5' region and the 3' region of the nucleic acid molecule provided in step (i) can be a fragment of synthetic DNA. For example, the 5' region of the nucleic acid molecule provided in step (i) can be a fragment of genomic DNA and the 3' region of the nucleic acid molecule provided in step (i) can be a fragment of synthetic DNA. For example, the 3' region of the nucleic acid molecule provided in step (i) can be a fragment of genomic DNA and the 5' region of the nucleic acid molecule provided in step (i) can be a fragment of synthetic DNA. The term "genomic DNA" refers to the heritable genetic information of an organism. Genomic DNA includes DNA of the cell nucleus (also referred to as chromosomal DNA, including cell-free DNA (cfDNA)), as well as DNA of plastids (e.g. chloroplasts) and other organelles (e.g. mitochondria, etc.). The term "genomic DNA" as referred to in the present application includes genomic DNA comprising sequences complementary to the sequences described herein.

[0059] The 5' and / or 3' region of the nucleic acid molecule can be a fragment of plasmid DNA or a single-stranded nucleic acid molecule (e.g. DNA, cDNA, mRNA).

[0060] DNA can be fragmented by any suitable method, including but not limited to mechanical stress (sonication, nebulization, cavitation, etc.), enzymatic fragmentation (enzymatic digestion with restriction enzymes, nicking enzymes, exonucleases, etc.) and chemical fragmentation (dimethyl sulfate, hydrazine, sodium chloride, piperidine, acid, etc.), or in the original organism (e.g. cell-free DNA). In principle, there is no limit to the length of the DNA fragments, but a relatively narrow length range is preferred. The appropriate fragment size can be selected prior to step (i) of the method of the present application. The optimal length ultimately depends on the available sequencing instruments and methods and the desired percentage of read overlap. In more preferred embodiments, the DNA molecule is a fragment of genomic DNA.

[0061] The nucleic acid molecule provided in step (i) of the method of the application can be a single stranded (ss) molecule (e.g. a ss DNA molecule). The nucleic acid molecule provided in step (i) can also be provided as part of a population of nucleic acid molecules or a plurality of nucleic acid molecules, e.g. a population or plurality of DNA molecules as described herein.

[0062] As described above, the nucleic acid molecule (or plurality of nucleic acid molecules) provided in step (i) of the method of the application comprises a 5' region and a 3' region. The 5' region and the 3' region of the nucleic acid molecule are covalently linked by a nucleotide sequence which can bind to a primer. Thus, the nucleic acid molecule of the application comprises a 5' region and a 3' region separated by a third region which is a nucleic acid sequence located between the 5' region and the 3' region of the nucleic acid molecule of the application. Thus, the nucleotide sequence located between the 5' region and the 3' region covalently links the 5' region and the 3' region of the nucleic acid molecule of the application. Thus, the nucleic acid molecule of the application comprises at least three regions: a 5' region, a linking region and a 3' region.

[0063] The nucleotide sequence located between the 5' and 3' regions of the nucleic acid molecule of the application ("linking region") is a region of the molecule which can (at least partially) bind (hybridise) to a primer. Thus, the nucleotide sequence located between the 5' and 3' regions should be long enough to enable a primer to (at least partially) bind (hybridise) thereto, preferably with sufficient specificity that the primer does not significantly bind to other regions of the molecule in order to sequence the 5' and / or 3' regions of the molecule of the application. Methods for designing linking regions as described herein are known to the skilled person.

[0064] For example, the nucleotide sequence located between the 5' region and the 3' region of the nucleic acid molecule of the application and covalently linking the 5' region and the 3' region (e.g. see Figure 1The length of the nucleotide sequence covalently linking the 5' region and the 3' region of the nucleic acid molecule of the application can be at least 5 nucleotides, such as at least 10 nucleotides, or at least 15 nucleotides, or at least 17 nucleotides, such as 17 nucleotides. For example, the length of the nucleotide sequence covalently linking the 5' region and the 3' region can be from 5 to 100 nucleotides, such as from 15 to 100 nucleotides, such as from 15 to 80 nucleotides, such as from 15 to 70 nucleotides, preferably from 15 to 80 nucleotides, more preferably from 17 to 70 nucleotides, even more preferably from 25 to 65 nucleotides, such as 17 nucleotides, or 29 nucleotides or 64 nucleotides. For example, the length of the nucleotide sequence covalently linking the 5' and 3' regions can be at least 20 nucleotides, such as at least 25, 26, 27, 28, 29 or 30 nucleotides. In a preferred embodiment, the length of the nucleotide sequence is at least 17 nucleotides, such as 17 nucleotides, or 18 nucleotides or 19 nucleotides. In another preferred embodiment, the length of the nucleotide sequence is 29 nucleotides. It can also have a longer length, such as at least 35, 40, 45, 50, 55 or at least 60 nucleotides. In another preferred embodiment, the length of the nucleotide sequence is 64 nucleotides, but it can be longer, such as at least 65, 70, 75, 80 or more nucleotides. Thus, the nucleotide sequence covalently linking the 5' region and the 3' region of the nucleic acid molecule of the application can comprise from 5 to 100 nucleotides, preferably from 15 to 80 nucleotides, more preferably from 25 to 70 nucleotides, even more preferably from 29 to 64 nucleotides. Of course, other lengths are possible for sequencing the 5' region and / or the 3' region of the nucleic acid molecule of the application, provided that the primer can bind (hybridize) to the nucleic acid molecule (at least partially); preferably with sufficient specificity that the primer does not substantially bind to other regions of the molecule.

[0065] In the context of the present application, "hybridization" (or "hybridizing") refers to the process in which two single-stranded polynucleotides non-covalently bind (at least partially) to form a stable double-stranded polynucleotide. In the context of the present application, the term "binding" can be used to refer to "hybridization" or "at least partial hybridization".

[0066] Conditions and buffers suitable for hybridization of two single-stranded polynucleotides as outlined above are well known to the skilled person. For example, "hybridization conditions" can include a salt concentration of about 1 M or less, typically less than about 500 mM, and can also be less than about 200 mM. "Hybridization buffer" refers to a buffered salt solution such as 5% SSPE or other such buffers known in the art. Hybridization temperatures can be as low as 5°C, but are typically above 22°C, more typically above 30°C, and often above 37°C. Hybridization is often performed under stringent conditions, i.e. conditions under which a primer is able to hybridize to its target sequence, but not to other non-complementary sequences. Exemplary stringent conditions include a sodium ion (or other salt) concentration of at least 0.01 M to no more than 1 M, a pH of about 7.0 to 8.3, and a temperature of at least 25°C.

[0067] As to composition, the nucleotide sequence covalently connecting the 5' region and the 3' region of the nucleic acid molecule of the application (also referred to as "linker" in the context of the present application) can consist of any base that can be present in a nucleic acid molecule (e.g. A, C, T, G, U, including any modifications thereof, such as methylated C (e.g. 5mC)), provided that a primer can (at least partially) bind (hybridize) thereto in order to sequence the 5' region and / or the 3' region of the nucleic acid molecule of the application. In one embodiment, the nucleotide sequence covalently connecting the 5' region and the 3' region of the nucleic acid molecule of the application comprises at least one nucleotide that can be modified, such as a modified nucleotide, preferably a modified cytosine, more preferably a methylated cytosine (e.g. 5-methylcytosine (5mC), 5-hydroxymethylcytosine (5hmC) or 5-formylcytosine (5fC)).

[0068] In the nucleic acid molecule of the application, the base identity in one of the 5' region or the 3' region and the base identity in the other region each independently provide information about the base identity at the corresponding locus in the original nucleic acid molecule.

[0069] The "original nucleic acid molecule" can also consist of the 5' region and / or the 3' region of the nucleic acid molecule of the application.

[0070] Thus, the 5' region and the 3' region of the nucleotide molecule of the application have a correlation in the sense that both have information about the base identity at the corresponding locus in the original nucleic acid molecule. The information provided by the 5' region sequence is independent from the information provided by the 3' region sequence. Thus, the molecule of the application contains information about the base identity at the corresponding position (locus) of the nucleic acid molecule Two source of information.

[0071] In the context of the present application, a locus refers to a physical site or position within a nucleic acid molecule.

[0072] For example, in the molecules of the application, the 5' region provides information about the identity of the bases at the corresponding loci of the nucleic acid sequence. In addition, the 3' region also independently provides information about the identity of the bases at the same loci of the same nucleic acid sequence.

[0073] Thus, in the nucleic acid molecules of the application, the identity of the bases in one of the 5' region or the 3' region provides information about the identity of the bases at the corresponding loci of the original nucleic acid molecule, while the identity of the bases in the other region (3' region or 5' region, respectively) provides information about the identity of the bases at the same loci of the same original nucleic acid molecule, so that for each locus of the original nucleic acid molecule there is information in both the 5' region and the 3' region of the molecules of the application.

[0074] In a preferred embodiment, the 3' region of the nucleic acid molecule provided in step (i) is at least partially complementary to the reverse strand of the 5' region. In another preferred embodiment, the 5' region of the nucleic acid molecule provided in step (i) is a genomic DNA fragment, and the 3' region of the nucleic acid molecule provided in step (i) is complementary synthesized to the reverse strand of the 5' region. In another embodiment, the 3' region of the nucleic acid molecule provided in step (i) is a genomic DNA fragment, and the 5' region of the nucleic acid molecule provided in step (i) is complementary synthesized to the reverse strand of the 3' region. In another embodiment, the 5' region of the nucleic acid molecule provided in step (i) is a genomic DNA fragment, and the 3' region of the nucleic acid molecule provided in step (i) is also a genomic DNA fragment and is complementary or reverse complementary to the 5' region. In another embodiment, the 3' region of the nucleic acid molecule provided in step (i) is a genomic DNA fragment, and the 5' region of the nucleic acid molecule provided in step (i) is also a genomic DNA fragment and is complementary or reverse complementary to the 3' region. Thus, in the last two examples, both the 3' and 5' regions are part of a genomic DNA fragment and they are complementary between them.

[0075] For example, the sequence of the 5' region of an exemplary nucleic acid molecule of the application can correspond to the sequence of the original nucleic acid molecule. Thus, the 5' region provides information about the identity of the bases in the original nucleic acid molecule. Conversely, the sequence of the 3' region of the nucleic acid molecule of the application can correspond to the sequence of the reverse complementary strand of the 5' region. Thus, the 3' region also provides independent information about the identity of the bases at the corresponding loci of the original nucleic acid molecule. For example, see Figure 4 and Figure 5 .

[0076] For example, the sequence of the 5' region of an exemplary nucleic acid molecule of the application can correspond to the sequence of the original nucleic acid molecule. Thus, the 5' region provides information about the identity of the bases in the original nucleic acid molecule. Conversely, the sequence of the 3' region of the nucleic acid molecule of the application can correspond to the sequence of the reverse complementary strand of the 5' region. Thus, the 3' region also provides independent information about the identity of the bases at the corresponding loci of the original nucleic acid molecule. For example, seeFigure 4 and Figure 5 .

[0077] For example, the sequence of the 5' region of the nucleic acid molecule of the application can correspond to the sequence of the original nucleic acid molecule after treatment of the original nucleic acid molecule with an agent capable of converting unmethylated cytosines into a base that is clearly different from cytosine (e.g. bisulphite), i.e. the sequence of the 5' region of the nucleic acid molecule of the application can correspond to the original nucleic acid molecule after conversion. Thus, the 5' region provides information about the identity of the bases in the original nucleic acid molecule. Conversely, the sequence of the 3' region of the nucleic acid molecule of the application can correspond to the sequence of the reverse complement of the 5' region before C to T conversion (e.g. before bisulphite treatment, see below) of the 5' region. Thus, the 3' region also provides independent information about the identity of the bases in the corresponding locus of the original nucleic acid molecule.

[0078] For example, the sequence of the 5' region of the nucleic acid molecule of the application can correspond to the sequence of the original nucleic acid molecule after treatment of the original nucleic acid molecule with an agent capable of converting unmethylated cytosines into a base that is clearly different from cytosine (e.g. bisulphite), i.e. the sequence of the 5' region of the nucleic acid molecule of the application can correspond to the original nucleic acid molecule after conversion. Thus, the 5' region provides information about the identity of the bases in the original nucleic acid molecule. Conversely, the sequence of the 3' region of the nucleic acid molecule of the application can correspond to the sequence of the reverse complement of the 5' region before C to T conversion (e.g. with bisulphite) of the 5' region. Thus, the 3' region also provides independent information about the identity of the bases in the corresponding locus of the original nucleic acid molecule.

[0079] For example, the sequence of the 5' region of the nucleic acid molecule of the application can correspond to the sequence of the original nucleic acid molecule after treatment of the original nucleic acid molecule with an agent capable of converting unmethylated cytosines into a base that is clearly different from cytosine (e.g. bisulphite), i.e. the sequence of the 5' region of the nucleic acid molecule of the application can correspond to the original nucleic acid molecule after conversion. Thus, the 5' region provides information about the identity of the bases in the original nucleic acid molecule. Conversely, the sequence of the 3' region of the nucleic acid molecule of the application can correspond to the sequence of the reverse complement of the 5' region before C to T conversion (e.g. before bisulphite treatment, see below) of the 5' region. Thus, the 3' region also provides independent information about the identity of the bases in the corresponding locus of the original nucleic acid molecule.

[0080] The skilled person is able to combine the information provided by the 5' region and the 3' region of the nucleic acid molecule of the application to assign information (base identity) to each locus in the original nucleic acid molecule.

[0081] In a preferred embodiment, the 5' region of the nucleic acid molecule of the application provides information about the identity of the bases in the same locus of the original nucleic acid molecule, and the 3' region of the nucleic acid molecule of the application independently provides information about the identity of the bases in the same locus of the original nucleic acid molecule.

[0082] The following scheme provides information and clarification on the nomenclature of the relevant sequences according to the present application: 5' - ATTTGGC - ATTTGGC - 3', both regions (5', ATTTGGC and 3', ATTTGGC) comprise the same sequence in tandem.

[0083] 5' - ATTTGGC - TAAACCG - 3', the 3' region (TAAACCG) has a sequence comprising the complementary base at the corresponding position of the 5' region ("identical complementary" sequence), i.e. the 3' region (TAAACCG) has a sequence complementary to the sequence of the 5' region (ATTTGGC).

[0084] 5' - ATTTGGC - CGGTTTA - 3', the 3' region (CGGTTTA) has a sequence reverse to the sequence of the 5' region (ATTTGGC).

[0085] 5' - ATTTGGC - GCCAAAT - 3', the 3' region (GCCAAAT) has a sequence which is the reverse complement of the 5' region (ATTTGGC); this is the so-called Watson-Crick strand covalently linked through the junction region.

[0086] 5' - ATTTGGU - GUUAAAT - 3', the 3' region (GUUAAAT) has a sequence which, before being treated with, for example, bisulfite or enzymes and thus "converted" (unmethylated C to U, also known as "C to T conversion"), is complementary to the reverse sequence of the 5' region (GCCAAAT). If a methylated C is present in the original molecule, the 5' region will have an A and the 3' region will have a G. Any other modification will be made in the same way, taking into account the original base and the converted base and the relationship between the 5' and 3' regions.

[0087] In the above cases, the "junction region" is indicated with "----".

[0088] Thus, as mentioned above, the nucleic acid molecule of the present application provides two independent sources of information on the true base identity of a specific locus of the original nucleic acid molecule.

[0089] The nucleic acid molecule of the present application also comprises one adaptor at the 5' end of the molecule and one adaptor at the 3' end of the molecule. In the present specification, "adaptor" and "linker" are used interchangeably in the present specification to refer to an oligonucleotide or nucleic acid fragment or segment which can be attached to a target nucleic acid molecule. The "linker molecule" of the method of the present application is preferably a DNA molecule which is compatible at one end with the end of the nucleic acid molecule of the present application (preferably DNA).

[0090] In genetic engineering, adapters refer to short, chemically synthesized single- or double-stranded oligonucleotides that can be ligated to the ends of other DNA or RNA molecules. Adapters can comprise "cleavage sites" (e.g. "restriction sites", i.e. oligonucleotide sequences that can be recognized by restriction enzymes). "Cleavage sites" add a way to adapt the final elements of the library to the needs of different sequencing platforms.

[0091] In one embodiment, at least a portion of the adapters has a sequence that is common to all adapters in the population of nucleic acid molecules of step (i) if this is the case. In this case, the same primer can be used to sequence all molecules.

[0092] Optionally, the adapters comprise unique and combinatorial barcodes (also referred to as "combinatorial sequences", "barcodes", "barcode sequences" or "combinatorial tags") enabling sample identification, multiplexing, pairing and quantitative analysis. The constructs obtained by the method of the application can have barcodes, thereby generating a unique identifier associated with the initial construct, which in turn enables the differentiation of different constructs. The unique identifier allows the identification of the specific construct comprising the identifier and its derivatives. Each unique identifier is associated with a single molecule or a fragment of a single molecule in the starting sample. Therefore, any amplification product of the initial single molecule with the unique identifier is assumed to be homologous. Combinatorial barcodes also allow the quantification of the percentage of each sequence in the sample and help to monitor biases and error control in the amplification step.

[0093] Barcodes sequences add a "bias control" function. When amplification occurs, some fragments can be selectively amplified for a variety of reasons. This undesirable effect is a major problem for quantitative purposes, which are essential for many applications of sequencing, in particular for the analysis of DNA methylation status (as each allele in each cell can have a different methylation status and even the sample can have a heterogeneous composition, which makes quantification and bias control essential for most applications). Therefore, there is at least one barcode sequence that enables bias control. Since each nucleic acid molecule in the plurality of molecules provided in step (i) if this is the case has one or more different barcode sequences, bias control can be performed and selective amplification of a given nucleic acid molecule can be detected.

[0094] Preferably, the adapter molecules and / or barcode sequences are provided as a library of molecules, respectively, wherein each member of the library can be distinguished from the others by a combinatorial sequence within the sequence, as described below.

[0095] As used herein, the term "pool of adaptor molecules and / or barcode sequences" and / or "combinatorial label" refers to a collection of adaptor molecules and / or barcode sequences, wherein each member of the collection can be distinguished from other members by a combinatorial sequence within the adaptor and / or hairpin sequence and / or barcode sequence.

[0096] The terms "combinatorial sequence", "barcode sequence", "barcode" and "combinatorial barcode" are used interchangeably in the present specification to refer to a unique identifier of an individual adaptor sequence or an individual nucleic acid (e.g. DNA) molecule (the barcode sequence itself, not belonging to the adaptor). Preferably, the barcode sequence is comprised in the adaptor. In embodiments, the combinatorial sequence within the adaptor sequence is a degenerate nucleic acid sequence. The combinatorial sequence can comprise any nucleotides, including adenine, guanine, thymine, cytosine, uracil, methylated cytosine (e.g. 5mC or 5hmC) and other modified nucleotides. The number of nucleotides in the combinatorial sequence is preferably designed such that the number of potential sequences and actual sequences represented by the combinatorial sequence is greater than the total number of adaptors in the library. The combinatorial sequence can be located in any region of the adaptor sequence.

[0097] In one particular embodiment, the molecules provided in step i. of the method of the application are molecules as described in WO 2015 / 104302 (also referred to as "GEUS molecules").

[0098] Step (ii) of the method of the application comprises sequencing the molecules provided in step (i) using at least two different primers, such as two, three or four different primers, preferably four different primers. Thus, step (ii) comprises using at least two different primers and sequencing the molecules provided in step (i) using said at least two different primers. Thus, step (ii) provides sequence information of the molecules provided in step (i) of the method of the application. Thus, step (ii) is a sequencing step. For example, see Figure 6 .

[0099] As used herein, the term "primer" refers to a short nucleic acid strand that is at least partially complementary to a sequence in another nucleic acid and serves as a starting point for nucleic acid (e.g. DNA) synthesis. Preferably, the primer has a length of at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 18, at least 20, at least 25, at least 30 or more bases.

[0100] The term "complementary" refers to the base pairing between nucleotides or nucleic acids that allows for the formation of a duplex, e.g., between the two strands of a double stranded DNA molecule, or between an oligonucleotide primer and a primer binding site on a single stranded nucleic acid, or between an oligonucleotide probe and its complementary sequence in a DNA molecule. Complementary nucleotides are typically A and T (or A and U), or C and G. Two single stranded DNA molecules are said to be substantially complementary when the nucleotides of one strand, optimally aligned and compared, are able to pair with the nucleotides of the other strand, with usually at least about 60%, at least 70%, at least 80%, at least 85%, typically at least about 90% to about 95%, and even about 98% to about 100% of the nucleotides of one strand pairing with the nucleotides of the other strand, after appropriate insertion or deletion of nucleotides. The degree of identity between two nucleotide regions is determined using algorithms implemented in computers and methods generally known to those skilled in the art. Preferably, the BLASTN algorithm is used to determine the identity between two nucleotide sequences (BLAST Manual, Altschul, S. et al., NCBI NLM NIH Bethesda, Md. 20894, Altschul, S., S. et al., J., 1990, Mol. Biol. 215:403-410).

[0101] at least two different primers, such as two, three or four different primers, preferably four different primers, bind to at least three, preferably at least four different regions in the nucleic acid molecule provided in (i), wherein: 1. at least one of the primers is capable of at least partially binding to at least a portion of the linker at the 5' end of the molecule to sequence at least a portion of the 5' region of the nucleic acid molecule provided in (i); 2. at least one of the primers is capable of at least partially binding to a region of the nucleotide sequence covalently linking the 5' region and the 3' region of the nucleic acid molecule provided in (i) to sequence at least a portion of the 3' region of the nucleic acid molecule provided in (i); 3. at least one of the primers is capable of at least partially binding to at least a portion of the linker at the 3' end of the molecule to sequence at least a portion of the 3' region of the nucleic acid molecule provided in (i); and / or 4. at least one of the primers is capable of at least partially binding to a region of the nucleotide sequence covalently linking the 5' region and the 3' region of the nucleic acid molecule provided in (i) to sequence at least a portion of the 5' region of the nucleic acid molecule provided in (i).

[0102] In one embodiment, at least two different primers, such as two, three or four different primers, preferably four different primers, bind to at least three, preferably at least four different regions in the nucleic acid molecule provided in (i), wherein: 1. At least one of the primers (e.g. the first primer) is capable of at least partially binding (hybridising) to at least part of the linker at the 5' end of the nucleic acid molecule provided in (i) to sequence at least part of the 5' region of the nucleic acid molecule provided in (i); 2. At least one of the primers (e.g. the second primer) is capable of at least partially binding to at least part of the linker at the 3' end of the molecule to sequence at least part of the 3' region of the nucleic acid molecule provided in (i); 3. At least one of the primers (e.g. the third primer) is capable of at least partially binding (hybridising) to: 3.1. a region of nucleotide sequence covalently linking the 5' region and the 3' region of the nucleic acid molecule provided in (i) to sequence at least part of the 3' region of the nucleic acid molecule provided in (i); or 3.2 a region of nucleotide sequence covalently linking the 5' region and the 3' region of the nucleic acid molecule provided in (a) to sequence at least part of the 5' region of the nucleic acid molecule provided in (i).

[0103] Preferably, at least two different primers, such as two, at least three or at least four different primers, preferably four different primers, bind to at least four, preferably four different regions in the nucleic acid molecule provided in (i), wherein: 1. At least one of the primers is capable of at least partially binding (hybridising) to at least part of the linker at the 5' end of the nucleic acid molecule provided in (i) to sequence at least part of the 5' region of the nucleic acid molecule provided in (i); 2. At least one of the primers is capable of at least partially binding (hybridising) to a region of nucleotide sequence covalently linking the 5' region and the 3' region of the nucleic acid molecule provided in (i) to sequence at least part of the 3' region of the nucleic acid molecule provided in (i); 3. At least one of the primers is capable of at least partially binding (hybridising) to at least part of the linker at the 3' end of the nucleic acid molecule provided in (i) to sequence at least part of the 3' region of the nucleic acid molecule provided in (i); 4. At least one of the primers is capable of at least partially binding (hybridising) to a region of nucleotide sequence covalently linking the 5' region and the 3' region of the nucleic acid molecule provided in (i) to sequence at least part of the 5' region of the nucleic acid molecule provided in (i).

[0104] The primers can at least partially bind (hybridise) to the above sequences under low stringency conditions, preferably under medium stringency conditions, most preferably under high stringency conditions.

[0105] The binding of the primers to at least three, preferably at least four, different regions in the nucleic acid molecule provided in (i) can be performed simultaneously (i.e. all two, three or four primers at the same time), or not simultaneously. Preferably, the binding of the primers to at least three, preferably at least four, different regions in the nucleic acid molecule provided in (i) is not performed simultaneously.

[0106] In preferred embodiments, the binding of the at least two different primers, e.g. two, three or four different primers, preferably four different primers, to at least three, preferably at least four, different regions in the nucleic acid molecule provided in (i) is specific binding. This means that the primers bind to the above-mentioned regions in a specific manner, i.e. it binds to the above-mentioned regions but does not substantially bind to any other region in the nucleic acid molecule provided in (i). The person skilled in the art knows how to design primers and to test their specificity. See, for example, the primer design tool provided by the National Library of Medicine (NIH) (Primer Design Tool (primerdesigntool.nih.gov), or "How To: Design PCR Primers and Check Their Specificity" from the National Institutes of Health (NIH) (designpcrprimersandchecktheirspecificity.nih.gov).

[0107] In preferred embodiments, the at least two different primers binding to at least four different regions in the nucleic acid molecule provided in (i) are four different primers, each primer specifically binding (hybridizing) to regions 1-3 and / or regions 1-4, preferably regions 1-4, as described herein.

[0108] For example, for Illumina sequencing, primer 1 (capable of at least partially binding to at least a portion of the linker at the 5' end of the molecule to sequence the 5' region of the nucleic acid molecule provided in (i)) and primer 2 (capable of at least partially binding to a region of the nucleotide sequence covalently linking the 5' region and the 3' region of the nucleic acid molecule provided in (i) to sequence the 3' region of the nucleic acid molecule provided in (i)) should be different, and primer 3 (capable of at least partially binding to at least a portion of the linker at the 3' end of the molecule to sequence at least a portion of the 3' region of the nucleic acid molecule provided in (i)) and primer 4 (capable of at least partially binding to a region of the nucleotide sequence covalently linking the 5' region and the 3' region of the nucleic acid molecule provided in (i) to sequence the 5' region of the nucleic acid molecule provided in (i)) should also be different. However, primer 1 and 3, or 1 and 4, or 2 and 3, or 2 and 4 as described above are not necessarily different.

[0109] Accordingly, the at least two different primers provided in step (ii) of the method of the application are used for sequencing the molecules provided in step (i). Sequencing can be performed using one or more of the currently available sequencing technologies (e.g. Illumina, Roche, Ion Torrent, etc. sequencing platforms).

[0110] Sequencing can be performed in a low throughput (which includes the analysis of selected fragments) or in a high throughput (also referred to as genome scale), which includes large scale analysis of all or most of the material, e.g. next generation sequencing (NGS) methods. The length of the fragments that can be analyzed depends on the sequencing method employed. The most advanced sequencing technologies currently available aim at performing genome scale sequencing and most locus specific sequencing can evaluate single stranded nucleic acid molecules (e.g. DNA strands) separately.

[0111] The term "sequencing" or the expression "determining a sequence" or the like, such as "determining the identity of a base" or "determining the identity of a base", refers to the determination of information related to the sequence of nucleotide bases of a nucleic acid, in particular to the determination and ordering of a plurality of consecutive nucleotides in a nucleic acid. The information can include the identification or determination of partial sequence information as well as complete sequence information of a nucleic acid molecule. The information refers, for example, to the primary sequence of a DNA molecule, such as a single stranded or double stranded DNA molecule, or epigenetic modifications (e.g. methylation or hydroxymethylation), or both. Sequence information can be determined with different degrees of statistical reliability or confidence. As mentioned above, the method of the application allows obtaining a high degree of confidence when sequencing a nucleic acid molecule.

[0112] The method of the application can be used for sequencing the primary sequence of a DNA molecule (e.g. a single stranded or double stranded DNA molecule) or a pool of DNA molecules. The determination of the primary sequence of a DNA molecule includes the detection of mutations or genetic variations, such as polymorphisms (SNPs, INDELs, etc.). Preferably, the method of the application allows determining the primary sequence and the epigenetic status of the original nucleic acid molecule, e.g. the methylation status of cytosines, in the same read. By analyzing the output of the sequencing, each read will provide information on the primary sequence of the original nucleic acid molecule (including mutations and SNPs) and the methylation sequence, as well as the information on the combined sequence comprised in the nucleic acid molecule adaptor provided in step (i) of the method of the application.

[0113] As described herein, and as will be appreciated by the skilled person, the at least two different primers (e.g. four different primers) bind to at least three, preferably at least four, different regions in the molecules provided in step (i) and at least part of the 5' and 3' regions are sequenced, meaning that: at least part of the 3' region of the nucleic acid molecule provided in step (i) is sequenced using a primer that binds (hybridizes) to at least part of the adaptor at the 3' end of the nucleic acid molecule provided in step (i); and sequencing the 5' region of the nucleic acid molecule provided in step (i) using primers that bind (hybridize) at least in part to at least part of the nucleotide sequence covalently linking the 5' region and the 3' region of the nucleic acid molecule provided in (i), complementary sequence of the nucleotide molecule provided in (i) (cluster amplification by synthesis for parallel sequencing). This is because, using primers, nucleic acid molecules are synthesized by so-called next generation sequencing "sequencing by synthesis" technologies (e.g. Illumina sequencing) which use the synthesis of a primary strand and a complementary strand to read the sequence of a particular nucleic acid molecule. For example, a primer is attached to the forward strand adapter primer binding site and a polymerase adds dNTPs with fluorescent labels to the DNA strand. Since the fluorophore acts as a blocking or synthesis termination group, only one base can be added per cycle; however, the blocking group is reversible. With four-color chemistry, each base has a unique emission spectrum, and after each round of reaction, the instrument automatically records which nucleotide was incorporated. After the color is recorded, the fluorophore is washed away, and the flow cell is then flushed with another dNTP, and the process is repeated. Since the polymerase adds nucleotides to the 3' end of the nucleic acid (DNA) strand, the nucleic acid molecule to be sequenced needs to be read in the 5' to 3' direction. Therefore, sequencing (preferably twice) the 5' and 3' regions of the nucleic acid molecule provided in step (i) using at least two different primers means that the 5' and 3' regions are sequenced using the following primers: (a) at least in part binds (hybridizes) to at least part of the adaptor at the 3' end of the nucleic acid molecule provided in step (i); and (b) at least in part binds (hybridizes) to at least part of the nucleotide sequence covalently linking the 5' region and the 3' region of the nucleic acid molecule provided in step (i), said sequencing occurring on a strand complementary to the nucleic acid molecule provided in step (i).

[0114] In certain cases, the method further comprises diagnosing a condition of the subject based at least in part on the sequencing information provided in step (ii) of the method of the application. The condition can be any condition, feature, or even aging, obesity, and the like. For example, the condition can be cancer, which can be selected from sarcoma, glioma, adenoma, leukemia (e.g. chronic lymphocytic leukemia (CLL)), bladder cancer, breast cancer, colorectal cancer (CRC), endometrial cancer, kidney cancer, liver cancer, lung cancer, melanoma, non-Hodgkin lymphoma, pancreatic cancer, prostate cancer, thyroid cancer, and the like. The condition can also be a neurodegenerative disease, such as Alzheimer's disease, frontotemporal dementia, amyotrophic lateral sclerosis, Parkinson's disease, spinocerebellar ataxia, spinal muscular atrophy, Lewy body dementia, or Huntington's disease. The condition can also be any genetic or environmental disease, or any rare or common disease, or any feature not necessarily associated with a disease.

[0115] The method of the application can further comprise the step of determining the true identity of the base at the particular position (locus) in the original nucleic acid molecule based on the information provided in step (ii). In the context of the present application, "true identity" refers to the identity of the base (e.g. A, C, G, T, U or any modification thereof, e.g. a modified nucleotide, e.g. a modified cytosine, e.g. a methylated cytosine, mC (e.g. 5mC, 5hmC and / or 5fC)) originally present at the particular position (locus) in the original nucleic acid molecule.

[0116] Thus, once the nucleic acid molecule provided in step (i) is sequenced according to step (ii), information about the sequence of at least part of (preferably all of) the 3' and 5' regions is provided. In particular, at least two sources of information for at least one (preferably each) of the 3' region and the 5' region. Since the 5' and 3' regions of the nucleic acid molecule provided in step (i) of the application independently provide information about the identity of the base at the corresponding locus in the original nucleic acid molecule, more than two, e.g. three, preferably all, of the following information is provided in total: Four Figure 1information sources about the base identity of each respective locus in the original nucleic acid molecule. Thus, the method of the present application can be used to reduce the uncertainty and overall error rate in determining the sequence of a polynucleotide (e.g. the original DNA polynucleotide) as described in step i), mainly before alignment with a reference genome (or reference nucleic acid sequence) is required. Thus, the method of the present application is able to provide more than two (e.g. three, preferably up to four) independent information sources (e.g. preferably up to eight independent information sources if a double stranded nucleic acid molecule is considered) from the single stranded nucleic acid molecule as described in step i) which are related to the base identity of each respective locus in the original nucleic acid molecule. The method of the present application provides more than two (e.g. three, preferably up to four) information sources about the base identity of each respective locus in the original nucleic acid molecule (preferably up to eight information sources if a double stranded nucleic acid molecule is considered) as described in step i). Since each of the four nucleotides in the nucleic acid provided in step (i) of the method of the present application can be read in a different sequence context, the method of the present application can reduce errors that are usually caused by the sequence before and after the base to be analyzed. Since each nucleotide of the original molecule is represented more than twice, e.g. three times, preferably up to four times, the original error probability of each base can be greatly reduced, mainly in the pre-alignment step, but also in the post-alignment step. Reducing the error rate in the pre-alignment step can also improve the alignment quality of each read, which again reduces alignment errors and in turn reduces variant calling errors. To know exactly the start and end position of each insert (3’ and 5’ region) can also improve the positioning and calling of SNPs, but mainly the calling of INDELs and other types of rearrangements. Adding UMIs (optional) at the beginning and end of each insert can improve the sequencing quality at the beginning of each read, allow for deduplication (which becomes crucial when enrichment is performed), and reduce the number of non-informative sequencing cycles and unnecessary bioinformatics resources. If the double stranded DNA of the original molecule is separated during the experiment, it can also re-ligate the double stranded DNA.

[0117] Thus, in the method of the present application, if the nucleic acid molecule of step (i) is a double stranded molecule, and the original base quality of each base is 10 -4 , then the minimum error rate of the pre-alignment inferred reads is 10 -16 . This is much lower than the minimum error rate (10 -4 ) obtained with current sequencing methods such as Illumina MiSeq.

[0118] Thus, taking the GEUS molecule as an example (as described in WO 2015 / 104302), if the identity of the first base of read 1 and read 3 and the identity of the second base of read 4 and read 2 do not match any of the following combinations, respectively, then the identity of the true base of the original nucleic acid molecule locus is determined to be misidentified: 1) Adenine and Adenine, corresponding to A, 2) Thymine and Thymine, corresponding to T, 3) Thymine and Cytosine, corresponding to unmodified Cytosine (e.g., unmethylated C), 4) Cytosine and Cytosine, corresponding to modified Cytosine (e.g., methylated C), and 5) Guanine and Adenine, corresponding to G.

[0119] If inferring from 4 reads, when the identity of the first base of read 1 and read 3 and the identity of the second base of read 4 and read 2 do not match any of the following combinations, respectively, then the identity of the true base of the original nucleic acid molecule locus is determined to be misidentified: 1) Adenine, Adenine, Thymine, and Thymine, corresponding to A; 2) Thymine, Thymine, Adenine, and Adenine, corresponding to T; 3) Thymine, Cytosine, Adenine, and Guanine, corresponding to unmodified Cytosine (e.g., unmethylated C); 4) Guanine, Adenine, Cytosine, Thymine, corresponding to G; 5) Cytosine, Cytosine, Guanine, and Guanine, corresponding to modified Cytosine, e.g., methylated C.

[0120] For example, if inferring the identity of the true base of the original nucleic acid molecule locus from 2 reads, there are 16 possible combinations, of which there are 5 possible match codes and 11 impossible match codes (errors that should not be considered). In this case, only one sequencing error would cause an A at the locus of the original nucleic acid molecule to be misidentified (recognized) as a G, or a T at the locus of the original nucleic acid molecule to be misidentified (recognized) as a C.

[0121] If inferring from 2 reads, there are 5 possible "match codes":

[0122] and 11 "non-match codes" (error detection and removal): A A A C C G G G T T C G T A G T C G T A G N N N N N N N N N N N However, if the true identity of the base at the locus of the original nucleic acid molecule is inferred from 4 reads, there are 256 possible combinations, of which there are 5 possible matching codes and 251 impossible matching codes (errors that should not be considered). In this case, at least two sequencing matching errors are required for the A at the locus of the original nucleic acid molecule to be identified as a G (e.g. A>G at R1 + T>C at R4) or a T to be identified as a C.

[0123] If the inference is made from 4 reads, there are 5 possible "matching codes":

[0124] and 251 "non-matching codes" (error detection and deletion or recovery of true identity):

[0125]

[0126]

[0127]

[0128] Furthermore, if the method of the application is performed on each strand of a double stranded nucleic acid molecule (e.g. a dsDNA molecule), a total of eight different sources of information can be provided for each locus in the original genome. This allows to determine the hemimethylation status (or methylation asymmetry of each strand) of the original double stranded molecule. The terms "hemimethylation" and "asymmetrical methylation" can be used interchangeably to refer to a stretch of sequence in a double stranded DNA, e.g. a CpG, wherein only one of the two strands is methylated.

[0129] Figure 2 A schematic of the method of the application is provided, which uses a GEUS molecule (e.g. as described in WO2015 / 104302) as template, wherein both strands of the dsDNA molecule are sequenced using the method of the application. As shown, for each of the 3' and 5' regions of one of the strands, there are two sources of information. Since both the 5' and 3' regions in one strand independently provide information about the identity of the base at the corresponding locus in the original nucleic acid molecule, there are four independent sources of information for the identity of the base at a certain position (locus) in the original nucleic acid molecule. Since the original molecule is a double stranded molecule, there are eight sources of information about the identity of the base at the corresponding locus in the original nucleic acid molecule. See also Figure 3 and Figure 1 and Table 1.

[0130] Table 1 provides an analysis of the sequencing information provided by the method of the application applied to a nucleic acid molecule Providing double stranded nucleic acid molecules .

[0131]

[0132]

[0133] C: un-methylated cytosine, M: methylated cytosine, A: adenine, T: thymine, G: guanine As shown in Table 1, there are 8 sources of information for each locus in the original molecule.

[0134] Thus, if there is only one source of information for each locus in the original molecule, the method of the application can call the identity of the true base with the same quality as the standard method (from 1 source). If there are 2 sources of information for each locus in the original molecule, the method of the application can better identify the true base than the standard method. If there are 3 sources of information for each locus in the original molecule, the method of the application can better identify the identity of the true base than the standard method. Finally, if there are 4 sources of information for each locus in the original molecule, the method of the application can identify the identity of the true base with the best quality.

[0135] Thus, the method of the application can determine the identity of the base, including its epigenetic state, such as hemi-methylation, in the original molecule with reduced error rate.

[0136] In other cases, the method of the application further comprises using a computer comprising a processor, a memory, and instructions stored thereon that, when executed, determine the identity of the base (e.g., the true base) at a particular position (locus) in the original nucleic acid molecule from the sequencing information provided in step (ii).

[0137] It is noted that the processor can comprise one or more processing units, such as microprocessors, GPUs, CPUs, multi-core processors, etc. Similarly, the memory can comprise one or more volatile or non-volatile storage devices, such as DRAM, SRAM, flash memory, read-only memory, ferroelectric RAM, hard disks, floppy disks, tapes, optical disks, etc.

[0138] Thus, the application further provides a computer program comprising instructions that, when executed by a computer, are capable of determining the identity of the true base and / or the BQ score (or error probability) at a particular position (locus) in the original nucleic acid molecule from the information provided in step (ii) of the method of the application.

[0139] Thus, once at least two different primers bind to at least three, preferably at least four, different regions in the nucleic acid molecule of step (ii) of the method of the application, the computer program comprises instructions to perform locus analysis between the different reads to determine the true base at the specific locus of the original nucleic acid to be sequenced. It is noted that different methods of locus analysis can be envisaged by the skilled person in the light of the present specification, all of which are comprised in the present application.

[0140] The present application also provides a computer program comprising instructions which, when executed by a computer, enable any of the methods disclosed in the present document. Such a computer program is thus communicatively connected to the electronic components of the sequencing machine.

[0141] It is noted that the sequencing machine typically comprises the sample, the tray, the incubator, the interchangeable components, the micropipetting system and many other components therein, enabling full automation of the sequencing of the specific nucleic acid molecule. The present embodiments are thus not limited to sequencing machines comprising only these components, but also encompass any device enabling any of the methods disclosed in the present document to be performed automatically, as can be envisaged by the skilled person.

[0142] The computer program product can be implemented in software, hardware or a combination of both. The computer program product can be stored in the memory of the sequencing machine or remotely, e.g. on a remote server communicatively connected to the device.

[0143] Method for obtaining the nucleic acid molecule provided in step (i) of the method of the application Methods for obtaining the nucleic acid molecule provided in step (i) of the method of the application are known to the skilled person. Some non-limiting examples are given below.

[0144] The nucleic acid molecule of step (i) can be generated by: a. Figure 5 such as a plurality of double stranded nucleic acid molecules or a population of double stranded nucleic acid molecules. In preferred embodiments, the plurality of double stranded DNA molecules are fragments of genomic DNA. For example, see b1. Covalently linking the forward and reverse single stranded nucleic acid molecules provided in step a. .

[0145] As used herein, “population or plurality of double stranded nucleic acid molecules” refers to a collection of double stranded nucleic acid molecules which can be, but are not limited to, genomic DNA (nuclear DNA, mitochondrial DNA, chloroplast DNA, cfDNA, etc.), plasmid DNA or double stranded DNA molecules obtained from a single stranded nucleic acid sample (e.g. DNA, cDNA, mRNA, etc.). In embodiments, the population is comprised of DNA fragments.

[0146] Preferably, the multiple double-stranded nucleic acid molecules are genomic DNA. This can be the entire genome or a simplified representation of the genome. Genomic DNA includes DNA from the cell nucleus (also known as chromosomal DNA), as well as DNA from plastids (e.g., chloroplasts) and other organelles (e.g., mitochondria) or circulating / free DNA (cfDNA). The term "genomic DNA" as used in this invention includes genomic DNA containing sequences complementary to those described herein.

[0147] DNA can be fragmented by any suitable method, including but not limited to mechanical stress (sonication, atomization, cavitation, etc.), enzymatic fragmentation (digestion by restriction endonucleases, nick endonucleases, exonucleases, etc.), and chemical fragmentation (dimethyl sulfate, hydrazine, sodium chloride, piperidine, acids, etc.), or natural fragmentation, such as cfDNA fragments that enter the bloodstream during apoptosis or necrosis. In principle, there is no limitation on the length of the fragmented DNA fragment, but a narrow length range is preferred for ease of sequencing. A suitable fragment size can be selected prior to step (a) above. The optimal length ultimately depends on the available sequencing method. In a more preferred embodiment, the double-stranded DNA molecule is a fragment of genomic DNA.

[0148] The multiple double-stranded nucleic acid molecules provided in step (a) above can be obtained in the following ways: [1] Provides a population of double-stranded nucleic acid molecules derived from genomic DNA or RNA, etc.; [2] Separate double-stranded nucleic acids (e.g., derived from genomic DNA) to obtain single-stranded nucleic acid molecules (e.g., derived from genomic DNA); [3] Use nucleotides A, G, C, T and U to provide complementary strands of ss nucleic acid molecules to obtain the double-stranded nucleic acid molecules (e.g. ds DNA molecules or RNA molecules) provided in step (a).

[0149] Multiple double-stranded nucleic acid molecules may include nucleic acid molecules in which both strands contain methylated cytosine and / or one of the strands contains methylated cytosine.

[0150] Typically, the ends of groups of double-stranded nucleic acid molecules (such as double-stranded DNA molecules) are processed so that the sample can enter a specific process of the sequencing platform. Optionally, the double-stranded linker ligated in step b (see below) may contain a “cleavage site” (e.g., a “restriction site”, i.e. an oligonucleotide sequence recognized by restriction endonucleases).

[0151] Preferably, the double-stranded nucleic acid molecule (e.g., dsDNA molecule) used in step (a) undergoes end repair prior to step (a). The term "end repair" as used herein refers to converting a nucleic acid (e.g., DNA) fragment containing damaged or incompatible 5' and / or 3' protruding ends into blunt-ended DNA containing a 5' phosphate group and a 3' hydroxyl group. Blunting of DNA ends can be achieved by a variety of enzymes, including but not limited to T4 DNA polymerase (with 5'→3' polymerase activity for filling 5' protruding DNA ends) and the Klenow fragment of E. coli DNA polymerase I (with 3'→5' exonuclease activity for removing 3' protruding ends). For efficient phosphorylation of DNA ends, any enzyme capable of adding a 5'-phosphate to the end of an unphosphorylated DNA fragment can be used, including but not limited to T4 polynucleotide kinase. Preferably, the method for obtaining the nucleic acid molecule provided in step (i) of the present invention further includes a dA tailing step of the nucleic acid (DNA) molecule after the end repair step.

[0152] As used herein, the term "dA tailing" refers to the addition of an A base to the 3' end of a blunt-ended phosphorylated DNA fragment. This treatment creates compatible overhangs for subsequent ligation. This step is performed using methods well known to those skilled in the art, such as using the Klenow fragment of *E. coli* DNA polymerase I. The multiple double-stranded nucleic acid molecules (e.g., double-stranded DNA molecules) used as starting materials in step (a) above can also be obtained from single-stranded molecules (e.g., ss DNA or cDNA). Groups of double-stranded nucleic acid molecules (e.g., double-stranded DNA molecules) can be obtained from cDNA. Double-stranded nucleic acid molecules (e.g., double-stranded DNA molecules) can also be obtained from mRNA (e.g., RNA from viruses) using methods well known in the art, including isolating mRNA, reverse transcribing the RNA into single-stranded cDNA, and processing the single-stranded DNA to obtain double-stranded DNA.

[0153] Samples used to obtain multiple double-stranded nucleic acid molecules (e.g., double-stranded DNA molecules) can be derived from biological or environmental sources. Biological samples include, but are not limited to, animal or human samples, liquid and solid food and feed products (dairy products, vegetables, meat, etc.). Samples being analyzed can originate from a single source (e.g., a single organism, tissue, cell, etc.) or from a nucleic acid library derived from multiple organisms, tissues, or cells. Samples used to obtain multiple double-stranded nucleic acid molecules (e.g., double-stranded DNA molecules) can also be synthetic nucleic acid molecules.

[0154] After completing step a., you can proceed to step b1. and / or step b2: b2. Step b2 of this particular embodiment comprises at least partially linking a double stranded (ds) adaptor to at least one end of the plurality of double stranded nucleic acid molecules,

[0155] wherein the covalent ligation in step b. is performed by means of a nucleotide sequence that can bind to a primer to obtain a nucleic acid molecule comprising a 5' region and a 3' region.

[0156] wherein the 5' region and the 3' region are covalently linked by means of a nucleotide sequence that can bind to a primer, wherein the base identity in one of the 5' region or the 3' region and the base identity in the other region each independently provide information about the base identity in the corresponding locus in the original nucleic acid molecule.

[0157] The "nucleotide sequence that can bind to a primer" has been described in the above description and the definition applies equally to this embodiment.

[0158] For example, the forward and reverse single-stranded nucleic acid molecules provided in step a. can be covalently linked by means of a "hairpin structure" or a "hairpin loop". For example, the forward and reverse single-stranded nucleic acid molecules provided in step a. can be the Watson strand and the Crick strand of a ds DNA molecule.

[0159] The generated molecules can be treated with a reagent that can convert unmethylated cytosines into a base that is clearly different from cytosine in terms of hybridization properties (preferably uracil) in order to analyze the methylation pattern of the sample as will be explained in detail below.

[0160] The nucleic acid molecules of step (i) of the method of the application can also be generated according to the methods described in WO 2015 / 104302. In this case, step a. is identical to the description in step a. above.

[0161] Figure 5 c. Synthesizing a complementary strand for each strand of the nucleic acid molecules obtained in step (b), Preferably, the ligation is to both ends. The ligation reaction is preferably performed under conditions sufficient to ligate the double-stranded adaptor to both ends of the double-stranded nucleic acid molecule (e.g. double-stranded DNA molecule) to obtain a plurality of adaptor-containing nucleic acid molecules (also referred to herein as "adaptor-modified nucleic acid molecules"). For example, see Physical .

[0162] In a preferred embodiment, at least a portion of the double-stranded adaptor has a sequence that is common to all double-stranded adaptors used in step (b).

[0163] In embodiments, the adaptor is a so-called "Y-shaped adaptor". A "Y-shaped adaptor" has a "Y" shape. The terms "Y-shaped adaptor" and "Y-shaped linker" are used interchangeably and in the context of the present invention refer to an adaptor formed by two nucleic acid (preferably DNA) strands, wherein a 3' region of a first nucleic acid strand and a 5' region of a second nucleic acid strand form a double-stranded region by sequence complementarity, wherein the ends of the double-stranded region formed by the 3' region of the first nucleic acid strand and the 5' region of the second nucleic acid strand of the Y-shaped adaptor are compatible with the ends of a double-stranded nucleic acid molecule. The term "3' region" as used herein refers to the region of a nucleotide strand where the 3' end of the strand is located.

[0164] The term "3' end" as used in the present invention refers to the end of a nucleotide strand which has a hydroxyl group on the third carbon of the deoxyribose sugar ring. The term "5' region" as used in the present invention refers to the region of a nucleotide strand which is located at the 5' end of the strand. The term "5' end" as used herein refers to the end of a nucleotide strand which has a fifth carbon atom on the deoxyribose sugar ring. The term "3' region" as used in the present invention refers to the region of a nucleotide strand which is located at the 3' end of the strand. The term "3' end" as used herein refers to the end of a nucleotide strand which is the third carbon atom on the deoxyribose sugar ring.

[0165] In one embodiment, the 3' region of the second nucleic acid (e.g. DNA) strand of the Y-shaped adaptor forms a hairpin loop by hybridization between a first segment and a second segment within the 3' region, wherein the first segment is located at the 3' end of the 3' region and the second segment is located near the 5' region of the second DNA strand. The term "hairpin loop" as used herein refers to a region of DNA formed by unpaired bases when a DNA strand folds and base pairs with another region or segment of the same strand.

[0166] Optionally, the 3' region of the second DNA strand of the Y-shaped adaptor does not form a hairpin loop by hybridization between a first segment and a second segment within the 3' region.

[0167] In the context of the present invention, the ds adaptor (e.g. DNA adaptor, which can or can not be a Y-shaped adaptor) comprises at least one barcode sequence in the ds region of the adaptor. This will provide at least for pairing between each of the original nucleic acid (e.g. DNA) strands of the original double-stranded nucleic acid molecule and enable de-duplication of reads, i.e. distinguishing between reads originating from the same original sequence or independent but starting and ending at the same site, which is crucial for enrichment, in particular for low input / low diversity libraries or high-depth whole genome sequencing. This enables tracking of both strands of each double-stranded nucleic acid fragment originally used in step (a) as described above.

[0168] The term "sequence complementarity" as used herein refers to the property shared by two nucleic acid sequences that, when arranged in antiparallel fashion, the nucleotide bases at each position will be complementary.

[0169] Thus, a plurality of linker-containing nucleic acid molecules is obtained. Since the linker is a double-stranded linker, the 3' region of the first strand and the 5' region of the second strand of the double-stranded linker form a double-stranded region at least by sequence complementarity, wherein the ends of the double-stranded region formed by the 3' region of the first strand and the 5' region of the second strand of the linker are compatible with the ends of the double-stranded nucleic acid molecule.

[0170] If the linker is a Y-shaped linker, the Y-shaped linker can comprise one or more barcode sequences in the 5' region of the first nucleic acid (DNA) strand and / or the 3' region of the second nucleic acid (DNA) strand (and / or the double-stranded region) of the Y-shaped linker formed by the two nucleic acid (DNA) strands. Thus, the barcode sequence can be located in the single-stranded region of the Y-shaped linker molecule and / or the double-stranded region of the Y-shaped linker. In this case, each original nucleic acid (DNA) strand and its synthetic complementary strand (see step (c)) will be paired.

[0171] In a preferred embodiment, prior to step (c), the plurality of paired linker-modified nucleic acid (e.g., DNA) molecules is separated to generate a library of paired linker-modified nucleic acid (e.g., DNA) molecules.

[0172] Figure 4 The "synthetic complementary strand" is obtained by polymerase extension from the 3' end of the second nucleic acid strand of the linker molecule, using each strand of the nucleic acid molecule obtained in step (b2) as a template. Thus, a double-stranded nucleic acid (e.g., DNA) molecule is obtained, wherein the original plus and minus strands of the nucleic acid (e.g., DNA) molecule are physically bound to each other by a hairpin region (i.e., a nucleotide sequence that can bind to the 5' region and 3' region of the covalently linked molecule of the primer). Figure 5 In this particular embodiment, each original strand of the nucleic acid (e.g., DNA) molecule is physically bound to the complementary strand obtained by synthesis. The original and complementary strands correspond to the 5' and 3' regions of the nucleic acid molecule provided in step (i) of the method of the application. For example, see Figure 5 .

[0173] In one embodiment, the linker-containing nucleic acid (e.g., DNA) molecule obtained in step (b2) is treated prior to step (c) under conditions sufficient to separate the strands of the linker-containing molecule. Suitable conditions to separate the strands of the linker-containing molecule can be, but are not limited to, conditions that denature the two strands, for example, heating the molecule to 94-98ºC (20 seconds to 2 minutes), resulting in the breakage of hydrogen bonds between complementary bases, thus generating single-stranded nucleic acid (e.g., DNA) molecules. The separation of the strands is achieved without heating the molecule using isothermal techniques, for example, using a strand-displacing DNA polymerase (for example, but not limited to, the large fragment of Phi29 DNA polymerase or Bacillus caldovelox DNA polymerase).

[0174] After ligation of the ds adaptor in step (b2), each strand of the nucleic acid (e.g. DNA) molecule obtained in step (b2) is converted into a paired double stranded nucleic acid (e.g. DNA) molecule by polymerase extension from the 3' end of the second strand of the adaptor molecule using each strand of the nucleic acid (e.g. DNA) molecule obtained in step (b2) as a template (step (c) above).

[0175] In the context of the present application, the expression "converting each strand into a paired double stranded nucleic acid (e.g. DNA) molecule" means that a strand complementary to each strand is synthesized, wherein both strands are paired. Thus, both strands are physically (covalently) linked (the 3' region of the second strand of the adaptor (e.g. Y-shaped adaptor) forms a hairpin loop by hybridization between a first segment and a second segment within the 3' region, the first segment being located at the 3' end of the 3' region and the second segment being located in the vicinity of the 5' region of the second strand), thereby forming a double stranded conformation, wherein one nucleic acid (e.g. DNA) strand is folded onto itself, as described above, thereby generating the nucleic acid molecule provided in step (i) of the method of the present application.

[0176] Thus, one end of the original nucleic acid strand of the nucleic acid (e.g. DNA) molecule and its synthesized complementary strand are physically linked by a looped structure, as described above, the loop being a nucleotide sequence that can be at least partially bound to a primer. Each of these molecules can also be in a linear conformation when the complementarity between both strands is partially or totally lost, see Amplification Thus, the nucleic acid molecule provided in step (i) of the method of the present application has been generated. In this case, the nucleic acid molecule comprises: a 5' region and a 3' region corresponding to the original strand and its synthesized complementary strand; a nucleotide sequence that can be bound to a primer, the nucleotide sequence being covalently linked to the 5' and 3' regions, the nucleotide sequence at least partially corresponding to the hairpin loop; one adaptor at the 5' end of the molecule, one adaptor at the 3' end of the molecule, the adaptors corresponding to the strands of the adaptor not comprising the hairpin loop.

[0177] For more details, see Step a. Providing double stranded nucleic acid molecules: .

[0178] In the context of the present embodiment, the terms "hairpin adaptor", "hairpin sequence" and / or "hairpin molecule" mean a double stranded formed by a single nucleic acid strand folding onto itself, a double stranded region maintained by base pairing between complementary base sequences on the same strand, a hairpin loop region formed by unpaired bases, and ends compatible with the ends of the molecule provided in step (a). The hairpin adaptor can comprise blunt ends or overhanging ends, preferably overhanging ends.

[0179] As used herein, "polymerase extension" refers to the synthesis of a complementary strand by a DNA polymerase that adds free nucleotides to the 3' end of a second DNA strand of a linker molecule. As described above, the linker molecule can serve as a primer for the extension step. In this step, the temperature is chosen depending on the optimal temperature of the particular DNA polymerase used.

[0180] After step (c), two double stranded nucleic acid (e.g., DNA) molecules are obtained from each nucleic acid (e.g., DNA) molecule containing a linker, and each of the double stranded nucleic acid (e.g., DNA) molecules is formed from the original nucleic acid (e.g., DNA) strand of the nucleic acid (e.g., DNA) molecule ("5' region or 3' region") and a synthetic complementary strand ("3' or 5' region") paired (at least physically) thereto, as described above, by nucleotide sequences that can pair to the primer. They can also pair by barcode sequences as described above.

[0181] Optionally, the complementary strands of the plurality of paired linker-modified DNA molecules obtained in step (c) can be provided using primers whose sequences are complementary to at least a portion of the double stranded linker. The complementary strands of the plurality of paired linker-modified DNA molecules obtained in step (c) can be provided using nucleotides A, G, C, T and modifications thereof, e.g., modified nucleotides, e.g., modified cytosine, mC (e.g., 5mC, 5hmC or 5fC).

[0182] Optionally, the paired double stranded nucleic acid molecules obtained in step (c) are subjected to Figure 5 to provide amplified paired double stranded nucleic acid molecules.

[0183] The pairing between the two strands of the original double stranded nucleic acid (e.g., DNA) molecule can track the two strands of each double stranded nucleic acid (e.g., DNA) fragment that was originally used.

[0184] Thus, each linker can comprise a unique combination barcode, enabling sample identification and multiplexing and quantitative analysis. In a preferred embodiment, the linkers (e.g., Y-shaped linkers) are provided in the form of a library of linkers, wherein each member of the library can be distinguished from the others by the combination sequence located within the double stranded region formed by the 3' region of the first strand and the 5' region of the second strand of the linker.

[0185] Optionally, the Y-shaped linkers can selectively comprise a base labeled with a binding pair second member, thereby restoring the original nucleic acid (e.g., DNA) template after the extension or amplification step. This provides the advantage that the sample used as the nucleic acid (e.g., DNA) template can be identified, preserved during the process, and can be recovered, stored, and subjected to multiple amplifications and sequencing under different conditions, without depleting the sample.

[0186] Optionally, the adaptor (e.g., Y-shaped adaptor) can comprise a "cleavage site" as described above.

[0187] In one embodiment, when the ds adaptor does not comprise a hairpin loop, the molecule generated in step (b2) is contacted with a hairpin adaptor under conditions sufficient to join the hairpin adaptor to the molecule generated in step (b2), as described in detail below. If the adaptor joined in step (b2) does not comprise a hairpin loop, a hairpin adaptor can be joined to the adaptor joined in step (b2). For example, the hairpin adaptor can be incorporated in the manner described in WO 2015 / 104302.

[0188] Optionally, in the present embodiment of the method of generating the nucleic acid molecule provided in step (i) of the method of the application, the method further comprises the following step (cl) after step (c): contacting each strand of the nucleic acid molecule comprising the adaptor with a complex of an extension primer and a hairpin adaptor under conditions sufficient for the extension primer to hybridize to the second strand of the adaptor molecule, wherein the extension primer comprises a 3' region complementary to the second strand of the adaptor molecule and which region forms an overhanging end upon hybridization to the second strand of the adaptor molecule, and wherein the hairpin adaptor comprises a hairpin loop region and an overhanging end compatible with the overhanging end formed upon hybridization of the extension primer to the second strand of the adaptor.

[0189] Optionally, the method of generating the nucleic acid molecule provided in step (i) of the method of the application further comprises the following steps (c21 and / or c22, respectively) after step (c): converting, in the paired adaptor-modified DNA molecule, unmodified nucleotides (e.g., unmethylated cytosines) in the paired adaptor-modified nucleic acid molecule, if present, to a base that is clearly different from the unmodified nucleotide (e.g., uracil) (c21); and / or converting, in the paired adaptor-modified DNA molecule, modified nucleotides (e.g., methylated cytosines) in the paired adaptor-modified nucleic acid molecule, if present, to a base that is clearly different from the modified nucleotide (e.g., cytosine) (c22).

[0190] For example, optionally, the method of generating the nucleic acid molecule provided in step (i) of the method of the application comprises the following steps (c21 and / or c22, respectively) after step (c): converting, in the paired adaptor-modified DNA molecule, unmethylated cytosines in the paired adaptor-modified nucleic acid molecule, if present, to a base that is clearly different from cytosine (e.g., uracil) (c21); and / or In DNA molecules modified with paired linkers, methylated cytosine (if present) in nucleic acid molecules modified with paired linkers is converted into a base (c22) that is distinctly different from cytosine.

[0191] Therefore, step c2 (in any variant thereof) specifies the conversion of unmodified (e.g., unmethylated) (c21) or modified (e.g., methylated) (c22) nucleotides (e.g., cytosine) in the nucleic acid molecule provided in step (i) of the method of the present invention.

[0192] For example, in this alternative embodiment, the nucleic acid molecules obtained in step (c) are treated with a reagent capable of converting (or transforming) one nucleotide into another nucleotide that is distinctly different from the original nucleotide (e.g., a reagent capable of converting unmethylated cytosine into a base (preferably uracil) that is detectable in hybridization properties to be distinctly different from cytosine) in order to analyze the epigenetic modification status (e.g., methylation pattern) of the sample (optional step c21).

[0193] In the context of this invention, the term "epigenetic modification" refers to any chemical modification that may be present in one or more nucleotides but does not alter the nucleotide sequence (does not change the genetic code sequence of the nucleotide). Epigenetic modifications, or "tags," such as DNA methylation, alter DNA accessibility and chromatin structure, thereby regulating gene expression patterns. Therefore, in the context of this invention, epigenetic modifications can be present in any nucleotide, i.e., A, T, U, G, and / or C in a DNA or RNA sequence. In the context of this invention, "epigenetic modification" may also be used interchangeably with the term "chemical modification" of a nucleotide. Therefore, "modified nucleotide," "chemically modified nucleotide," or "epigenetically modified nucleotide" refers to a nucleotide chemically modified with an epigenetic modification or "tag." Therefore, in the context of this invention, "modified nucleotide" refers to a nucleotide whose structure differs from that of a primary nucleotide (guanine, cytosine, thymine, uracil, or adenine), for example, because it contains an epigenetically modified base. Therefore, a modified nucleotide can be a nucleotide carrying "epigenetic information," i.e., a nucleotide carrying "epigenetic modification," as described above. Preferably, the epigenetically modified base is a methylated base, a hydroxymethylated base, a formylated base, an acetylated base, or a carboxylic acid-containing base. Preferably, the epigenetic modification is methylation, and the modified base is a modified (e.g., methylated) cytosine.

[0194] The epigenetic modification can be cytosine methylation. In this case, the nucleotide sequence (C) is not altered, but the nucleotide cytosine is chemically modified by the introduction of a methyl group. Thus, the cytosine is chemically modified by being methylated. In differentiated mammalian cells, the main epigenetic tag found in DNA is the covalent attachment of a methyl group to the C5 position of cytosine residues in CpG dinucleotide sequences (see, for example, Handy DE. et al., “Epigenetic modifications: basic mechanisms and role in cardiovascular disease”, Circulation, 2011, 123(19):2145-56). Together with histone modifications, DNA methylation regulates chromatin structure and influences the expression of homologous genes by maintaining diverse expression patterns in different cell types. The presence of promoter region DNA methylation is directly associated with transcriptional repression. Conversely, DNA methylation in the body of a gene is positively correlated with gene expression. Notably, epigenetic modifications can also be associated with the onset of disease. For example, 5mC oxidation derivatives can be used as markers for cancer diagnosis and prognosis (see, for example, Chen K., Zhao BS. and He C., “Nucleic acid modifications in regulation of gene expression”, Cell Chem Biol., 2016; 23(1): 74-85). In differentiated mammalian cells, the main epigenetic modification found in DNA is the covalent attachment of a methyl group to the C5 site of cytosine residues in CpG dinucleotide sequences (referred to as CpG), but other cytosines can also be methylated in addition to cytosines in CpG. See, for example, Handy DE. et al., “Epigenetic modifications: basic mechanisms and role in cardiovascular disease”, Circulation, 2011; 123(19):2145-56. Chemical or epigenetic modifications that occur in nucleotides are, for example, 5-methylcytosine (5mC) and its oxidation derivatives (for example 5-hydroxymethylcytosine (5hmC), 5-formylcytosine (5fC) and 5-carboxylcytosine (5caC)) and N 6 - methyladenine (6mA); in messenger RNA and long non-coding RNA 6 - methyladenine (m6A), pseudouridine (psi, Ψ) and 5-methylcytosine (m5C); or in bacterial genomes 4 - methylcytosine (4mC or m 4dC). See Chen K., Zhao BS. and He C., “Nucleic acid modifications in regulation of gene expression”, Cell Chem Biol., 2016; 23(1): 74-85. As mentioned above, besides genomic DNA, RNA molecules also have similar modifications. For example, N 6 - methyladenosine (m 6 A), for example, see Chen K., Zhao BS. and He C., “Nucleic acid modifications in regulation of gene expression”, Cell Chem Biol., 2016; 23(1): 74-85. In addition, pseudouridine is a common component of structural RNAs (transfer RNA, ribosomal RNA, small nuclear RNA, and small nucleolar RNA), for example, see Charette M and Gray MW, “Pseudouridine in RNA: what, where, how, and why”, IUBMB Life, 2000; 49(5): 341-51. Cytosine can also be methylated in RNA to form 5mC. tRNA modifications are known to affect translation, and in turn, different physiological processes. For example, in S. cerevisiae, there are 74 genes involved in installing about 25 chemically distinct modifications at 36 positions in cytoplasmic tRNAs, see Chen K., Zhao BS. and He C., “Nucleic acid modifications in regulation of gene expression”, Cell Chem Biol., 2016; 23(1): 74-85.

[0195] The term “modified cytosine” as used herein refers to a cytosine base modified by replacing or adding one or more atoms or chemical groups, such as a methyl group.

[0196] Modified cytosines (e.g. methylated cytosines) are resistant to treatment with reagents such as bisulfite and A3A because the cytosines remain unchanged (e.g. they are still cytosines) after treatment with these reagents, or because they are converted to a base complementary to guanine after treatment or after replication (e.g. by PCR amplification), and are read as (unmodified, e.g. unmethylated) cytosines (e.g. 5-hydroxymethylcytosine converted to cytosine-5-methylsulfonate, or 5mC / 5fC converted to 5hmC / 5caC after treatment with TET-methylcytosine dioxygenase 2, respectively) in polymerase base amplification and sequencing. Preferably, the modified cytosine is 5-methylcytosine (5mC), 5-hydroxymethylcytosine (5hmC), or 5-formylcytosine (5fC).

[0197] In the context of the present application, the term "nucleotide which can be modified" refers to any nucleotide which can carry an epigenetic modification, as described above. For example, cytosine is a preferred nucleotide which can be modified. Cytosine can be methylated (mC) or unmethylated (uC), i.e. cytosine can carry at least one epigenetic modification (e.g. methylation). Preferably, the nucleotide which can be modified is cytosine.

[0198] In the context of the present application, the term “an agent capable of transforming (or converting) one nucleotide into another nucleotide different from the original nucleotide” refers to any agent (e.g. a reactant, reagent or enzyme) or method or process capable of transforming (i.e. changing the chemical structure of) a particular nucleotide such that the transformed nucleotide is recognized by an enzyme of a replicating nucleic acid molecule (e.g. a polymerase) as another nucleotide different from the original nucleotide. Thus, once the agent has transformed / converted the modified / unmodified nucleotide, the enzyme of a replicating nucleic acid molecule (e.g. a polymerase) will introduce at that position a nucleotide identical to the one it would have introduced if the original nucleotide had not been modified. For example, an agent capable of transforming the base of a nucleotide into another base clearly different from the original base can be bisulfite. Bisulfite is capable of transforming cytosine (C) into uracil (U) by deamination. A polymerase reads uracil (U) differently than cytosine (C), i.e. a polymerase will introduce adenine (A) when reading uracil (U) and not guanine (G) as it would do when reading cytosine (C). There are other agents capable of transforming one nucleotide into another nucleotide in such a way that the reading of the transformed nucleotide is clearly different from the reading of the original nucleotide. Examples of other agents include, but are not limited to, deaminating agents, metabisulfite or cytosine deaminases such as activation-induced cytosine deaminase (AID). For example, beta-glycosyltransferase is capable of glycosylating 5hmCs, while APOBEC3A cytosine deaminase (A3A) is capable of deaminating uCs to Us. TET methylcytosine dioxygenase 2 is capable of oxidizing 5mC to 5hmC or 5fC or 5caC. For example, an agent capable of transforming the base of a nucleotide into another base clearly different from the original base can be the AID / APOBEC enzyme family, see for example, Berney, M. and McGouran, J.F., “Methods for detection of cytosine and thymine modifications in DNA”, Nat Rev Chem, 2018 , 2, 332-348. The AID / APOBEC enzyme family can deaminate mC to T (Nabel CS. et al., “AID / APOBEC deaminases disfavor modified cytosines implicated in DNA demethylation”, Nat Chem Biol., 2012, 8(9):751-8).

[0199] In one embodiment, the agent capable of converting a base of a nucleotide to another base that is significantly different from the original base is bisulfite. As noted above, sodium bisulfite (commonly referred to as "bisulfite") selectively converts unmethylated cytosines to uracils through deamination, while methylated cytosines (5-methylcytosine and 5-hydroxymethylcytosine) remain unchanged. As used herein, the bisulfite ion has its usual meaning HSO3 - Typically, bisulfite is used in the form of an aqueous solution of bisulfite, such as sodium bisulfite (chemical formula NaHSO3) or magnesium bisulfite (chemical formula Mg(HSO3)2). Suitable counterions for the bisulfite compound can be monovalent or divalent. Examples of monovalent cations include, but are not limited to, sodium, lithium, potassium, ammonium, and tetraalkylammonium. Suitable divalent cations include, but are not limited to, magnesium, manganese, and calcium. Treatment of DNA with bisulfite converts unmethylated cytosine bases to uracils, but does not affect 5-methylcytosine bases. The conversion is performed according to standard procedures (Frommer et al., 1992, Proc Natl Acad Sci USA, 89: 1827-31; Olek, 1996, Nucleic Acid Res. 24:5064-6; EP 1394172). Methods of obtaining samples include methods for reduced representation bisulfite sequencing (RRBS).

[0200] In another embodiment, the agent capable of converting a base of a nucleotide to another base that is significantly different from the original base is A3A. In another embodiment, the agent capable of converting a base of a nucleotide to another base that is significantly different from the original base is a beta-glycosyltransferase.

[0201] As used herein, the expression "a base that is detectably different from a certain nucleotide in terms of hybridization properties" means a base that cannot hybridize (will not form hydrogen bonds) with the base in the complementary strand that would otherwise be complementary to it (e.g., if the original base is guanine, then adenine; if the original base is cytosine, then uracil).

[0202] As used herein, the expression "a base that is detectably different in hybridization properties from cytosine" means a base that is not capable of hybridizing (no hydrogen bond present) to guanine in a complementary strand (e.g. uracil). Preferably, the base that is detectably different in hybridization properties from cytosine is thymine or uracil, more preferably uracil. The reagent used in this step can be a reagent that is capable of converting an unmethylated cytosine to a base that is detectably different in hybridization properties from cytosine, but does not act on a methylated cytosine. As mentioned above, examples of such reagents include, but are not limited to, a deaminating agent, bisulfite, metabisulfite or a cytidine deaminase, such as activation-induced cytidine deaminase (AID). In a preferred embodiment, the reagent is bisulfite.

[0203] Preferably, the conversion of (unmethylated) cytosine to uracil in the paired DNA molecules is achieved by a deaminating agent such as bisulfite, but any other reagent or enzymatic treatment as described above (e.g. oxidation of the cytosine with TET followed by deamination of the non-modified cytosine with APOBEC) can also be used.

[0204] In one particular embodiment of step (c2), if a cytosine is present in one of the strands of the paired double stranded nucleic acid molecule obtained in step (c2) or step (d or e) and a guanine is present in the corresponding position of the other strand of the paired double stranded nucleic acid molecule, it is determined that a modified cytosine (e.g. 5C modified cytosine, e.g. methylated cytosine) is present at the given position; and / or if a uracil or thymine is present in one of the strands of the paired double stranded nucleic acid molecule obtained in step (c2) or step (d or e) and a guanine is present in the corresponding position of the other strand of the paired double stranded nucleic acid molecule, it is determined that an unmethylated cytosine is present at the given position.

[0205] WO 2015 / 104302 describes one exemplary embodiment of the molecules provided in step i. of the present application.

[0206] Kits of the application The present application also provides a kit comprising at least two different primers, wherein the at least two different primers are capable of binding at least partially to at least three, preferably at least four different regions in the nucleic acid molecule provided in step (i) of the method of the present application.

[0207] In one embodiment, the at least two, e.g. three or preferably four different primers are capable of binding at least partially to at least three, preferably at least four different regions in the nucleic acid molecule provided in step (i) of the method of the present application, wherein: 1. at least one of the primers is capable of at least partially binding to at least a portion of the linker at the 5' end of the molecule to sequence at least a portion of the 5' region of the nucleic acid molecule provided in (i); 2. at least one of the primers is capable of at least partially binding to a region of the nucleotide sequence covalently linking the 5' region and the 3' region of the nucleic acid molecule provided in (i) to sequence at least a portion of the 3' region of the nucleic acid molecule provided in (i); 3. at least one of the primers is capable of at least partially binding to at least a portion of the linker at the 3' end of the molecule to sequence at least a portion of the 3' region of the nucleic acid molecule provided in (i); and / or 4. at least one of the primers is capable of at least partially binding to a region of the nucleotide sequence covalently linking the 5' region and the 3' region of the nucleic acid molecule provided in (a) to sequence at least a portion of the 5' region of the nucleic acid molecule provided in (i).

[0208] In one embodiment, at least one of the primers (e.g., a first primer) is capable of at least partially binding (hybridizing) to at least a portion of the linker at the 5' end of the nucleic acid molecule provided in (i) to sequence at least a portion of the 5' region of the nucleic acid molecule provided in (i). In addition, at least one of the primers (e.g., a second primer) is capable of at least partially binding to at least a portion of the linker at the 3' end of the molecule to sequence at least a portion of the 3' region of the nucleic acid molecule provided in (i). Finally, at least one of the primers (e.g., a third primer) is capable of at least partially binding (hybridizing) to: - a region of the nucleic acid molecule covalently linking the 5' region and the 3' region of the nucleic acid molecule provided in (i) to sequence the 3' region of the nucleic acid molecule provided in (i); or - a region of the nucleic acid molecule covalently linking the 5' region and the 3' region of the nucleic acid molecule provided in (a) to sequence at least a portion of the 5' region of the nucleic acid molecule provided in (i).

[0209] In a preferred embodiment, the present application provides a kit comprising at least two different primers, e.g., at least three different primers, preferably four different primers, wherein the at least two different primers, e.g., at least three different primers, preferably four different primers, are capable of at least partially binding to at least four different regions of the nucleic acid molecule provided in step (i) of the method of the present application, wherein: 1. at least one of the primers is capable of at least partially binding (hybridizing) to at least a portion of the linker at the 5' end of the nucleic acid molecule provided in (i) to sequence at least a portion of the 5' region of the nucleic acid molecule provided in (i); 2. At least one of the primers is capable of at least partially binding (hybridizing) to a region of the nucleotide sequence covalently linking the 5' region and the 3' region of the nucleic acid molecule provided in (i) to sequence at least a portion of the 3' region of the nucleic acid molecule provided in (i); 3. At least one of the primers is capable of at least partially binding (hybridizing) to at least a portion of the linker at the 3' end of the nucleic acid molecule provided in (i) to sequence at least a portion of the 3' region of the nucleic acid molecule provided in (i); 4. At least one of the primers is capable of at least partially binding (hybridizing) to a region of the nucleotide sequence covalently linking the 5' region and the 3' region of the nucleic acid molecule provided in (i) to sequence at least a portion of the 5' region of the nucleic acid molecule provided in (i).

[0210] In a preferred embodiment, the kit of the application further comprises instructions for its use. In another preferred embodiment, the kit of the application further comprises a method for generating a double-stranded linker for the nucleic acid molecule provided in step (i) of the method of the application, wherein the linker comprises a first nucleic acid strand and a second nucleic acid strand, wherein the 3' region of the first nucleic acid strand and the 5' region of the second nucleic acid strand form a double-stranded region by sequence complementarity, wherein the ends of the double-stranded region formed by the 3' region of the first nucleic acid strand and the 5' region of the second nucleic acid strand of the linker are compatible with the ends of the double-stranded nucleic acid molecule, wherein the double-stranded region of the linker comprises one or more barcode sequences, and wherein the 3' region of the second strand of the linker forms a hairpin loop by hybridization between a first segment within the 3' region and a second segment near the 5' region of the second strand, the first segment being located at the 3' end of the 3' region and the second segment being located near the 5' region of the second strand. And / or, wherein the linker comprises at least one barcode sequence in a single-stranded region of the linker, the barcode sequence consisting of a unique identifier that allows the identification of the specific construct comprising the identifier and its amplification products, and wherein compatible means that the ends of the double-stranded region of the linker molecule are capable of being ligated to one or both ends of the double-stranded nucleic acid molecule.

[0211] In another preferred embodiment, the 5' region of the first strand of the linker has a restriction site. In another preferred embodiment, the linker comprises at least one barcode sequence in a single-stranded region of the linker, and wherein the 3' region of the second strand of the linker forms a hairpin loop by hybridization between a first segment within the 3' region and a second segment near the 5' region of the second strand, the first segment being located at the 3' end of the 3' region and the second segment being located near the 5' region of the second strand.

[0212] In another preferred embodiment, the kit further comprises: (i) a library of double-stranded adapters, the adapters comprising a first strand and a second strand, wherein the 3' region of the first strand and the 5' region of the second strand form a double-stranded region by sequence complementarity, and the ends of the double-stranded region are compatible with the ends of the double-stranded nucleic acid molecule; (ii) a plurality of extension primers, wherein each extension primer comprises a 3' region that is complementary to the second strand of the adapter molecule defined in (i), and which region forms an overhanging end upon hybridization to the second strand of the adapter molecule; (iii) a plurality of hairpin adapters, wherein each hairpin adapter comprises a hairpin loop region and an overhanging end that is compatible with the overhanging end formed upon hybridization of the extension primer defined in (ii) to the second strand of the Y adapter defined in (i), wherein the extension primers of (ii) and the hairpin adapters of (iii) can be provided as a complex; wherein the adapters of (i), the extension primers of (ii) and the hairpin adapters of (iii) are suitable for obtaining a library of adapters for a method of generating a nucleic acid molecule provided in step (i) of the method of the application.

[0213] In the context of the present application, the term "sequence identity" refers to the relatedness between two nucleotide sequences or between two amino acid sequences. For the purposes of the present application, the degree of sequence identity between two nucleotide sequences or between two amino acid sequences is determined using the multiple sequence alignment tool Clustal Omega (https: / / www.ebi.ac.uk / Tools / msa / clustalo / ; Sievers, F. et al. 2011 "Fast, scalable generation of high-quality protein multiple sequence alignments using Clustal Omega", Mol. Syst. Biol., 7: 539 Syst.) with standard parameters.

[0214] As used herein, "about" or "approximately" means ±1% of the indicated value, or "about" means ±2% of the indicated value, or "about" means ±5% of the indicated value, or "about" means ±10% of the indicated value, or "about" means ±20% of the indicated value, or "about" means ±30% of the indicated value; preferably, "about" means the exact indicated value (±0%).

[0215] In the description and in the claims, the word "comprising" and its variants, such as "comprise" and "comprises", are used generally and are not to be construed as limiting, unless otherwise indicated by the context. However, whenever the words "including" or "comprising" or variants thereof are used in either the description or the claims, this is meant to also include, in addition to the recited steps, features, components, elements or steps, any additional steps, features, components, elements or steps which are stated in the special embodiment in which the term "including" or "comprising" is used.

[0216] In describing the present application, (especially in the context of the following claims) the use of the terms "a", "an" and "the" and similar referents are to be construed to cover both the singular and plural unless otherwise indicated by the context. The use of the terms "comprising", "comprise" and "comprises" and the like in the description and in the claims are used to mean that the methods, compositions and processes so described contain the recited elements, but not necessarily to the exclusion of others. The use of the terms "including" and / or "having" and the like in the description and in the claims are used to mean that the methods, compositions and processes so described contain the recited elements, but not necessarily to the exclusion of others. The use of the term "or" in the description and in the claims is used to mean "and / or" unless otherwise indicated by the context.

[0217] The use of any example, or exemplary language (e.g., "for instance", "such as") provided herein, is intended merely to better illuminate the present application and does not in any way limit the scope of the present application, unless otherwise indicated by the context. No language in the specification should be construed as indicating any non-claimed element as essential to the practice of the present application.

[0218] While the foregoing application has been described in some detail for purposes of clarity and understanding, it will be appreciated by one skilled in the art, upon reading this disclosure, that various changes in form and detail can be made without departing from the true scope of the application and the appended claims.

[0219] The present application is described below by way of the following examples, which are provided for purposes of illustration only, and do not limit the scope of the present application in any way.

[0220] Examples Step i. of the method of the present application: This section provides an example of how to provide a molecule of step i of the claimed method. In particular, the molecule exemplified herein is a GEUS molecule described in WO 2015 / 104302. As described above, the molecule is an example of a molecule defined in step i of the method of the present application. The advantages and effects of the GEUS molecule described herein apply equally to any molecule defined in step i of the method of the present application.

[0221] The nucleic acid molecule of step i. can be generated by: Step b2. At least partially linking a double stranded (ds) adaptor to at least one end, preferably to both ends, wherein one of the adaptors comprises a hairpin structure (referred to as "hairpin" in the following representation): For example, see Figure 5 A: ATCGAAMGMT TAGCTTGMGA ("M" denotes a methylated cytosine, C denotes a non-methylated cytosine) Step c. Synthesizing a complementary strand for each strand of the nucleic acid molecules obtained in step (b), Figure 5 See, for example, Step c. 21 : Converting unmethylated cytosines in the molecules obtained after step c) into a base that is clearly different from cytosine (e.g. uracil / thymine), e.g. by bisulfite treatment: B.

[0222] Linker - ATCGAAMGMT - Hairpin Hairpin - TAGCTTGMGA - Linker Figure 5 "synthetic complementary strand", by polymerase extension from the 3' end of the "hairpin", using each strand of the nucleic acid molecule obtained in step (b2) as a template. See, for example, AG D.

[0223] Linker - ATCGAAMGMT - Hairpin - AGCGTTCGAT Linker - AGMGTTCGAT - Hairpin - ATCGAACGCT GTT GAT For simplicity, we will continue this example using just one of the molecules obtained after step c. However, it is worth noting that the method can continue using both, see Figure 6 E.

[0224] Linker - AT U GAACGCT - Hairpin - Primer 1 (p1): U Primer 2 (p2): U Primer 3 (p3): - Linker (underlined are the converted but not methylated cytosines) The resulting molecule is a "GEUS molecule" as described in WO 2015 / 104302, comprising a 5' region (bold) and a 3' region (italic), wherein these two regions are covalently linked by a nucleotide sequence that can bind to a primer (comprising a linker hairpin structure as described above, also referred to as a "junction region"), wherein the base identity of one of the 5' region or the 3' region and the base identity of the other region each independently provide information about the base identity in the corresponding locus in the original nucleic acid molecule (in this example, they provide information about the base identity of the original strand "ATCGAAMGMT" of step a), and wherein the molecule comprises one linker at the 5' end and another linker at the 3' end. The molecule is an example of the molecule provided in step i. of the method of the application, as shown in Primer 4 (p3): A.

[0225] The method step ii. to which the application relates The molecule provided in step i is sequenced using at least two primers, which can bind to at least three, preferably at least four, different regions. In the present example, four different primers will be used to bind to four different regions of the nucleic acid molecule provided in (i), as described below. However, it will be clear to the skilled person that at least some of the advantages and effects described herein are equally applicable to methods using at least two different primers (e.g. three different primers) to bind to at least three different regions in the nucleic acid molecule provided in (i).

[0226] As mentioned above, four different primers are used in the present example: Figure 6 capable of binding to a part of the 5’ end adaptor of the GEUS molecule to sequence at least part of the 5’ region of the provided GEUS molecule (corresponding to the primer of item 1, step ii. of the claimed method).

[0227] Figure 7 capable of binding to the region of the hairpin structure to sequence the 3’ region of the GEUS molecule (corresponding to the primer of item 2, step ii. of the claimed method).

[0228] Figure 6 capable of binding to the 3’ end adaptor of the GEUS molecule to sequence at least part of the 3’ region of the GEUS molecule (corresponding to the primer of item 3, step ii of the claimed method).

[0229] Figure 7 capable of binding to the region of the hairpin structure to sequence the 5’ region of the GEUS molecule (corresponding to the primer of item 4, step ii. of the claimed method).

[0230] The above four primers will result in four reads: see also Figure 7 B.

[0231] p1 : read 1 ATTGAACGCT p3: read 3 ATCAAACACT p4: read 4 TAACTTGCGA p2: read 2 TAGTTTGTGA Importantly, if we only consider read 1 and read 3, which originate from pi and p3, respectively, then in fact, the GEUS molecule will be sequenced based on single-end (SE) sequencing, as read 1 and read 3 provide base information located at the same locus of the original molecule. Since the GEUS molecule comprises two regions (the 5' region shown in bold and the 3' region shown in italics), they provide relevant information, i.e. base identity information about the corresponding locus of the original nucleic acid molecule (strand ATCGAAMGMT in step a. above), thus, using only two regular primers hybridizing to the molecule adaptor (i.e. "regular PE sequencing") would in fact result in SE sequencing, as Figure 7 A.

[0232] However, in the claimed method, two additional read results are obtained, named read 2 and read 4, respectively. The reads originate from primers hybridizing to the GEUS molecule hairpin region (p2 and p4, respectively, see Figure 7 B). By obtaining these two additional reads, true paired-end (PE) sequencing of the 3' and 5' regions of the GEUS molecule is obtained, as now the 5' and 3' regions are read from both ends of each region, as shown in the above figure and ​ B.

[0233] Thus, the method of the present invention allows PE sequencing of a molecule as defined in step i. of the method (e.g. the GEUS molecule described in WO 2015 / 104302). As mentioned above, the advantages of PE sequencing over SE sequencing, see background of the invention.

[0234] Moreover, and more importantly, when these four reads are able to at least partially, preferably completely, cover the 5' and 3' regions of the GEUS molecule, the method claimed herein provides two or more, e.g. at least three, preferably at least four, information sources (i.e. four reads) for each original single-stranded molecule (e.g. if the original molecule is a double-stranded molecule, as in the present example, then a total of up to eight information sources can be provided). Having two or more, e.g. three, preferably four, information sources for each original strand is advantageous, as it allows detection of sequencing errors that would not be detected if only two reads would be obtained, as described above and below: If only pi and p3 are used (see above), the following reads would be obtained (see also ​ A):

[0235] The underlined bases in the read results indicate that they correspond to the true bases in the original DNA template. For example, in read 1, when an "A" is read, it indicates that the corresponding position of the original DNA template is an "A".

[0236] The underlined bases in the reads refer to bases that provide two possible choices upon reading, thus the true identity of the base at that position in the DNA template cannot be inferred from the information source. For example, in read 1, when a "T" is read, it can be either a "T" or an "unmethylated C" in the original DNA template.

[0237] The ambiguity or redundancy problem with the white background bases in read 1 is resolved by reading read 3. For example, the underlined second base "T" in read 1 can be inferred when a "T" appears in read 3. This is because in read 3, "T" indicates a true "T" in the DNA template. Therefore, even if read 1 does not provide information of the second base of the DNA template, read 3 can provide the true identity of the base, thus the sequence of the template DNA can be inferred.

[0238] However, when read 1 has a sequencing error in one of the underlined bases, read 3 cannot resolve the ambiguity and thus incorrectly infers the true identity of the base. The case of the error "G" highlighted with a thick black border in read 1. Although read 3 shows an "A", the base "A" in read 3 can be either a true "G" or a true "A" because read 3 is ambiguous for these bases. Therefore, if read 1 includes a base error that cannot be directly inferred by read 3, the true base identity in the DNA template will be misjudged: the error "G" highlighted with a thick black border will be inferred as a true "G" because "G" in read 1 indicates a true "G" in the DNA template, thus leading to a sequencing error of the DNA template.

[0239] The method of the present application can eliminate these sequencing errors, wherein at least one, preferably at least two additional reads, i.e. read 2 and / or 4, are obtained, see also ​ B:

[0240] In the above table, read 1 and read 3 are similar to read 1 and read 3 shown above. However, two additional reads are added: read 4 and read 2 are reads generated by primers 2 and 4, see ​ B.

[0241] In this case, due to the presence of reads 2 and 4, it can be seen that a base difference occurs because read 4 shows an underlined "T", but the "T" in read 4 indicates a true "T" in the DNA template. Therefore, thanks to true paired-end sequencing, two additional reads are obtained and a sequencing error due to a base difference in the reads can be detected.

[0242] Thus, when the molecules defined in step i of the method of the application are sequenced, the claimed method can achieve very low error rates. This is because each molecule gets more than two reads (e.g. at least three, preferably at least four reads).

[0243] Projects of the application The present application provides the following projects: 1. A method comprising: i. providing a nucleic acid molecule comprising a 5' region and a 3' region, wherein the 5' region and the 3' region are covalently linked by a region of nucleotide sequence that can bind to a primer, wherein the base identity of one of the 5' region or the 3' region independently of the other provides information about the base identity in the corresponding locus in the original nucleic acid molecule, wherein the molecule further comprises: one adaptor at the 5' end of the molecule; one adaptor at the 3' end of the molecule; ii. sequencing the molecule provided in step (i) using at least two different primers that bind to four different regions in the nucleic acid molecule provided in (i), wherein: 1. at least one of the primers is capable of at least partially binding to at least a portion of the adaptor at the 5' end of the molecule to sequence at least a portion of the 5' region of the nucleic acid molecule provided in (i); 2. at least one of the primers is capable of at least partially binding to a region of the nucleotide sequence covalently linking the 5' region and the 3' region of the nucleic acid molecule provided in (i) to sequence the 3' region of the nucleic acid molecule provided in (i); 3. at least one of the primers is capable of at least partially binding to at least a portion of the adaptor at the 3' end of the molecule to sequence at least a portion of the 3' region of the nucleic acid molecule provided in (i); 4. at least one of the primers is capable of at least partially binding to a region of the nucleotide sequence covalently linking the 5' region and the 3' region of the nucleic acid molecule provided in (a) to sequence the 5' region of the nucleic acid molecule provided in (i).

[0244] 2. The method according to project 1, wherein the 3' region of the nucleic acid molecule provided in step (i) is at least partially complementary to the reverse strand of the 5' region.

[0245] 3. The method according to any one of the preceding items, wherein the nucleic acid molecule provided in step (i) is a DNA molecule.

[0246] 4. The method according to any one of the preceding items, wherein the 5' region and / or the 3' region of the nucleic acid molecule provided in step (i) is a fragment of genomic DNA.

[0247] 5. The method according to any one of the preceding items, wherein the 5' region of the nucleic acid molecule provided in step (i) is a fragment of genomic DNA and the 3' region of the nucleic acid molecule provided in step (i) is complementary to the reverse strand of the 5' region.

[0248] 6. The method according to any one of items 1-4, wherein the 3' region of the nucleic acid molecule provided in step (i) is a fragment of genomic DNA and the 5' region of the nucleic acid molecule provided in step (i) is complementary to the reverse strand of the 3' region.

[0249] 7. The method according to any one of the preceding items, wherein the nucleotide sequence covalently linking the 5' region and the 3' region of the nucleic acid molecule provided in step (i) has a length of at least 5 nucleotides, preferably a length of at least 10 nucleotides, more preferably a length of at least 17 nucleotides.

[0250] 8. The method according to any one of the preceding items, wherein in step (ii) the molecule provided in step (i) is sequenced using four different primer pairs.

[0251] 9. The method according to any one of the preceding items, wherein the nucleic acid molecule in step (i) is generated by: a. providing a population of double-stranded nucleic acid molecules, preferably wherein the plurality of double-stranded DNA molecules are fragments of genomic DNA; b. at least partially ligating a double-stranded (ds) adaptor to at least one end of the strands of the plurality of double-stranded nucleic acid molecules, thereby obtaining a plurality of adaptor-containing nucleic acid molecules; c. using each strand of the nucleic acid molecules obtained in step (b) as a template, synthesizing a complementary strand, "synthesizing a complementary strand", to each strand of the nucleic acid molecules obtained in step (b) by polymerase extension from the 3' end of the second nucleic acid strand in the adaptor molecule, thereby pairing each strand of the nucleic acid molecules obtained in step (b) with its synthesized complementary strand, to provide a plurality of adaptor-modified nucleic acid molecules.

[0252] wherein the original nucleic acid strand and its synthesized complementary strand obtained in step (c) are covalently linked by a nucleotide sequence to which a primer can at least partially bind. d. optionally, providing a complementary strand of the plurality of adaptor-modified DNA molecules obtained in step (c), optionally using primers complementary to at least a portion of the sequence of the double-stranded adaptor; e. optionally, amplifying the paired double-stranded nucleic acid molecules obtained in step (d) to provide amplified paired double-stranded nucleic acid molecules.

[0253] 10. The method of item 9, wherein at least a portion of the double-stranded adaptor has a sequence common to all double-stranded adaptors used in step (b).

[0254] 11. The method of one or more of items 9 to 10, wherein prior to step (c), a plurality of paired adaptor-modified DNA molecules are isolated to generate a library of paired adaptor-modified DNA molecules.

[0255] 12. The method of one or more of items 9 to 11, wherein the 3' region of the second DNA strand of the adaptor forms a hairpin loop through hybridization between a first segment within the 3' region and a second segment near the 5' region of the second DNA strand, the first segment being located 3' to the 3' region, the second segment being located near the 5' region of the second DNA strand.

[0256] 13. The method of one or more of items 9 to 12, wherein after step (c), the method further comprises the following step (cl): contacting each strand of the adaptor-containing nucleic acid molecules with a complex of an extension primer and a hairpin adaptor under conditions sufficient for the extension primer to hybridize to the second strand of the adaptor molecule, wherein the extension primer comprises a 3' region complementary to the second strand of the adaptor molecule and the extension primer forms an overhang upon hybridization to the second strand of the adaptor molecule, and wherein the hairpin adaptor comprises a hairpin loop region and an overhang compatible with the overhang formed upon hybridization of the extension primer to the second strand of the adaptor molecule.

[0257] 14. The method of one or more of items 9 to 13, wherein the adaptor has a first barcode sequence in the double-stranded region and / or a second barcode sequence in the 3' region of the second strand of the adaptor.

[0258] 15. The method of one or more of items 9 to 14, wherein the adaptor has a restriction site in the 5' region of the first strand of the adaptor.

[0259] 16. The method according to one or more of items 9 to 15, wherein in step (b) a plurality of adaptor-modified nucleic acid molecules is provided, wherein the strands of the genomic DNA fragments are further paired by using barcode sequences.

[0260] 17. The method according to item 16, wherein the pairing can be performed prior to or after the ligation, or simultaneously with the ligation.

[0261] 18. The method according to one or more of items 9 to 17, wherein at least a portion of the double-stranded adaptors has a sequence in common with all double-stranded adaptors used in step (b).

[0262] 19. The method according to one or more of items 9 to 18, wherein after step (c) the method further comprises the following step (c2): converting unmethylated cytosines in the paired modified adaptor nucleic acid molecules into uracils in the paired modified adaptor DNA molecules.

[0263] 20. The method according to item 19, wherein if a cytosine occurs in one strand of the paired double-stranded nucleic acid molecule obtained in step (d) or step (e) and a guanine occurs at the corresponding position of the other strand of the paired double-stranded nucleic acid molecule, it is determined that a 5C-modified cytosine is present at the given position; and / or if a uracil or thymine occurs in one strand of the paired double-stranded nucleic acid molecule obtained in step (d) or step (e) and a guanine occurs at the corresponding position of the other strand of the paired double-stranded nucleic acid molecule, it is determined that an unmodified cytosine is present at the given position. 21. The method according to any one of items 1 to 8, wherein the nucleic acid molecules in step (i) are generated by: a. providing double-stranded nucleic acid molecules, preferably wherein the double-stranded DNA molecules are fragments of genomic DNA; b. covalently linking the forward and reverse single-stranded nucleic acid molecules provided in step a, wherein the covalent linking in step b is performed by nucleotide sequences that can bind to primers to obtain nucleic acid molecules comprising a 5' region and a 3' region, wherein the 5' region and the 3' region are covalently linked by nucleotide sequences that can bind to primers, wherein the base identity in one of the 5' region or the 3' region and the base identity in the other region each independently provide information about the base identity in the corresponding locus in the original nucleic acid molecule.

[0264] 22. The method according to item 21, wherein the double stranded nucleic acid molecules provided in step a are provided as a population of double stranded nucleic acid molecules, preferably the plurality of double stranded DNA molecules are fragments of genomic DNA.

[0265] 23. The method according to one or more of items 1 to 21, wherein the method further comprises determining the true identity of a base at a specific locus in the original nucleic acid molecule based on the information provided in step (ii).

[0266] 24. The method according to item 23, wherein the identity of the true base at the locus of the original nucleic acid molecule is determined to be false if the identity of the first base from reads 1 and 3 and the identity of the second base from reads 4 and 2 do not match any of the following combinations, respectively: 1) Adenine, Adenine, Thymine and Thymine corresponding to A, 2) Thymine, Thymine, Adenine and Adenine corresponding to T, 3) Thymine, Guanine, Adenine and Guanine corresponding to unmethylated C, 4) Guanine, Adenine, Guanine, Thymine corresponding to G, 5) Cytosine, Cytosine, Guanine and Guanine corresponding to methylated C.

[0267] 25. The method according to one or more of the preceding items, wherein the method further comprises the use of a computer, the computer comprising a processor, a memory and instructions stored thereon, which when executed determine the identity of a base at a specific position and / or the associated BQ in the original nucleic acid molecule based on the information provided in step (ii).

[0268] 26. A computer program comprising instructions which when executed by a computer are capable of determining the identity of a determined / inferred base at a specific position and / or the associated BQ in the original nucleic acid molecule based on the information provided in step (ii) of the method defined in any one of the preceding claims.

[0269] 27. A kit comprising at least two different primers, such as at least three different primers, preferably four different primers, wherein the at least two different primers, preferably four different primers, are capable of at least partially binding to at least four different regions in the nucleic acid molecule provided in step (i) of the method defined in any one of items 1 to 25, wherein: 1. at least one of the primers is capable of at least partially binding (hybridizing) to at least a portion of the adaptor at the 5' end of the nucleic acid molecule provided in (i) for sequencing at least a portion of the 5' region of the nucleic acid molecule provided in (i); 2. At least one of the primers is capable of binding (hybridizing) at least partially to regions of the nucleotide sequences of the 5' and 3' regions of the nucleic acid molecule provided in (i) to sequence at least a portion of the 3' region of the nucleic acid molecule provided in (i); 3. At least one of the primers is capable of binding (hybridizing) at least partially to at least a portion of the linker at the 3' end of the nucleic acid molecule provided in (i) to sequence at least a portion of the 3' region of the nucleic acid molecule provided in (i); and / or 4. At least one of the primers is capable of binding (hybridizing) at least partially to regions of the nucleotide sequences of the 5' and 3' regions of the nucleic acid molecule provided in (i) to sequence at least a portion of the 5' region of the nucleic acid molecule provided in (i).

[0270] 28. The kit according to item 26, wherein the kit further comprises a double-stranded linker for use in the method as defined in any one of claims 9-24, wherein the linker comprises a first nucleic acid strand and a second nucleic acid strand. In this configuration, the 3' region of the first nucleic acid strand and the 5' region of the second nucleic acid strand form a double-stranded region through sequence complementarity. The ends of the double-stranded region formed by the 3' region of the first nucleic acid strand and the 5' region of the second nucleic acid strand of the connector are compatible with the ends of the double-stranded nucleic acid molecule. The double-stranded region of the connector includes one or more barcode sequences, and Wherein, the 3' region of the second chain of the connector forms a hairpin loop through hybridization between a first segment and a second segment within the 3' region, the first segment being located at the 3' end of the 3' region, and the second segment being located near the 5' region of the second chain, and / or The connector includes at least one barcode sequence in a single-chain region of the connector. The barcode sequence consists of a unique identifier, which allows identification of a specific structure including the identifier and its amplification product. Here, compatibility means that the ends of the double-stranded regions of the linker molecule can be connected to one or both ends of the double-stranded nucleic acid molecule.

[0271] 29. The kit according to any one of items 27, wherein the linker has a limiting site in the 5' region of the first chain of the linker.

[0272] 30. The kit of any one of items 27 to 29, wherein the adaptor comprises at least one barcode sequence in the single-stranded region of the adaptor, and wherein the 3' region of the second strand of the adaptor forms a hairpin loop through hybridization between a first fragment and a second fragment within the 3' region, the first fragment being at the 3' end of the 3' region, the second fragment being proximal to the 5' region of the second strand.

[0273] 31. The kit of any one of items 27 to 30, wherein the kit further comprises: (i) a library of double-stranded adaptors, the adaptors comprising a first strand and a second strand, wherein a 3' region of the first strand and a 5' region of the second strand form a double-stranded region through sequence complementarity, and the ends of the double-stranded region are compatible with the ends of the double-stranded nucleic acid molecule; (ii) a plurality of extension primers, wherein each extension primer comprises a 3' region that is complementary to the second strand of the adaptor molecule defined in (i), and the extension forms an overhanging end upon hybridization to the second strand of the adaptor molecule; (iii) a plurality of hairpin adaptors, wherein each hairpin adaptor comprises a hairpin loop region and an overhanging end that is compatible with the overhanging end formed upon hybridization of the extension primer defined in (ii) and the second strand of the Y adaptor defined in (i), wherein the extension primers of (ii) and the hairpin adaptors of (iii) can be provided as a complex; wherein the adaptors of (i), the extension primers of (ii), and the hairpin adaptors of (iii) are suitable for obtaining a library of adaptors for use in the method defined in any one of claims 9 to 24.

Claims

1. A method comprising: i. providing a nucleic acid molecule comprising a 5' region and a 3' region, wherein the 5' region and the 3' region are covalently linked by a nucleotide sequence that can bind to a primer, wherein the base identity of one of the 5' region or the 3' region independently of the base identity in the other region provides information about the base identity in the corresponding locus in the original nucleic acid molecule, wherein the molecule further comprises: - one linker located at the 5' end of the molecule; - one linker located at the 3' end of the molecule; ii. sequencing the molecule provided in step (i) using at least two different primers, such as at least three different primers, preferably at least four different primers, wherein the at least two different primers, such as at least three different primers, preferably at least four different primers, bind to at least three, preferably at least four, different regions in the nucleic acid molecule provided in (i), wherein:

1. at least one of the primers is capable of at least partially binding to at least a portion of the adaptor at the 5' end of the molecule to sequence at least a portion of the 5' region of the nucleic acid molecule provided in (i); 2. at least one of the primers is capable of at least partially binding to a region of the nucleotide sequence covalently linking the 5' region and the 3' region of the nucleic acid molecule provided in (i) to sequence at least a portion of the 3' region of the nucleic acid molecule provided in (i); 3. at least one of the primers is capable of at least partially binding to at least a portion of the adaptor at the 3' end of the molecule to sequence at least a portion of the 3' region of the nucleic acid molecule provided in (i); and / or 4. at least one of the primers is capable of at least partially binding to a region of the nucleotide sequence covalently linking the 5' region and the 3' region of the nucleic acid molecule provided in a. to sequence at least a portion of the 5' region of the nucleic acid molecule provided in (i).

2. The method according to claim 1, wherein the 3' region of the nucleic acid molecule provided in step (i) is at least partially complementary to the reverse strand of the 5' region.

3. The method according to any of the preceding claims, wherein, The nucleic acid molecule provided in step (i) is a DNA molecule.

4. The method according to any of the preceding claims, wherein, The 5' region and / or the 3' region of the nucleic acid molecule provided in step (i) is a fragment of genomic DNA.

5. The method according to any of the preceding claims, wherein, The 5' region of the nucleic acid molecule provided in step (i) is a fragment of genomic DNA and the 3' region of the nucleic acid molecule provided in step (i) is complementary to the reverse strand of the 5' region.

6. The method of any one of claims 1-4, wherein, The 3' region of the nucleic acid molecule provided in step (i) is a fragment of genomic DNA and the 5' region of the nucleic acid molecule provided in step (i) is complementary to the reverse strand of the 3' region.

7. The method according to any of the preceding claims, wherein, The nucleotide sequence covalently linking the 5' region and the 3' region of the nucleic acid molecule provided in step (i) has a length of at least 5 nucleotides, preferably a length of at least 10 nucleotides, even more preferably a length of at least 17 nucleotides.

8. The method of any of the preceding claims, wherein, In step (ii), the molecule provided in step (i) is sequenced using four different primers.

9. The method according to any of the preceding claims, wherein, The nucleic acid molecule in step (i) is generated by: a. providing a population of double-stranded nucleic acid molecules, preferably wherein the plurality of double-stranded DNA molecules are fragments of genomic DNA; b. at least partially ligating a double-stranded (ds) adaptor to at least one end of the strands of the plurality of double-stranded nucleic acid molecules, thereby obtaining a plurality of adaptor-containing nucleic acid molecules; c. using each strand of the nucleic acid molecules obtained in step (b) as a template, synthesizing a complementary strand, "synthesizing a complementary strand", to each strand of the nucleic acid molecules obtained in step (b) by polymerase extension from the 3' end of the second nucleic acid strand in the adaptor molecule, thereby pairing each strand of the nucleic acid molecules obtained in step (b) with its synthesized complementary strand to provide a plurality of adaptor-modified nucleic acid molecules, wherein the original nucleic acid strand obtained in step (c) and its synthetic complementary strand are covalently linked via a nucleotide sequence to which a primer can at least partially bind; d. optionally, providing a complementary strand of the plurality of adaptor-modified DNA molecules obtained in step (c), optionally using a primer whose sequence is complementary to at least a portion of the double-stranded adaptor; e. optionally amplifying the paired double-stranded nucleic acid molecules obtained in step (d) to provide amplified paired double-stranded nucleic acid molecules.

10. The method of claim 9, wherein, at least a portion of the double-stranded adaptor has a sequence common to all double-stranded adaptors used in step (b).

11. The method according to one or more of claims 9 to 10, wherein, prior to step (c), the plurality of paired adaptor-modified DNA molecules are separated to generate a library of paired adaptor-modified DNA molecules.

12. The method according to one or more of claims 9 to 11, wherein, the 3' region of the second DNA strand of the adaptor forms a hairpin loop via hybridization between a first segment and a second segment within the 3' region, the first segment being located at the 3' end of the 3' region and the second segment being located near the 5' region of the second DNA strand.

13. The method according to one or more of claims 9 to 12, wherein, after step (c), the method further comprises the following step (cl): contacting each strand of the adaptor-containing nucleic acid molecules with a complex of an extension primer and a hairpin adaptor under conditions sufficient for the extension primer to hybridize to the second strand of the adaptor molecule, wherein the extension primer comprises a 3' region complementary to the second strand of the adaptor molecule and the extension primer forms an overhang upon hybridization to the second strand of the adaptor molecule, and wherein the hairpin adaptor comprises a hairpin loop region and an overhang compatible with the overhang formed upon hybridization of the extension primer to the second strand of the adaptor.

14. The method according to one or more of claims 9 to 13, wherein, the adaptor has a first barcode sequence in a double-stranded region and / or a second barcode sequence in a 3' region of the second strand of the adaptor.

15. The method according to one or more of claims 9 to 14, wherein, the adaptor has a restriction site in a 5' region of the first strand of the adaptor.

16. The method according to one or more of claims 9 to 15, wherein, in step (b), a plurality of adaptor-modified nucleic acid molecules are provided, wherein the strands of the genomic DNA fragments are further paired using barcode sequences.

17. The method of claim 16, wherein, The pairing can be performed before or after the ligation, or simultaneously with the ligation.

18. The method according to one or more of claims 9 to 17, wherein, at least a portion of the double-stranded adaptor has a sequence common to all double-stranded adaptors used in step (b).

19. The method according to one or more of claims 9 to 18, wherein, after step (c), the method further comprises the following steps c21 and / or c22: c21) in the paired adaptor-containing nucleic acid molecules, converting unmodified nucleotides, if present, in the paired adaptor-containing nucleic acid molecules to another nucleotide that is visibly different from the nucleotide; and / or c22) in the paired adaptor-containing nucleic acid molecules, converting modified nucleotides, if present, in the paired adaptor-containing nucleic acid molecules to another nucleotide that is visibly different from the nucleotide.

20. The method of any one of claims 1 to 8, wherein the nucleic acid molecules in step (i) are generated by: a. providing a double stranded nucleic acid molecule, preferably wherein the double stranded DNA molecule is a fragment of genomic DNA; b. covalently linking the forward and reverse single stranded nucleic acid molecules provided in step a, wherein the covalent linking in step b is by means of the nucleotide sequence that can bind to a primer to obtain a nucleic acid molecule comprising a 5' region and a 3' region, wherein the 5' region and the 3' region are covalently linked by means of the nucleotide sequence that can bind to a primer, wherein the base identity of one of the 5' region or the 3' region and the base identity in the other region, each independently provide information about the base identity in the corresponding locus in the original nucleic acid molecule.

21. The method of claim 20, wherein, The double stranded nucleic acid molecule provided in step a. is provided as a plurality of double stranded nucleic acid molecules, preferably the plurality of double stranded DNA molecules are fragments of genomic DNA.

22. The method according to one or more of claims 1 to 21, wherein, The method further comprises determining the true identity of a base at a particular locus in the original nucleic acid molecule based on the information provided in step (ii).

23. The method of claim 22, wherein, The true identity of the base at the locus of the original nucleic acid molecule is determined to be false if the identity of the first base from reads 1 and 3 and the identity of the second base from reads 4 and 2 do not match any of the following combinations, respectively: 1) Adenine, Adenine, Thymine and Thymine corresponding to A; 2) Thymine, Thymine, Adenine and Adenine corresponding to T; 3) Thymine, Guanine, Adenine and Guanine corresponding to unmethylated C; 4) Guanine, Adenine, Guanine, Thymine corresponding to G; 5) Cytosine, Cytosine, Guanine and Guanine corresponding to methylated C.

24. The method according to one or more of the preceding claims, wherein, The method further comprises using a computer, the computer comprising a processor, a memory and instructions stored thereon that, when executed, determine the identity of a base and / or associated BQ at a particular position in the original nucleic acid molecule based on the information provided in step (ii).

25. A computer program comprising instructions that, when executed by a computer, are capable of determining the identity of a determined / inferred base and / or associated BQ at a particular position in the original nucleic acid molecule based on the information provided in step (ii) of the method of any one of the preceding claims.

26. A kit comprising at least two different primers, such as at least three different primers, preferably four different primers, wherein the at least two different primers, such as at least three different primers, preferably four different primers, are capable of binding at least partially to at least three, preferably at least four different regions in the nucleic acid molecule provided in step (i) of the method according to any one of claims 1 to 25, wherein:

1. at least one of the primers is capable of at least partially binding (hybridizing) to at least a portion of the adaptor at the 5' end of the nucleic acid molecule provided in (i) to sequence at least a portion of the 5' region of the nucleic acid molecule provided in (i); 2. at least one of the primers is capable of at least partially binding (hybridizing) to a region of the nucleotide sequence covalently linking the 5' region and the 3' region of the nucleic acid molecule provided in (i) to sequence at least a portion of the 3' region of the nucleic acid molecule provided in (i); 3. at least one of the primers is capable of at least partially binding (hybridizing) to at least a portion of the adaptor at the 3' end of the nucleic acid molecule provided in (i) to sequence at least a portion of the 3' region of the nucleic acid molecule provided in (i); and / or 4. at least one of the primers is capable of at least partially binding (hybridizing) to a region of the nucleotide sequence covalently linking the 5' region and the 3' region of the nucleic acid molecule provided in (i) to sequence at least a portion of the 5' region of the nucleic acid molecule provided in (i).

27. The kit of claim 26, wherein, The kit further comprises a double-stranded adaptor for use in a method according to any one of claims 9 to 24, wherein the adaptor comprises a first nucleic acid strand and a second nucleic acid strand, wherein a 3' region of the first nucleic acid strand and a 5' region of the second nucleic acid strand form a double-stranded region by sequence complementarity, wherein the ends of the double-stranded region formed by the 3' region of the first nucleic acid strand and the 5' region of the second nucleic acid strand of the adaptor are compatible with the ends of a double-stranded nucleic acid molecule, wherein the double-stranded region of the adaptor comprises one or more barcode sequences, and wherein a 3' region of the second strand of the adaptor forms a hairpin loop by hybridization between a first segment and a second segment within the 3' region, the first segment being located at the 3' end of the 3' region and the second segment being located proximal to the 5' region of the second strand, and / or wherein the adaptor comprises at least one barcode sequence in a single-stranded region of the adaptor, wherein the barcode sequence consists of a unique identifier that allows the identification of a specific structure comprising the identifier and its amplification products, and wherein compatible means that the ends of the double-stranded region of the adaptor molecule are capable of ligating to one or both ends of a double-stranded nucleic acid molecule.

28. The kit of any one of claims 26 or 27, wherein, The adaptor has a restriction site in the 5' region of the first strand of the adaptor.

29. The kit of any one of claims 26 to 28, wherein, The adaptor comprises at least one barcode sequence in a single-stranded region of the adaptor, and wherein a 3' region of the second strand of the adaptor forms a hairpin loop by hybridization between a first segment and a second segment within the 3' region, the first segment being located at the 3' end of the 3' region and the second segment being located proximal to the 5' region of the second strand.

30. The kit of any one of claims 26 to 29, wherein, The kit further comprises: (i) a library of double-stranded adaptors, the adaptors comprising a first strand and a second strand, wherein a 3' region of the first strand and a 5' region of the second strand form a double-stranded region by sequence complementarity, and wherein the ends of the double-stranded region are compatible with the ends of a double-stranded nucleic acid molecule; (ii) a plurality of extension primers, wherein each extension primer comprises a 3' region that is complementary to the second strand of the adaptor molecule defined in (i) and forms an overhanging end upon hybridization to the second strand of the adaptor molecule; and (iii) a plurality of sequencing primers, wherein each sequencing primer comprises a 3' region that is complementary to the first strand of the adaptor molecule defined in (i) and forms an overhanging end upon hybridization to the first strand of the adaptor molecule. (iii) a plurality of hairpin adaptors, wherein each hairpin adaptor comprises a hairpin loop region and an overhanging end compatible with the overhanging end formed upon hybridization of the extension primer defined in (ii) and the second strand of the Y adaptor defined in (i), wherein the extension primer of (ii) and the hairpin adaptor of (iii) can be provided as a complex; wherein the adaptors of (i), the extension primer of (ii) and the hairpin adaptor of (iii) are suitable for obtaining a library of adaptors for use in the method of any one of claims 9 to 24.

Citation Information

Patent Citations

  • Improved method for bisulfite treatment

    EP1394172A1

  • Method for generating double stranded DNA libraries and sequencing methods for the identification of methylated cytosines

    WO2015104302A1