DNA sequencing methods

The novel sequencing method using linked 5' and 3' regions with multiple primers enhances nucleic acid sequencing accuracy and error detection, addressing errors in conventional NGS and improving molecular diagnosis and cancer treatment.

JP2026514126APending Publication Date: 2026-05-01ANILING SL
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
ANILING SL
Filing Date
2024-04-15
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Conventional nucleic acid sequencing methods, particularly Next-Generation Sequencing (NGS), introduce errors at various stages, leading to inaccuracies in detecting low-frequency gene variants, especially when methylation is considered, which complicates molecular diagnosis and cancer treatment.

Method used

A novel sequencing methodology that utilizes nucleic acid molecules with 5' and 3' regions covalently linked by nucleotide sequences, allowing for sequencing with at least two different primers to provide multiple reads per molecule, enhancing error detection and correction, especially for cytosine methylation analysis.

Benefits of technology

This approach significantly reduces sequencing errors, improves accuracy, and enables faster, more efficient processing with lower costs by providing multiple sources of information for each nucleotide read, allowing for precise base identification and quality determination.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026514126000017
    Figure 2026514126000017
  • Figure 2026514126000018
    Figure 2026514126000018
  • Figure 2026514126000019
    Figure 2026514126000019
Patent Text Reader

Abstract

The present invention relates to a method for determining the sequence of a nucleic acid molecule. In particular, the present invention provides: i. a nucleic acid molecule comprising a 5' region and a 3' region, wherein the 5' region and the 3' region are covalently linked by a nucleotide sequence to which a primer can bind, and the base recognition in one of the 5' region or the 3' region and the base recognition in the other region together independently provide information regarding the base recognition at the corresponding locus in the original nucleic acid molecule, and the molecule further comprises: - one adapter at the 5' end of the molecule; - one adapter at the 3' end of the molecule; ii. sequencing the molecule provided in step (i) using at least two different primers, e.g., at least three different primers, preferably at least four different primers, wherein at least two different primers, e.g., at least three different primers, preferably at least four different primers, bind to at least three different regions, preferably at least four different regions in the nucleic acid molecule provided in (i): 1. At least one of the primers binds at least partially to at least a portion of the adapter at the 5' end of the molecule, thereby sequencing at least a portion of the 5' region of the nucleic acid molecule provided in (i); 2. The present invention provides a method comprising: 3. At least one of the primers at least partially binds to a region of a nucleotide sequence covalently linking the 5' and 3' regions of the nucleic acid molecule provided in (i), thereby enabling sequencing of at least a portion of the 3' region of the nucleic acid molecule provided in (i); 4. At least one of the primers at least partially binds to at least a portion of the adapter at the 3' end of the molecule, thereby enabling sequencing of at least a portion of the 3' region of the nucleic acid molecule provided in (i); and / or 5. At least one of the primers at least partially binds to a region of a nucleotide sequence covalently linking the 5' and 3' regions of the nucleic acid molecule provided in (i), thereby enabling sequencing of at least a portion of the 5' region of the nucleic acid molecule provided in (i).
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a method for determining the sequence of a nucleic acid molecule, including the identification of methylated cytosine (e.g., cytosine nucleic acid methylation (5mCs)) in any sequence context (e.g., CpG, CHG, CHH (H=A, C, or T)). The present invention also relates to computer programs and kits related to the method of the present invention. [Background technology]

[0002] The analysis of the primary structure of nucleic acids (as DNA and RNA), including epigenetic modifications (e.g., methylation of DNA or RNA), can be addressed by using different techniques commonly referred to as "sequencing."

[0003] Single-end (SE) sequencing sequences nucleic acids from only one end of an insert; this is the first and simplest method for high-throughput sequencing. A simple modification to the standard SE library preparation process made it easy to read both the forward and reverse template strands of each cluster during a single paired-end read.

[0004] In library preparation, in addition to producing twice the number of reads from the same amount of nucleic acid (input) as single-ended reads in the same time and effort, sequences aligned as read pairs enable more accurate read alignment and allow for the detection of inversions and insertion-deletion (indel) variants, which is not possible with single-ended read data (Nakazato T, Ohta T, Bono H. Experimental design-based functional mining and characterization of high-throughput sequencing data in the sequence read archive. PLoS One. 2013;8(10):e77910).

[0005] Paired-end (PE) sequencing allows sequencing of both ends of a nucleic acid template. Because both reads contain long-range positional information, highly accurate read alignment is possible. Once the sample is prepared, the distance between each paired read is known, allowing the alignment algorithm to use this information to more accurately map reads on repeating regions. This results in significantly better read alignment, especially in repeating regions of the genome where sequencing is difficult.

[0006] If one PE read maps to a different region of the genome while the other does not, the first read should likely be relatively close to the other read (following the known distribution of insert sizes in library preparation). This is also true when both reads are multi-mapped, and even if the marginal alignment probability for each read is the same, a particular pairwise alignment may be more likely than other pairwise alignments.

[0007] PE not only provides a double amount of DNA sequenced to search for structural variant breakpoints, but can also estimate whether there is a breakpoint between two paired reads when the two reads are aligned much further apart than expected, or when they are aligned on separate chromosomes. For SE reads, structural variant breakpoints are detected only if the read overlaps with the breakpoint. This concept also applies to RNA containing chromosomal conformational capture and ligation junction sites, or splice sites.

[0008] In summary, the advantages of paired ends over single ends are as follows: • High accuracy: Enables the detection of errors and gaps in sequencing data. Errors in one read can be corrected using the other read. • Longer read length: This allows us to use the combined length of both reads to infer the arrangement of the original fragments. • Better genome assembly: Provides information on the orientation and distance between fragments, aiding in the identification of insertions and deletions within the genome. • Improved detection of structural variations: Information regarding the location and size of these variants (insertions, deletions, and inversions) can be obtained.

[0009] Next-generation sequencing (NGS) technology is highly accurate, but not perfect, and introduces errors at various steps throughout the workflow. See, for example, Figure 1 in Ma, X., Shao, Y., Tian, ​​L. et al. Analysis of error profiles in deep next-generation sequencing data. Genome Biol 20, 50 (2019): • Sample handling: during sample collection, storage, and nucleic acid isolation; Library preparation: due to fragmentation, ligation, adapter contamination, index hopping, amplification PCR error and bias, and capture bias; • Sequencing itself: cluster amplification, inaccuracies in detecting emitted fluorescence signals, or base determination errors due to incorrectly incorporated reversible terminators); • Bioinformatics data analysis: misalignment during read mapping, identification of duplicates, and incorrect variant determination and INDEL assembly errors.

[0010] The raw NGS error rate averages 10 -3 However, this varies depending on the input sample type, technique, equipment, sequence context, base position along the read, and nucleotide substitution type.

[0011] For example, C>T / G>A and A>G / T>C substitutions account for approximately 70% of errors in terms of the total number of mutated nucleotides. See, for example, Figure 1 in Zhang Z, Gerstein M. Patterns of nucleotide substitution, insertion and deletion in the human genome inferred from pseudogenes. Nucleic Acids Res. 2003 Sep 15;31(18):5338-48.

[0012] When methylation is considered, the use of deamination processes, the generation of C>T / G>A ambiguity (WGBS), or inappropriate or failed conversion rates increase the potential for errors.

[0013] These errors are significant confounding factors for detecting low-frequency gene variants that are important for the molecular diagnosis, treatment, and surveillance of cancer using NGS.

[0014] A new type of sequencing technique called Genomic and Epigenomic Unified Sequencing (GEUS), described, for example, in International Publication No. 2015 / 104302, can take methylation into account and significantly reduce the error rate by examining the same position of the original sequence in different contexts from two related strands.

[0015] In the GEUS technique (as described, for example, in International Publication No. 2015 / 104302), the five true bases of the original molecule, e.g., A, T, G, unmethylated C(C), and methylated C(M), are inferred from the following two-letter matching codes:

[0016] [Table 1]

[0017] Any other combination that does not match the two-letter code can only be generated by an error at some point in the entire process, and thus can be detected and removed (e.g., N is assigned to ignore that particular position), or a certain base or two-base ambiguity can be tentatively assigned according to a decision tree based on the base quality (BQ) of each nucleotide, and if the read is mapped to the reference and the reference nucleotide is known, it can be finally determined (the recalibration step of the inferred base).

[0018] [Table 2]

[0019] In this high-precision method, most types of nucleotide substitutions and especially short INDELs are detected and removed, and the error rate is significantly reduced.

[0020] For example, when T is misjudged as G, two errors are required to generate another matching two-letter code at a specific position in both reads that exactly corresponds to the same base in the original template. Therefore, this type of error is significantly reduced (about 10,000 times) under optimal conditions. To cause a misjudgment in one inferred read, an error from T to C (T>C) in read 1 and an error from T to A (T>A) in read 2 are required.

[0021] [Table 3]

[0022] If only one of the two errors occurs, that error will result in a non-matching two-letter code and will be detected and removed.

[0023] [Table 4] TIFF2026514126000005.tif14160

[0024] Another example:

[0025] [Table 5]

[0026] If two errors occur but do not lead to the generation of a matching two-character code, those errors will also be detected and removed. For example:

[0027] [Table 6]

[0028] However, because there is ambiguity in the two-letter matching codes, some of the most common types of substitution errors in DNA, such as G>A / C>T and T>C / A>G, need to be further reduced.

[0029] Furthermore, in the GEUS technology described in, for example, International Publication No. 2015 / 104302, the molecule being sequenced (e.g., referred to as the "GEUS molecule") has two regions that independently provide information regarding base identification at the corresponding locus in the original nucleic acid molecule. See, for example, the molecule described in claim 1. However, when this molecule is subjected to conventional PE sequencing, single-end sequencing is actually obtained. This is because there are two regions within the same molecule that contain related information. As a result, only one end of each region of the GEUS molecule is read, and the information from the other end that would be provided in conventional paired-end sequencing is missed. See, for example, Figure 7A. [Overview of the Initiative] [Problems that the invention aims to solve]

[0030] Therefore, further sequencing methodologies are needed to reduce the error rate when sequencing nucleic acid molecules, such as genomic DNA or RNA molecules, particularly the nucleic acid molecules described in step i of the present invention. [Means for solving the problem]

[0031] The present invention addresses the above requirements and provides a novel sequencing methodology that enables a very low error rate. In particular, by using the method described herein, it is possible to obtain more than two, for example three, or preferably four, reads per molecule that partially or completely cover the insert to be sequenced. The methodology described herein makes it possible to confirm the start and / or end of all inserts to be sequenced, regardless of the size or structural rearrangement of the insert or the insertion or deletion (INDEL) of bases contained in the insert. The information provided by the methodology of the present invention has higher quality and / or a lower error rate compared to prior art sequencing methodologies. The improved quality also leads to improved process efficiency, resulting in a faster process associated with lower costs. [Effects of the Invention]

[0032] Currently, NGS workflows introduce errors that must be anticipated. As mentioned above, errors are introduced at various steps in conventional NGS workflows and result from sample handling, library preparation, enrichment PCR, sequencing, mapping, duplication, and variant determination. Most of these errors are not depicted in the final results and remain hidden. This has a negative effect on subsequent processing and analysis of the results. The method of the present invention provides more than two, for example three, preferably up to four (up to eight if considering the original Watson-Crick double-stranded molecule; see below for further details) sources of information for each nucleic acid molecule, and enables verification of each nucleotide read, as all reads must be consistent. Thus, the method of the present invention makes it possible to detect and further correct errors in sequencing (both primary sequencing and analysis of modified nucleotides, e.g., cytosine methylation analysis).

[0033] The method of the present invention is suitable for sequencing nucleic acid molecules including 5' and 3' regions. The 5' and 3' regions are covalently linked by nucleotide sequences to which primers can bind. Base recognition in either the 5' or 3' region, and base recognition in the other region, independently provide information about base recognition at the corresponding locus in the original nucleic acid molecule, and the molecule is as follows: - A single adapter at the 5' end of the molecule; - A single adapter at the 3' end of the molecule It also includes.

[0034] Such molecules are referred to herein as "the molecule defined in step i of the present invention," "the molecule according to the present invention," and so on. A particular example of this is the so-called "GEUS molecule" described in International Publication No. 2015 / 104302.

[0035] In particular, the present invention is as follows: i. To provide a nucleic acid molecule containing a 5' region and a 3' region, The 5' and 3' regions are covalently linked by known nucleotide sequences to which primers can bind. Base recognition in either the 5' or 3' region, and base recognition in the other region, independently provide information about base recognition at the corresponding locus in the original nucleic acid molecule. The molecules are as follows: - A single adapter at the 5' end of the molecule; - A single adapter at the 3' end of the molecule Further includes; ii. Sequencing the molecule provided in step (i) using at least two different primers, preferably at least three different primers, more preferably four different primers, wherein at least two different primers bind to at least three different regions, preferably at least four different regions, in the nucleic acid molecule provided in (i): 1. At least one primer can at least partially bind to at least a portion of the adapter at the 5' end of the molecule to sequence at least a portion of the 5' region of the nucleic acid molecule provided in (i); 2. At least one of the primers can at least partially bind to a region of a nucleotide sequence that covalently links the 5' and 3' regions of the nucleic acid molecule provided in (i), thereby sequencing at least a portion of the 3' region of the nucleic acid molecule provided in (i); 3. At least one primer can at least partially bind to at least a portion of the adapter at the 3' end of the molecule to sequence at least a portion of the 3' region of the nucleic acid molecule provided in (i); and / or 4. A method is provided in which at least one primer can be at least partially bound to a region of a nucleotide sequence that covalently links the 5' and 3' regions of the nucleic acid molecule provided in a., thereby sequencing at least a portion of the 5' region of the nucleic acid molecule provided in (i).

[0036] In one embodiment, step ii. includes sequencing the molecule provided in step (i) using at least two different primers, preferably at least three different primers, wherein at least two different primers bind to at least three different regions in the nucleic acid molecule provided in (i): 1. At least one of the primers (e.g., the first primer) can at least partially bind (hybridize) to at least a portion of the adapter at the 5' end of the nucleic acid molecule provided in (i) so that at least a portion of the 5' region of the nucleic acid molecule provided in (i) can be sequenced; 2. At least one of the primers (e.g., the second primer) can at least partially bind to at least a portion of the adapter at the 3' end of the molecule so that at least a portion of the 3' region of the nucleic acid molecule provided in (i) can be sequenced; and 3. At least one of the primers (e.g., the third primer) is at least partially: 3.1. A region of nucleotide sequence covalently linking the 5' and 3' regions of the nucleic acid molecule provided in (i) in order to sequence at least a portion of the 3' region of the nucleic acid molecule provided in (i); or 3.2 (i) To sequence at least a portion of the 5' region of the nucleic acid molecule provided in a., a region of nucleotide sequence covalently linking the 5' and 3' regions of the nucleic acid molecule provided in a. It can be combined (hybridized) with any of the following.

[0037] In a preferred embodiment, the present invention is as follows: i. To provide a nucleic acid molecule containing a 5' region and a 3' region, The 5' and 3' regions are covalently linked by known nucleotide sequences to which primers can bind. Base recognition in either the 5' or 3' region, and base recognition in the other region, independently provide information about base recognition at the corresponding locus in the original nucleic acid molecule. The molecules are as follows: - A single adapter at the 5' end of the molecule; - A single adapter at the 3' end of the molecule Further includes; ii. Sequencing the molecule provided in step (i) using at least two different primers, for example, at least three different primers, preferably four different primers, wherein at least two different primers, preferably four different primers, bind to at least four different regions, preferably four different regions, in the nucleic acid molecule provided in (i): 1. At least one of the primers can at least partially bind (hybridize) to at least a portion of the adapter at the 5' end of the nucleic acid molecule provided in (i) so that at least a portion of the 5' region of the nucleic acid molecule provided in (i) can be sequenced; 2. At least one of the primers can at least partially bind (hybridize) to a region of a nucleotide sequence that covalently links the 5' and 3' regions of the nucleic acid molecule provided in (i), thereby sequencing at least a portion of the 3' region of the nucleic acid molecule provided in (i); 3. At least one of the primers can at least partially bind (hybridize) to at least a portion of the adapter at the 3' end of the nucleic acid molecule provided in (i) so that at least a portion of the 3' region of the nucleic acid molecule provided in (i) can be sequenced; 4. A method is provided in which at least one of the primers can at least partially bind (hybridize) to a region of a nucleotide sequence that covalently links the 5' and 3' regions of the nucleic acid molecule provided in (i), thereby enabling sequencing of at least a portion of the 5' region of the nucleic acid molecule provided in (i).

[0038] Therefore, in any embodiment of the present invention, it is possible to determine the true base identification (e.g., A, G, C, T, U, M, or any modification thereof) and / or base quality (BQ, i.e., the probability that the assigned identification is true or false) at a specific position in the original nucleic acid molecule. Accordingly, in any embodiment of the present invention, preferably, a further step is to determine the true base identification (e.g., A, G, C, T, U, M, or any modification thereof) and / or base quality (BQ, i.e., the probability that the assigned identification is true or false) at a specific position in the nucleic acid molecule provided in step i, based on the information provided in step (ii).

[0039] The present invention also provides a computer program comprising instructions, the instructions which, when executed by a computer, can determine the true base identification (e.g., A, G, C, T, U, M, or any modification thereof) and / or base quality (BQ, i.e., the probability that the assigned identification is true or false) at a particular position in the original nucleic acid molecule, based on information provided in step (ii) of the method of the present invention.

[0040] Furthermore, a kit is provided comprising at least two different primers, preferably at least three different primers, more preferably four different primers, the at least two different primers, preferably three different primers, more preferably four different primers, which can at least partially bind to at least three different regions, preferably at least four different regions, in the nucleic acid molecule provided in step (i) of the method of the present invention: 1. At least one primer can at least partially bind to at least a portion of the adapter at the 5' end of the molecule to sequence at least a portion of the 5' region of the nucleic acid molecule provided in (i); 2. At least one of the primers can at least partially bind to a region of a nucleotide sequence that covalently links the 5' and 3' regions of the nucleic acid molecule provided in (i), thereby sequencing at least a portion of the 3' region of the nucleic acid molecule provided in (i); 3. At least one primer can at least partially bind to at least a portion of the adapter at the 3' end of the molecule to sequence at least a portion of the 3' region of the nucleic acid molecule provided in (i); and / or 4. At least one primer can at least partially bind to a region of a nucleotide sequence that covalently links the 5' and 3' regions of the nucleic acid molecule provided in a., thereby sequencing at least a portion of the 5' region of the nucleic acid molecule provided in (i).

[0041] A kit is further provided comprising at least two different primers, preferably at least three different primers, which can at least partially bind to at least three different regions in the nucleic acid molecule provided in step (i) of the method of the present invention: 1. At least one of the primers (e.g., the first primer) can at least partially bind (hybridize) to at least a portion of the adapter at the 5' end of the nucleic acid molecule provided in (i) so that at least a portion of the 5' region of the nucleic acid molecule provided in (i) can be sequenced; 2. At least one of the primers (e.g., the second primer) can at least partially bind to at least a portion of the adapter at the 3' end of the molecule so that at least a portion of the 3' region of the nucleic acid molecule provided in (i) can be sequenced; and 3. At least one of the primers (e.g., the third primer) is at least partially: 3.1. A region of nucleotide sequence covalently linking the 5' and 3' regions of the nucleic acid molecule provided in (i) in order to sequence at least a portion of the 3' region of the nucleic acid molecule provided in (i); or 3.2 (i) To sequence at least a portion of the 5' region of the nucleic acid molecule provided in a., a region of nucleotide sequence covalently linking the 5' and 3' regions of the nucleic acid molecule provided in a. It can be combined (hybridized) with any of the following.

[0042] In a preferred embodiment, the present invention provides a kit comprising at least two different primers, preferably four different primers, which can at least partially bind to at least four different regions in a nucleic acid molecule provided in step (i) of the method of the present invention: 1. At least one of the primers can at least partially bind (hybridize) to at least a portion of the adapter at the 5' end of the nucleic acid molecule provided in (i) so that at least a portion of the 5' region of the nucleic acid molecule provided in (i) can be sequenced; 2. At least one of the primers can at least partially bind (hybridize) to a region of a nucleotide sequence that covalently links the 5' and 3' regions of the nucleic acid molecule provided in (i), thereby sequencing at least a portion of the 3' region of the nucleic acid molecule provided in (i); 3. At least one of the primers can at least partially bind (hybridize) to at least a portion of the adapter at the 3' end of the nucleic acid molecule provided in (i) so that at least a portion of the 3' region of the nucleic acid molecule provided in (i) can be sequenced; 4. At least one of the primers can at least partially bind (hybridize) to a region of a nucleotide sequence that covalently links the 5' and 3' regions of the nucleic acid molecule provided in (i), thereby sequencing at least a portion of the 5' region of the nucleic acid molecule provided in (i).

[0043] Preferably, any of the above primers at least partially hybridize to a specific location (region) of the nucleic acid molecule provided in step (i) of the method of the present invention at the time of sequencing by synthesis. [Brief explanation of the drawing]

[0044] [Figure 1] Schematic representation of a nucleic acid molecule as defined in step i of the present invention, for example, the GEUS molecule as a template (e.g., described in International Publication No. 2015 / 104302). The 5' region is represented from left to right as the nucleotides ATTGAACGCT within a gradient-colored (grayscale) line. The darker side of the gradient represents the beginning (5') of the original molecule, and the lighter side represents its end (3'). The 3' region is represented from left to right as the nucleotides AGTGTTTGAT. As before, the darker side of the gradient represents the beginning of the original molecule, and the lighter side represents the end. Adapters at the 5' and 3' ends of the molecule are represented by solid gray lines. The nucleotide sequence covalently linking the 5' and 3' regions is represented in black, connecting the 5' and 3' regions of the molecule. Optionally selected unique molecular barcodes (UMIs) are represented by white lines. At least two different primers are represented by thin black arrows indicating the direction of sequencing synthesis. The thick white arrows containing the sequences correspond to the leads synthesized after primer hybridization in the 5' to 3' direction. The leads can be of the same length or different lengths. R1 represents lead 1, R2 represents lead 2, R3 represents lead 3, and R4 represents lead 4. [Figure 2] Figure 1 shows an analysis of four reads of a nucleic acid molecule when all reads are completely overlapping. The inferred read (labeled "GEUS inferred" in the figure) is based on information from four sequenced bases. Reads can have the same or different lengths (number of cycles in the NGS instrument). The DNA template corresponds to the Watson strand of the original dsDNA sequence. See Figure 1. [Figure 3]Analysis of the four reads of the nucleic acid molecule shown in Figure 1, where the reads do not overlap. In this case, the inferred reads are based on the information of two sequenced bases from each end. The DNA template corresponds to the Watson strand of the original dsDNA sequence. [Figure 4] Analysis of four reads of the nucleic acid molecule shown in Figure 1, when both strands (Watson and Crick) of the original double-stranded DNA sequence are sequenced. In this case, the inferred sequence (reference genome format) is based on information from up to eight sequenced bases derived from a single double-stranded insertion processed through two GEUS molecules paired by barcodes (e.g., described in International Publication No. 2015 / 104302). Note that because the strands of the dsDNA molecule are antiparallel and complementary, the last base at the inferred position on the Crick strand provides information about the first base on the reference inferred in the GEUS molecule, shown in gray gradient. [Figure 4-2] This is a continuation of the above. [Figure 5] A schematic diagram of a method for obtaining a nucleic acid sequence provided in step (i) of the present invention is shown, for example, as described in International Publication No. 2015 / 104302. As before, the darker side of the gradient represents the beginning of the original molecule, and the lighter side represents its end. [Figure 6]This is a schematic diagram illustrating sequencing by steps (i)(A) ​​and (ii)(B) of the method of the present invention in an NGS instrument (e.g., an Illumina MiSeq NGS instrument). (A) is a tiled representation of the flow cell on which cluster amplification is performed, with a nucleic acid molecule according to the present invention, in this case a DNA GEUS molecule (e.g., described in International Publication No. 2015 / 104302) attached. From bottom to top, the end of the molecule contains an NGS adapter, in this case an Illumina P7, followed by the first sample index (for multiplexing samples in the same lane), then the external unique molecular index (UMI), then the 5' region of the nucleic acid molecule, followed by the internal UMI, then the known sequence (i.e., a nucleotide sequence to which primers can bind, which may be a hairpin in the case of a GEUS molecule), the synthetic internal UMI, followed by the 3' region of the nucleic acid molecule (in this case the synthetic 3' region of the molecule), the synthetic external UMI, and finally an Illumina P5 NGS sequencing adapter having a second sample index to prevent index hopping. [Figure 6-2] In section B), the sequence of the first three leads by synthesis is described step by step after each step of primer hybridization (hybridization of primer 1 and sequencing of lead 1, hybridization of primer 2 and sequencing of lead 2, hybridization of sample index primer and first index lead). Next, complementary molecules are synthesized and amplified by clustering, and the last three other leads are generated in the same step by step through primer hybridization and synthesis of lead 3, primer hybridization and synthesis of lead 4, and second index sample primer hybridization and synthesis of second sample index. Depending on the instrument, the leads may have a different synthesis order, and the NGS instrument protocol must be adapted accordingly. [Figure 7]A) Paired-end (PE) sequencing of the molecule provided in step i of the method of the present invention (in practice, this is single-end (SE) sequencing). B) Paired-end (PE) sequencing of the molecule provided in step i of the method of the present invention. [Modes for carrying out the invention]

[0045] The present invention relates to a method for determining (e.g., identifying) the sequence of a nucleic acid molecule, and includes determining (e.g., identifying) variations of a nucleic acid molecule (e.g., fragments of genomic DNA) obtained from a subject, wherein the nucleic acid molecule comprises a 5' region and a 3' region, the 5' region and the 3' region being covalently linked by a nucleotide sequence to which a primer can bind, and the base identification in one of the 5' or 3' regions and the base identification in the other region independently provide information regarding the base identification at the corresponding locus in the original nucleic acid molecule, wherein the molecule is as follows: - One adapter at the 5' end of the molecule; and - A single adapter at the 3' end of the molecule It also includes.

[0046] Nucleic acid molecule sequencing may involve determining the identification of bases present at specific gene loci within the nucleic acid molecule (e.g., adenine (A), cytosine (C), thymine (T), guanine (G), uracil (U), and their modifications, such as methylcytosine (5mC, 5hmC)).

[0047] The method of the present invention, in any of the embodiments described, ensures sequence fidelity, increases sequencing quality, and in particular enables a very low error rate. By using the method described herein, a total of four reads (partially or completely covering the insert to be sequenced) can be obtained per molecule. Therefore, the information provided by the methodology of the present invention is of higher quality and has a lower error rate than conventional synthetic methodologies.

[0048] Furthermore, more accurate sequencing allows for shallower coverage depths and less material to be required to obtain reliable readings.

[0049] Method of the present invention In a first aspect, the present invention provides a method comprising two steps (i) and (ii), such as a nucleic acid sequencing method.

[0050] Step (i) This includes the 5' region and the 3' region. nucleic acid molecule This includes providing ("the nucleic acid molecule of the present invention," "the molecule defined in step i. of the present invention," "the molecule according to the present invention," or the same hereinafter). In a preferred embodiment, the nucleic acid molecule is a DNA molecule, but it may be any nucleic acid molecule, such as RNA. The nucleic acid molecule may be provided as a plurality of nucleic acid molecules, as detailed below.

[0051] The nucleic acid molecule provided in step (i) of the present invention includes a 5' region and a 3' region. In a preferred embodiment, the 5' region and / or 3' region of the nucleic acid molecule are obtained from an organism, e.g., a human or non-human animal, or a plant, bacterium, fungus, yeast, and / or virus, i.e., the 5' region and / or 3' region of the nucleic acid molecule are preferably fragments of genomic DNA (e.g., nuclear DNA, mitochondrial DNA, and chloroplast DNA). The 5' region and / or 3' region of the nucleic acid molecule of the present invention may also be synthetic nucleic acid, e.g., synthetic DNA. Therefore, in a preferred embodiment, the 5' region and / or 3' region of the nucleic acid molecule provided in step (i) is a fragment of genomic DNA. For example, the 5' region and 3' region of the nucleic acid molecule provided in step (i) may be a fragment of genomic DNA. For example, the 5' region and 3' region of the nucleic acid molecule provided in step (i) may be a fragment of synthetic DNA. For example, the 5' region of the nucleic acid molecule provided in step (i) may be a fragment of genomic DNA, and the 3' region of the nucleic acid molecule provided in step (i) may be a fragment of synthetic DNA. For example, the 3' region of the nucleic acid molecule provided in step (i) may be a fragment of genomic DNA, and the 5' region of the nucleic acid molecule provided in step (i) may be a fragment of synthetic DNA. The term "genomic DNA" refers to the heritable genetic information of an organism. Genomic DNA includes not only nuclear DNA (also called chromosomal DNA, including cell-free DNA (cfDNA)) but also the DNA of plastids (e.g., chloroplasts) and other organelles (e.g., mitochondria). The term "genomic DNA" as intended in this invention includes genomic DNA containing sequences complementary to the sequences described in this document.

[0052] The 5' and / or 3' regions of a nucleic acid molecule may be a fragment of plasmid DNA or a single-stranded nucleic acid molecule (e.g., DNA, cDNA, mRNA).

[0053] DNA can be fragmented by any suitable method, including, but not limited to, mechanical stress (e.g., ultrasound, nebulization, cavitation), enzymatic fragmentation (enzymatic digestion by restriction endonucleases, nickel endonucleases, exonucleases, etc.), and chemical fragmentation (e.g., dimethyl sulfate, hydrazine, NaCl, piperidine, acid), or it can be fragmented in the original organism (e.g., cell-free DNA). In principle, there is no limit to the length of DNA fragments, but a narrow range of lengths is preferred. Fragments of appropriate size can be selected before step (i) of the method of the present invention. The optimal length ultimately depends on the available sequencing equipment and methods, as well as the desired percentage of read overlap. In a more preferred embodiment, the DNA molecule is a fragment of genomic DNA.

[0054] The nucleic acid molecule provided in step (i) of the method of the present invention may be a single-stranded (ss) molecule (e.g., an ssDNA molecule). The nucleic acid molecule provided in step (i) may also be provided as a group or part of a group of nucleic acid molecules, for example, as a group or part of DNA molecules described herein.

[0055] As described above, the nucleic acid molecule (or a plurality thereof) provided in step (i) of the method of the present invention includes a 5' region and a 3' region. The 5' region and the 3' region of the nucleic acid molecule are covalently linked by a nucleotide sequence to which a primer can bind. Thus, the nucleic acid molecule of the present invention includes a 5' region and a 3' region separated by a third region which is a nucleic acid sequence located between the 5' region and the 3' region of the nucleic acid molecule of the present invention. Thus, the nucleic acid sequence located between the 5' region and the 3' region covalently links the 5' region and the 3' region of the nucleic acid molecule of the present invention. Thus, the nucleic acid molecule of the present invention includes at least three regions: a 5' region, a linking region, and a 3' region.

[0056] The nucleotide sequence located between the 5' and 3' regions of the nucleic acid molecule of the present invention ("linking region") is a region of the molecule to which the primer can (at least partially) bind (hybridize). Therefore, the nucleotide sequence located between the 5' and 3' regions should be long enough for the primer to (at least partially) bind (hybridize) to it in order to sequence the 5' and / or 3' regions of the molecule of the present invention, and preferably have sufficient specificity so that the primer does not substantially bind to other regions of the molecule. Those skilled in the art will know the means for designing the linking region described herein.

[0057] For example, a nucleotide sequence located between the 5' and 3' regions of the nucleic acid molecule of the present invention, covalently linking the 5' and 3' regions (see, for example, the black region in the molecule in Figure 1), may have a length of at least 5 nucleotides, for example, at least 10 nucleotides, or at least 15 nucleotides, or at least 17 nucleotides, for example, 17 nucleotides. For example, a nucleotide sequence covalently linking the 5' and 3' regions may have a length of 5 to 100 nucleotides, for example, 15 to 100 nucleotides, for example, 15 to 80 nucleotides, for example, 15 to 70 nucleotides, preferably 15 to 80 nucleotides, more preferably 17 to 70 nucleotides, and even more preferably 25 to 65 nucleotides, for example, 17 nucleotides, or 29 nucleotides, or 64 nucleotides. For example, a nucleotide sequence covalently linking the 5' and 3' regions may have a length of at least 20 nucleotides, for example, at least 25, 26, 27, 28, 29, or 30 nucleotides. In a preferred embodiment, the nucleotide sequence has a length of at least 17 nucleotides, for example, 17, 18, or 19 nucleotides. In another preferred embodiment, the nucleotide sequence has a length of 29 nucleotides. It may also have a longer length, for example, at least 35, 40, 45, 50, 55, or at least 60 nucleotides. In another preferred embodiment, it has a length of 64 nucleotides, but it may be a longer length, for example, at least 65, 70, 75, 80, or more nucleotides. Thus, the nucleotide sequence covalently linking the 5' and 3' regions of the nucleic acid of the present invention may comprise 5 to 100 nucleotides, preferably 15 to 80 nucleotides, more preferably 25 to 70 nucleotides, and even more preferably 29 to 64 nucleotides. Of course, other lengths are possible, but only if the primer can (at least partially) bind (hybridize) to the 5' and / or 3' regions of the nucleic acid molecule of the present invention, and preferably has sufficient specificity so that the primer does not substantially bind to other regions of the molecule.

[0058] In the context of the present invention, “hybridization” (or “hybridize”) refers to the process by which two single-stranded polynucleotides are joined (at least partially) by non-covalent bonds to form a stable double-stranded polynucleotide. In the context of the present invention, the term “joining” may be used to mean “hybridize” or “at least partially hybridize.”

[0059] Those skilled in the art are familiar with suitable conditions and buffers for the hybridization of two single-stranded polynucleotides, as described above. For example, “hybridization conditions” may include a salt concentration of about 1 M or less, more typically less than about 500 mM, and may be less than about 200 mM. The “hybridization buffer” is a buffer salt solution, e.g., 5% SSPE, or other such buffers known in the art. The hybridization temperature may be as low as 5°C, but is typically above 22°C, more typically above about 30°C, and typically above 37°C. Hybridization is often carried out under stringent conditions, i.e., under conditions in which the primer hybridizes to its target subsequence but not to other (non-complementary sequences). Exemplary stringent conditions include a pH of about 7.0 to about 8.3 and a temperature of at least 25°C, with a sodium ion concentration (or other salt) of at least 0.01 M to 1 M.

[0060] With respect to composition, the nucleotide sequence covalently linking the 5' and 3' regions of the nucleic acid of the present invention (also referred to in the context of the present invention as the "linking region") may consist of any base that may be present in the nucleic acid molecule (e.g., A, C, T, G, U, and any modification thereof, e.g., methylated C (e.g., 5mC)), but only if a primer can (at least partially) bind (hybridize) to it for sequencing the 5' and / or 3' regions of the nucleic acid molecule of the present invention. In one embodiment, the nucleotide sequence covalently linking the 5' and 3' regions of the nucleic acid of the present invention includes at least one easily modifiable nucleotide, e.g., a modified nucleotide, preferably a modified cytosine, more preferably a methylated cytosine (e.g., 5-methylcytosine (5mC), 5-hydroxymethylcytosine (5hmC), or 5-formylcytosine (5fC)).

[0061] In the nucleic acid molecule of the present invention, base recognition in either the 5' region or the 3' region and base recognition in the other region independently provide information regarding base recognition at the corresponding gene locus in the original nucleic acid molecule.

[0062] The "original nucleic acid molecule" may also be composed of the 5' and / or 3' regions of the nucleic acid molecule of the present invention.

[0063] Therefore, both the 5' and 3' regions of the nucleic acid molecule of the present invention are related in the sense that they both contain information regarding base recognition at the corresponding locus in the original nucleic acid molecule. The information provided by the sequence of the 5' region is independent of the information provided by the sequence of the 3' region. Therefore, the molecule of the present invention relates to base recognition at the corresponding position (locus) of the nucleic acid molecule. two Includes sources.

[0064] In the context of this invention, a locus (plural: loci) is a physical site or location within a nucleic acid molecule.

[0065] For example, in the molecule of the present invention, the 5' region provides information regarding base identification at the corresponding locus of the nucleic acid sequence. Furthermore, the 3' region independently provides information regarding base identification at the same locus of the same nucleic acid sequence.

[0066] Therefore, in the nucleic acid molecule of the present invention, base identification in either the 5' or 3' region provides information regarding the base identification of the corresponding locus of the original nucleic acid molecule, and base identification in the other region (3' or 5', respectively) provides information regarding the base identification of the same locus of the same original nucleic acid molecule, thereby providing concentrated information in the 5' and 3' regions of the molecule of the present invention for each locus of the original nucleic acid molecule.

[0067] In a preferred embodiment, the 3' region of the nucleic acid molecule provided in step (i) is at least partially complementary to the reverse strand of the 5' region. In a further preferred embodiment, the 5' region of the nucleic acid molecule provided in step (i) is a fragment of genomic DNA, and the 3' region of the nucleic acid molecule provided in step (i) is a synthetic complementary strand of the reverse strand of the 5' region. In another embodiment, the 3' region of the nucleic acid molecule provided in step (i) is a fragment of genomic DNA, and the 5' region of the nucleic acid molecule provided in step (i) is a synthetic complementary strand of the reverse strand of the 3' region. In another embodiment, the 5' region of the nucleic acid molecule provided in step (i) is a fragment of genomic DNA, and the 3' region of the nucleic acid molecule provided in step (i) is also a fragment of genomic DNA and is complementary or inversely complementary to the 5' region. In another embodiment, the 3' region of the nucleic acid molecule provided in step (i) is a fragment of genomic DNA, and the 5' region of the nucleic acid molecule provided in step (i) is also a fragment of genomic DNA and is complementary or inversely complementary to the 3' region. Therefore, in these last two examples, both the 3' and 5' regions are parts of the genomic DNA fragment, and they are complementary to each other.

[0068] For example, the sequence of the 5' region of an exemplary nucleic acid molecule of the present invention may correspond to the sequence of the original nucleic acid molecule. Therefore, the 5' region provides information regarding base recognition of the original nucleic acid molecule. Conversely, the sequence of the 3' region of the nucleic acid molecule of the present invention may correspond to the sequence of the inverse complementary strand of the 5' region. Therefore, the 3' region also provides independent information regarding base recognition at the corresponding locus of the original nucleic acid molecule. See, for example, Figures 4 and 5.

[0069] For example, the sequence of the 5' region of the nucleic acid molecule of the present invention may correspond to a sequence complementary to the sequence of the original nucleic acid molecule. Therefore, the 5' region provides information regarding base recognition in the original nucleic acid molecule. Conversely, the sequence of the 3' region of the nucleic acid molecule of the present invention may correspond to the sequence of the inverse complementary strand of the 5' region. Therefore, the 3' region also provides independent information regarding base recognition at the corresponding locus of the original nucleic acid molecule. See, for example, Figures 4 and 5.

[0070] For example, the sequence of the 5' region of the nucleic acid molecule of the present invention may correspond to the sequence of the original nucleic acid molecule after treatment with a drug (e.g., bisulfite) capable of converting unmethylated cytosines(or more) to bases that can be read separately from cytosines; that is, the sequence of the 5' region of the nucleic acid molecule of the present invention may correspond to the converted original nucleic acid molecule. Therefore, the 5' region provides information regarding base recognition of the original nucleic acid molecule. Conversely, the sequence of the 3' region of the nucleic acid molecule of the present invention may correspond to the sequence of the inverse complementary strand of the 5' region before C-to-T conversion (e.g., before bisulfite treatment; see below for further details). Therefore, the 3' region also provides independent information regarding base recognition at the corresponding locus in the original nucleic acid molecule.

[0071] For example, the sequence of the 5' region of the nucleic acid molecule of the present invention may correspond to the sequence of the original nucleic acid molecule after treatment with an agent (e.g., bisulfite) capable of converting unmethylated cytosines(or more) to bases that can be read separately from cytosines; that is, the sequence of the 5' region of the nucleic acid molecule of the present invention may correspond to the converted original nucleic acid molecule. Thus, the 5' region provides information regarding base recognition of the original nucleic acid molecule. Conversely, the sequence of the 3' region of the nucleic acid molecule of the present invention may correspond to the sequence of the inverse complementary strand of the C-to-T converted (e.g., treated with bisulfite) 5' region. Thus, the 3' region also provides independent information regarding base recognition at the corresponding locus in the original nucleic acid molecule.

[0072] In another embodiment, the sequence of the 3' region of the nucleic acid molecule of the present invention may correspond to the sequence of the original nucleic acid molecule after treatment with an agent (e.g., bisulfite) capable of converting unmethylated cytosines(or more) to bases that can be read separately from cytosines; that is, the sequence of the 3' region of the nucleic acid molecule of the present invention may correspond to the converted original nucleic acid molecule. Thus, the 3' region provides information regarding base recognition of the original nucleic acid molecule. Conversely, the sequence of the 5' region of the nucleic acid molecule of the present invention may correspond to the sequence of the inverse complementary strand of the 3' region before C-to-T conversion (e.g., before bisulfite treatment, further details of which will be discussed later). Thus, the 5' region also provides independent information regarding base recognition at the corresponding locus in the original nucleic acid molecule.

[0073] Those skilled in the art can combine the information obtained from both the 5' and 3' regions of the nucleic acid molecule of the present invention to assign information (base recognition) to each gene locus of the original nucleic acid molecule.

[0074] In a preferred embodiment, the 5' region of the nucleic acid molecule of the present invention provides information regarding the identification of the same locus of bases in the original nucleic acid molecule, and the 3' region of the nucleic acid molecule of the present invention independently provides information regarding the identification of the same locus of bases in the original nucleic acid molecule.

[0075] The following schema provides information and clarification regarding the nomenclature of related sequences according to the present invention: 5'--ATTTGGC----ATTTGGC--3' Both regions (5', ATTTGGC, and 3', ATTTGGC) contain the same sequence in tandem. 5'--ATTTGGC----TAAACCG--3' The 3' region (TAAACCG) has a sequence containing complementary bases at the corresponding locus of the 5' region ("identical complementary" sequence), that is, the 3' region (TAAACCG) has a sequence complementary to the sequence of the 5' region (ATTTGGC). 5'--ATTTGGC----CGGTTTA--3' The 3' region (CGGTTTA) has a sequence that is the reverse of the sequence of the 5' region (ATTTGGC). 5'--ATTTGGC----GCCAAAT--3' The 3' region (GCCAAAT) has a sequence that is the inverse complementary sequence of the 5' region (ATTTGGC); this is a so-called Watson chain and a Crick chain covalently linked by a linking region. The 3' region (GUUAAAT) has a sequence that was complementary to the reverse sequence of the 5' region (ATTTGGC) (GCCAAAT) before both the 3' and 5' regions were treated, for example, with bisulfite or enzymes, and thus "converted" (unmethylated C to U, also called "C-to-T conversion"). If methylated C was present in the original molecule, the 5' region would have A and the 3' region would have G. Any other modifications would function similarly, taking into account the original bases and the converted bases, and the relationship between the 5' and 3' regions.

[0076] In the above case, the "connected area" is represented by "----".

[0077] Therefore, as described above, the nucleic acid molecule of the present invention provides two independent sources of information regarding true base identification at a specific gene locus of the original nucleic acid molecule.

[0078] The nucleic acid molecule of the present invention further comprises one adapter at the 5' end of the molecule and one adapter at the 3' end of the molecule. The terms “adapter” and “adaptor” are used interchangeably herein and refer to oligonucleotides or nucleic acid fragments or segments that can be ligated to the nucleic acid molecule of interest. The “adapter molecule” of the method of the present invention is preferably a DNA molecule having one end that is compatible with the end of the nucleic acid molecule of the present invention (preferably DNA).

[0079] In genetic engineering, an adapter is a short, chemically synthesized single- or double-stranded oligonucleotide that can be ligated to the ends of other DNA or RNA molecules. Adapters may contain "cleavage sites" (e.g., "restriction sites," sequences of oligonucleotides recognized by restriction enzymes). These "cleavage sites" provide an additional method for adapting the final elements of a library to the requirements of different sequencing platforms.

[0080] In one embodiment, at least a portion of the adapter has a sequence common to all adapters present in the population of nucleic acid molecules in step (i). In this case, the same primer can be used to sequence all the molecules.

[0081] Optionally, the adapter includes a unique combinatorial barcode (also called a “combinatorial sequence,” “barcode,” “barcode sequence,” or “combinatorial label”) that enables the identification, multiplexing, pairing, and quantitative analysis of the sample. Constructs obtained by the method of the present invention may have barcodes that enable the generation of unique identifiers associated with the initial construct, thus giving the constructs the ability to distinguish between them. The unique identifier enables the identification of a particular construct containing the identifier and its descendants. Each unique identifier is associated with an individual molecule or fragment of an individual molecule in the starting sample. Thus, any amplification product of the initial individual molecule having a unique identifier is assumed to be identical by its descendants. The combinatorial barcode also enables the quantification of the percentage of individual sequences in the sample and is useful for monitoring bias and error control during the amplification step.

[0082] The barcode sequence adds a "bias control" feature. When amplification occurs, some fragments may be selectively amplified for various reasons. This undesirable effect is a major problem for quantitative purposes, which is important for many applications of sequencing, especially in the analysis of DNA methylation states (quantification and bias control are essential in most applications because each allele in each cell may have a different methylation state, and even the sample may have a heterogeneous composition). Therefore, the presence of at least one barcode sequence makes it possible to have bias control. In this case, each nucleic acid molecule from the multiple molecules provided in step (i) may have one or more different barcode sequences, so it may be possible to perform bias control and detect the selective amplification of a given nucleic acid molecule.

[0083] Preferably, the adapter molecules and / or barcode sequences are provided as a library of molecules, and each member in the library can be distinguished from others by a combinatorial sequence within the sequence, as described below.

[0084] The terms “library of adapter molecules and / or barcode sequences” and / or “combinatorial label” as used in this document refer to an assembly of adapter molecules and / or barcode sequences, each member of the assembly being distinguishable from others by a combinatorial sequence within the adapter and / or hairpin sequence and / or barcode sequence.

[0085] The terms “combinatorial sequence,” “barcode sequence,” “barcode,” and “combinatorial barcode” are all used interchangeably herein and refer to identifiers (barcode sequences of their own, not belonging to an adapter) that are unique to individual adapter sequences or separate nucleic acid (e.g., DNA) molecules. Preferably, the barcode sequence is contained within the adapter. In one embodiment, the combinatorial sequence within the adapter sequence is a degenerate nucleic acid sequence. The combinatorial sequence may contain any nucleotides, including adenine, guanine, thymine, cytosine, uracil, methylated cytosine (e.g., 5mC or 5hmC), and other modified nucleotides. The number of nucleotides in the combinatorial sequence is preferably designed such that the number of potential and actual sequences represented by the combinatorial sequence is greater than the total number of adapters in the library. The combinatorial sequence may be located in any region of the adapter sequence.

[0086] In one specific embodiment, the molecule provided in step i. of the present invention is, for example, the molecule described in International Publication No. 2015 / 104302 (also known as the “GEUS molecule”).

[0087] Method of the present invention Step (ii)Step (ii) includes using at least two different primers, e.g., two, three, or four different primers, preferably four different primers, to sequence the molecule provided in step (i). Thus, step (ii) includes the use of at least two different primers and sequencing the molecule provided in step (i) using at least two different primers. Thus, step (ii) provides sequence information of the molecule provided in step (i) of the method of the present invention. Thus, step (ii) is a sequencing step. See, for example, Figure 6.

[0088] As used in this document, the term "primer" refers to a short strand of nucleic acid that is at least partially complementary to a sequence in another nucleic acid and serves as a starting point for nucleic acid (e.g., DNA) synthesis. Preferably, the primer has a base length of at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 18, at least 20, at least 25, at least 30 or more bases.

[0089] The term "complementary" refers to base pairings that enable the formation of a double helix between nucleotides or nucleic acids, for example, between the two strands of a double-stranded DNA molecule, between an oligonucleotide primer and a primer binding site on a single-stranded nucleic acid, or between an oligonucleotide probe and its complementary sequence in a DNA molecule. Complementary nucleotides are generally A and T (or A and U), or C and G. Two single-stranded DNA molecules are said to be substantially complementary if the nucleotides of one strand are optimally aligned, compared, and paired with approximately 60%, at least 70%, at least 80%, at least 85%, usually at least 90% to 95%, and even 98% to 100% of the nucleotides of the other strand, with appropriate nucleotide insertions or deletions. The degree of identity between two nucleotide regions is determined using computer-implemented algorithms and methods widely known to those skilled in the art. The identity between two nucleotide sequences is preferably determined using the BLASTN algorithm (BLAST Manual, Altschul, S. et al., NCBI NLM NIH Bethesda, Md. 20894, Altschul, S., et al., J., 1990, Mol. Biol. 215:403-410).

[0090] At least two different primers, for example two, three, or four different primers, preferably four different primers, bind to at least three, preferably at least four different regions in the nucleic acid molecule provided in (i): 1. At least one primer can at least partially bind to at least a portion of the adapter at the 5' end of the molecule to sequence at least a portion of the 5' region of the nucleic acid molecule provided in (i); 2. At least one of the primers can at least partially bind to a region of a nucleotide sequence that covalently links the 5' and 3' regions of the nucleic acid molecule provided in (i), thereby sequencing at least a portion of the 3' region of the nucleic acid molecule provided in (i); 3. At least one primer can at least partially bind to at least a portion of the adapter at the 3' end of the molecule to sequence at least a portion of the 3' region of the nucleic acid molecule provided in (i); and / or 4. At least one primer can at least partially bind to a region of a nucleotide sequence that covalently links the 5' and 3' regions of the nucleic acid molecule provided in a., thereby sequencing at least a portion of the 5' region of the nucleic acid molecule provided in (i).

[0091] In one embodiment, at least two different primers, for example two, three, or four different primers, preferably four different primers, bind to at least three different regions, preferably at least four different regions, in the nucleic acid molecule provided in (i): 1. At least one of the primers (e.g., the first primer) can at least partially bind (hybridize) to at least a portion of the adapter at the 5' end of the nucleic acid molecule provided in (i) so that at least a portion of the 5' region of the nucleic acid molecule provided in (i) can be sequenced; 2. At least one of the primers (e.g., the second primer) can at least partially bind to at least a portion of the adapter at the 3' end of the molecule so that at least a portion of the 3' region of the nucleic acid molecule provided in (i) can be sequenced; and 3. At least one of the primers (e.g., the third primer) is at least partially: 3.1. A region of nucleotide sequence covalently linking the 5' and 3' regions of the nucleic acid molecule provided in (i) in order to sequence at least a portion of the 3' region of the nucleic acid molecule provided in (i); or 3.2 (i) To sequence at least a portion of the 5' region of the nucleic acid molecule provided in a., a region of nucleotide sequence covalently linking the 5' and 3' regions of the nucleic acid molecule provided in a. It can be combined (hybridized) with any of the following.

[0092] Preferably, at least two different primers, for example two, or at least three, or at least four different primers, preferably four different primers, bind to at least four different regions, preferably four different regions, in the nucleic acid molecule provided in (i): 1. At least one of the primers can at least partially bind (hybridize) to at least a portion of the adapter at the 5' end of the nucleic acid molecule provided in (i) so that at least a portion of the 5' region of the nucleic acid molecule provided in (i) can be sequenced; 2. At least one of the primers can at least partially bind (hybridize) to a region of a nucleotide sequence that covalently links the 5' and 3' regions of the nucleic acid molecule provided in (i), thereby sequencing at least a portion of the 3' region of the nucleic acid molecule provided in (i); 3. At least one of the primers can at least partially bind (hybridize) to at least a portion of the adapter at the 3' end of the nucleic acid molecule provided in (i) so that at least a portion of the 3' region of the nucleic acid molecule provided in (i) can be sequenced; 4. At least one of the primers can at least partially bind (hybridize) to a region of a nucleotide sequence that covalently links the 5' and 3' regions of the nucleic acid molecule provided in (i), thereby sequencing at least a portion of the 5' region of the nucleic acid molecule provided in (i).

[0093] The primer can at least partially bind (hybridize) to the above sequence under low stringent conditions, preferably medium stringent conditions, and most preferably high stringent conditions.

[0094] The binding of primers to at least three, preferably at least four, different regions in the nucleic acid molecule provided in (i) may occur simultaneously (i.e., all two, three, or four primers simultaneously) or not simultaneously. Preferably, the binding of primers to at least three, preferably at least four different regions in the nucleic acid molecule provided in (i) is not performed simultaneously.

[0095] In a preferred embodiment, the binding of at least two different primers, e.g., two, three, or four different primers, preferably four different primers, to at least three different regions, preferably at least four different regions, in the nucleic acid molecule provided in (i) is specific binding. This means that the primer(s) bind to the above-mentioned regions in the molecule in a specific manner, i.e., they bind to the above-mentioned regions but substantially do not bind to other regions in the nucleic acid molecule provided in (i). Those skilled in the art know means of designing primers and checking their specificity. For example, see the Primer designing tool (nih.gov) provided by the National Library of Medicine (NIH), or "How to: Design PCR primers and check them for specificity (nih.gov)" from the National Library of Medicine (NIH).

[0096] In a preferred embodiment, the at least two different primers that bind to at least four different regions in the nucleic acid molecule provided in (i) are four different primers, each primer specifically and at least partially binding (hybridizing) to regions 1-3 and / or 1-4, preferably 1-4, as described above.

[0097] For example, in the case of Illumina sequencing, primer 1 (which can at least partially bind to at least a portion of the adapter at the 5' end of the molecule, allowing sequencing of at least a portion of the 5' region of the nucleic acid molecule provided in (i)) and primer 2 (which can at least partially bind to the region of the nucleotide sequence that covalently links the 5' and 3' regions of the nucleic acid molecule provided in (i), allowing sequencing of the 3' region of the nucleic acid molecule provided in (i)) should be different, and primer 3 (which can at least partially bind to at least a portion of the adapter at the 3' end of the molecule, allowing sequencing of at least a portion of the 3' region of the nucleic acid molecule provided in (i)) and primer 4 (which can at least partially bind to the region of the nucleotide sequence that covalently links the 5' and 3' regions of the nucleic acid molecule provided in a., allowing sequencing of the 5' region of the nucleic acid molecule provided in (i)) should also be different. However, it is not essential that primers 1 and 3, or 1 and 4, or 2 and 3, or 2 and 4 as defined above are different.

[0098] Therefore, at least two different primers provided in step (ii) of the method of the present invention are used to sequence the molecule provided in step (i). Sequencing can be performed using one or more of the sequencing technologies currently available (e.g., sequencing platforms such as Illumina, Roche, and Ion Torrent).

[0099] Sequencing can be performed either at a low scale, consisting of the analysis of selected fragments, or at a high-throughput level (also called genome-scale), consisting of the analysis of large quantities of representatives of all or most of all substances, such as next-generation sequencing (NGS) approaches. The length of the fragments that can be analyzed depends on the sequencing methodology used. Current state-of-the-art sequencing techniques, aiming for genome scale and being mostly locus-specific, evaluate ss nucleic acid molecules (e.g., DNA strands) individually.

[0100] Terms such as “sequencing,” “determining a sequence,” or “sequencing,” for example, “determining base identification,” or “determining base identification,” mean determining information related to the nucleotide sequence of a nucleic acid, particularly involving the determination and ordering of multiple consecutive nucleotides in the nucleic acid. Such information may include the identification or determination of partial and complete sequence information of a nucleic acid molecule. Such information refers, for example, to the primary sequence of a DNA molecule, such as an ssDNA molecule or a dsDNA molecule, or to epigenetic modifications (e.g., methylation or hydroxymethylation), or both. Sequence information can be determined with varying degrees of statistical reliability or confidence. As described above, the method of the present invention provides high reliability when sequencing nucleic acid molecules.

[0101] The method of the present invention can be used to sequence the primary sequences of DNA molecules, such as ss or dsDNA molecules, or libraries of DNA molecules. Determining the primary sequence of a DNA molecule involves detecting mutations or gene variants, such as polymorphisms (SNPs, INDELs, etc.). Preferably, the method of the present invention allows for the simultaneous determination of both the primary sequence and epigenetic state, such as the cytosine methylation state, in the original nucleic acid molecule using the same read. By analyzing the sequencing output, each read provides information about the primary sequence (including mutations and SNPs) and methylation sequence of the original nucleic acid molecule, as well as the combinatorial sequence contained in the adapter of the nucleic acid molecule provided in step (i) of the method of the present invention.

[0102] As described in this document and as will be understood by those skilled in the art, the fact that at least two different primers (e.g., four different primers) bind to at least three, preferably at least four different regions in the molecule provided in step (i), and that both the 5' and 3' regions are the subject of at least partial sequencing, is as follows: - The 3' region of the nucleic acid molecule provided in step (i), using a primer that at least partially binds (hybridizes) to at least a portion of the adapter at the 3' end of the nucleic acid molecule provided in (i); and - Using a primer that at least partially binds (hybridizes) to at least a portion of the nucleotide sequence that covalently links the 5' and 3' regions of the nucleic acid molecule provided in (i), the 5' region of the nucleic acid molecule provided in step (i) Regarding (i), this means that the complementary sequence of the nucleotide molecule provided in (i) needs to be synthesized and amplified (cluster amplification for synthetic parallel sequencing). This is because nucleic acid molecules are synthesized using so-called "synthetic sequencing" techniques of next-generation sequencing, such as Illumina sequencing, which utilize the synthesis of the original and complementary strands in order to read the base sequence of a particular nucleic acid molecule using primers. For example, a primer attaches to the adapter-primer binding site of the forward strand, and polymerase adds a fluorescently tagged dNTP to the DNA strand. Only one base can be added per round because the fluorescent dye acts as a blocking or synthetic terminator group; however, the blocking group is reversible. Using tetrachromatic chemistry, each of the four bases exhibits a unique emission, and after each round, the instrument records which nucleotide was added. Once the color is recorded, the fluorescent dye is washed away, another dNTP is washed on the flow cell, and the process is repeated. Since polymerase adds nucleotides to the 3' end of a nucleic acid (DNA) strand, the nucleic acid molecule to be sequenced needs to be read from 5' to 3'. Therefore, using at least two different primers (preferably twice) to sequence the 5' and 3' regions of the nucleic acid molecule provided in step (i) means that sequencing of the 5' and 3' regions is performed on a strand complementary to the nucleic acid molecule provided in step (i), using (a) a primer that at least partially hybridizes to at least a portion of the adapter at the 3' end of the nucleic acid molecule provided in (i); and (b) a primer that at least partially hybridizes to at least a portion of the nucleotide sequence covalently linking the 5' and 3' regions of the nucleic acid molecule provided in (i).

[0103] In some cases, the method further includes diagnosing a condition of interest based at least in part on sequencing information provided in step (ii) of the method of the present invention. The condition may be any condition, trait or aging, obesity, etc. For example, the condition may be a cancer, which may be selected from sarcomas, gliomas, adenomas, leukemias such as chronic lymphocytic leukemia (CLL), bladder cancer, breast cancer, colorectal cancer (CRC), endometrial cancer, kidney cancer, liver cancer, lung cancer, melanoma, non-Hodgkin lymphoma, pancreatic cancer, prostate cancer, thyroid cancer, etc. The condition may also be a neurodegenerative condition, such as Alzheimer's disease, frontotemporal dementia, amyotrophic lateral sclerosis, Parkinson's disease, spinocerebellar ataxia, spinal muscular atrophy, Lewy body dementia, Huntington's disease. The condition may also be a genetic or environmental disease, or any rare or common disease, or a trait not necessarily associated with a disease.

[0104] The method of the present invention may further include the step of determining the true identification of a base present at a specific location (locus) in the original nucleic acid molecule based on the information provided in step (ii). In the context of the present invention, “true identification” is the identification of a base originally present at a specific location (locus) in the original nucleic acid molecule (e.g., A, C, G, T, U, or any modified thereof, e.g., modified nucleotide, e.g., modified cytosine, e.g., methylated cytosine, mC (e.g., 5mC, 5hmC and / or 5fC)).

[0105] Therefore, once the nucleic acid molecule provided in step (i) is sequenced as described in step (ii), information regarding the sequence of at least a portion (and preferably) all of the 3' and 5' regions is provided. In particular, at least two sources of information are provided for at least one of the 3' and 5' regions, preferably one of each. Since the 5' and 3' regions of the nucleic acid molecule provided in step (i) of the present invention independently provide information regarding base identification at the corresponding locus in the original nucleic acid molecule, a total of more than two, for example three, preferably four, sources of information regarding base identification at the corresponding locus in the original nucleic acid molecule are provided. Therefore, the method of the present invention can be used primarily to reduce uncertainty and the overall error rate in sequencing the polynucleotide (e.g., original DNA polynucleotide) described in step i) before requiring alignment to a reference genome (or reference nucleic acid sequence). Therefore, the method of the present invention provides more than two, for example three, preferably up to four independent sources of information (e.g., up to eight independent sources of information when a double-stranded nucleic acid molecule is considered) with respect to base identification at each corresponding locus in the original nucleic acid molecule from the single-stranded nucleic acid molecule described in step i). The method of the present invention provides more than two, for example three, preferably up to four (preferably up to eight, when considering double-stranded nucleic acid molecules) sources of information for base identification of each corresponding locus in the original nucleic acid molecule described in step i). Since each of the four nucleotides can be read in a different sequence context in the nucleic acid provided in step (i) of the method of the present invention, the error bias that would normally occur in the base to be analyzed from preceding and succeeding sequences is reduced in the method of the present invention. Since all nucleotides of the original molecule are expressed more than two, for example three, preferably up to four times, the raw probability of error for each base can be greatly reduced, mainly in the pre-mapping step, but also after the mapping step. Reducing the error rate in the pre-mapping step also improves the mapping quality of each read, which reduces mapping errors and therefore variant determination errors.Knowing precisely where all inserts (3' and 5' regions) begin and end improves mapping and SNP detection, but primarily improves the detection of indels and other types of rearrangements. Having UMIs (optional) at the beginning and end of all inserts improves sequencing at the beginning of all reads, enables deduplication which is crucial during enrichment, and reduces the number of unhelpful sequencing cycles and unnecessary bioinformatics resources. It also allows for the ligation of the original molecule's dsDNA strands if they are separated during the method.

[0106] Therefore, in the method of the present invention, the nucleic acid molecule of step (i) is provided as a ds molecule, and the raw base quality of each base is 10 -4 If so, the lowest error rate in pre-mapped inferred reads is 10 -16 These error rates are the lowest error rates (10) obtainable with current sequencing methodologies, such as Illumina MiSeq. -4 It is much lower compared to [another factor].

[0107] Therefore, using the GEUS molecule as an example (as described, for example, in International Publication No. 2015 / 104302), if the first base identification from read 1 and read 3, and the second base identification from read 4 and read 2, do not match any of the following combinations: 1) adenine and adenine corresponding to A, 2) thymine and thymine corresponding to T, 3) thymine and cytosine corresponding to unmodified cytosine (e.g., unmethylated C), 4) cytosine and cytosine corresponding to modified cytosine (e.g., methylated C), or 5) guanine and adenine corresponding to G, then the true base identification at the locus of the original nucleic acid molecule is determined to be a misidentification.

[0108] When inferring from four reads, if the first base identification from read 1 and read 3, and the second base identification from read 4 and read 2, do not match any of the following combinations: 1) adenine, adenine, thymine, and thymine corresponding to A; 2) thymine, thymine, adenine, and adenine corresponding to T; 3) thymine, cytosine, adenine, and guanine corresponding to unmodified cytosine (e.g., unmethylated C); 4) guanine, adenine, cytosine, and thymine corresponding to G; or 5) cytosine, cytosine, guanine, and guanine corresponding to modified cytosine (e.g., methylated C), then the true base identification at the gene locus of the original nucleic acid molecule is determined to be a misidentification.

[0109] For example, if the true base identification at a locus of the original nucleic acid molecule is inferred from two reads, there are 16 possible combinations, of which 5 are possible matching codes and 11 are impossible matching codes (errors that should not be considered). In this case, just one sequencing error could cause A at the original nucleic acid molecule's locus to be determined to be G (identified as G), or T at the original nucleic acid molecule's locus to be determined to be C (identified as C).

[0110] Based on the two leads, there are five possible "matching codes":

[0111] [Table 7] And 11 types of "mismatched codes" (error detection and removal): AAACCCGGGTT CGTAGTCGTAG NNNNNNNNNNN It exists.

[0112] However, when inferring the true base identification at the original nucleic acid molecule's locus from four reads, there are 256 possible combinations, of which 5 are possible matching codes and 251 are impossible matching codes (errors that should not be considered). In this case, at least two sequencing matching errors are required for A at the original nucleic acid molecule's locus to be determined (identified as) as G (e.g., A>G at R1 + T>C at R4), or for T to be determined (identified as) as G.

[0113] Based on the four leads, there are five possible "matching codes":

[0114] [Table 8] And 251 different "mismatch codes" (error detection and removal, or recovery of true identification):

[0115] [Table 9] The file TIFF2026514126000011.tif206170 exists.

[0116] Furthermore, when the method of the present invention is applied to each strand of a double-stranded nucleic acid molecule (e.g., a dsDNA molecule), a total of eight different sources of information can be provided for each gene locus in the original genome. This makes it possible to determine the hemimethylation state (or methylation symmetry of each strand) of the original ds molecule. The terms "hemimethylation" and "asymmetric methylation" are used interchangeably and refer to sequences in double-stranded DNA, such as CpG, where only one of the two strands is methylated.

[0117] Figure 1 provides a schema of the present invention's method, starting with a GEUS molecule as a template (e.g., as described in International Publication No. 2015 / 104302), where both strands of the dsDNA molecule are sequenced using the present invention's method. As shown therein, there are two sources of information for each of the 3' and 5' regions of one strand. Since both the 5' and 3' regions in one of the strands independently provide information about base identification at the corresponding locus in the original nucleic acid molecule, there are four independent sources of information about base identification at a specific location (locus) in the original nucleic acid molecule. Because the original molecule is a ds molecule, there are a total of eight sources of information about base identification at the corresponding locus in the original nucleic acid molecule. See also Figures 2 and 3 and Table 1.

[0118] [Table 10] TIFF2026514126000013.tif94161

[0119] As shown in Table 1, there are eight sources of information for each gene locus in the original molecule.

[0120] Therefore, if there is one source of information for each locus in the original molecule, the method of the present invention can determine true base identification with the same quality as the standard method (from one source). If there are two sources of information for each locus in the original molecule, the method of the present invention can determine true base identification with better quality than the standard method. If there are three sources of information for each locus in the original molecule, the method of the present invention can determine true base identification with even better quality than the standard method. Finally, if there are four sources of information for each locus in the original molecule, the method of the present invention can determine true base identification with the best quality.

[0121] Therefore, the method of the present invention makes it possible to determine the epigenetic state, such as the base recognition in the original molecule including hemimethylation, with reduced error.

[0122] In other cases, the method of the present invention further includes using a computer comprising a processor, memory, and instructions stored therein, which, when executed, determine the identification of a base (e.g., a true base) at a particular location (locus) in the original nucleic acid molecule based on the sequencing information provided in step (ii).

[0123] It should be noted that a processor may include one or more processing units, such as a microprocessor, GPU, CPU, or multi-core processor. Similarly, memory may include one or more volatile or non-volatile memory devices, such as DRAM, SRAM, flash memory, read-only memory, ferroelectric RAM, hard disk drives, floppy disks, magnetic tapes, or optical disks.

[0124] Accordingly, the present invention further provides a computer program that, when executed by a computer, can determine the probability of a true base identification and / or BQ score or error at a particular location (locus) in the original nucleic acid molecule, based on the information provided in step (ii) of the method of the present invention.

[0125] Therefore, once at least two different primers have bound to at least three, preferably at least four, different regions in a nucleic acid molecule, as described in step (ii) of the present invention, the computer program includes instructions to perform locus analysis between different readings to determine the true base at a certain locus of the original nucleic acid to be sequenced. Those skilled in the art will note that different methods by which locus analysis may be performed can be envisioned from this specification, but all of them are within the scope of the present invention.

[0126] The present invention also further provides a computer program that, when executed by a computer, can carry out any of the methods disclosed herein. Such a computer program is therefore transmitted communicably to the electronic components of a sequencing device.

[0127] It should be noted that sequencing equipment typically includes a sample, trays, incubators, funnels, micropipette systems, and many other elements that enable the complete automation of sequencing of specific nucleic acid molecules. Therefore, this embodiment is not limited to sequencing equipment comprising only these elements, nor is it limited to any other equipment that can automate any of the methods disclosed herein, as can be imagined by those skilled in the art.

[0128] Computer program products can be implemented as software, hardware, or a combination of both. Computer program products can be stored in the memory of a sequencing device, or remotely, for example, on a remote server that is communicatively connected to the device.

[0129] A method for obtaining a nucleic acid molecule provided in step (i) of the present invention. Those skilled in the art know methods for obtaining the nucleic acid molecule provided in step (i) of the present invention. Some non-limiting examples are as follows:

[0130] The nucleic acid molecule in step (i) can be produced by: a. To provide double-stranded nucleic acid molecules, for example, a plurality or group of double-stranded nucleic acid molecules. In a preferred embodiment, the multiple double-stranded DNA molecules are fragments of genomic DNA. See, for example, Figure 5.

[0131] As used in this document, "a group or number of double-stranded nucleic acid molecules" refers to an aggregate of ds nucleic acid molecules, which may be genomic DNA (nuclear DNA, mitochondrial DNA, chloroplast DNA, cfDNA, etc.), plasmid DNA, or dsDNA molecules obtained from ss nucleic acid samples (DNA, cDNA, mRNA, etc.). In one embodiment, the group is formed by DNA fragments.

[0132] Preferably, multiple ds nucleic acid molecules are genomic DNA. This may be the whole genome or a reduced representative of the genome. Genomic DNA includes not only nuclear DNA (also called chromosomal DNA) but also DNA of plastids (e.g., chloroplasts) and other organelles (e.g., mitochondria), or circulating / cell-free DNA (cfDNA). As intended in this invention, the term “genomic DNA” includes genomic DNA containing sequences complementary to the sequences described in this document.

[0133] DNA can be fragmented by any suitable method, including, but not limited to, mechanical stress (ultrasound, nebulization, cavitation, etc.), enzymatic fragmentation (enzymatic digestion by restriction endonucleases, nickel endonucleases, exonucleases, etc.), and chemical fragmentation (dimethyl sulfate, hydrazine, NaCl, piperidine, acid, etc.), or it can be fragmented spontaneously, for example, cfDNA fragments that enter the bloodstream during apoptosis or necrosis. In principle, there is no limit to the length of the fragmented DNA fragments, but a narrow range of lengths is preferred for sequencing. The appropriate size of the fragments can be selected before step (a) above. The optimal length ultimately depends on the available sequencing method. In a more preferred embodiment, the double-stranded DNA molecule is a fragment of genomic DNA.

[0134] The multiple ds nucleic acid molecules provided in step (a) above are as follows: [1] For example, to provide a population of RNA-derived ds nucleic acid molecules from genomic DNA; [2]. Separating ds nucleic acids (e.g., derived from genomic DNA) to provide ss nucleic acid molecules (e.g., derived from genomic DNA); [3]. Provide a complementary strand of ss nucleic acid molecules using nucleotides A, G, C, T, and U to obtain the ds nucleic acid molecules (e.g., dsDNA molecules or RNA molecules) provided in step (a). It can be obtained by [method].

[0135] Multiple ds nucleic acid molecules may include nucleic acid molecules containing methylated cytosine in both strands and / or methylated cytosine in one of the strands.

[0136] Typically, the ends of a population of ds nucleic acid molecules (such as dsDNA molecules) are processed so that the sample can enter a specific protocol on a sequencing platform. Optionally, the double-stranded adapter ligated in step b (see below) may contain a "cleavage site" (e.g., a "restriction site," a sequence of oligonucleotides recognized by restriction enzymes).

[0137] Preferably, the ds nucleic acid molecule (e.g., dsDNA molecule) used in step (a) is end-repaired before step (a). The term “end repair,” as used in this document, refers to the conversion of a nucleic acid (e.g., DNA) fragment containing damaged or incompatible 5'- and / or 3'-overhangs into blunt-ended DNA containing a 5'-phosphate group and a 3'-hydroxyl group. DNA end blunting can be achieved by enzymes including, but are not limited to, T4 DNA polymerase (having 5'→3' polymerase activity to fill 5' overhangs) and the Klenow fragment of E. coli DNA polymerase I (having 3'→5' exonuclease activity to remove 3'-overhangs). For efficient phosphorylation of DNA ends, any enzyme capable of adding 5'-phosphate to the ends of an unphosphorylated DNA fragment can be used, including, but is not limited to, T4 polynucleotide kinase. Preferably, the method for obtaining the nucleic acid molecule provided in step (i) of the present invention further comprises a step of dA-tailing the nucleic acid (DNA) molecule after the end repair step.

[0138] The term "dA-tailing," as used in this document, refers to the addition of an A base to the 3' end of a blunt phosphorylated DNA fragment. This procedure creates an overhang that is suitable for subsequent ligation. This step is carried out in a manner well known to those skilled in the art, for example, by using a Klenow fragment of E. coli DNA polymerase I. The multiple ds nucleic acid molecules (e.g., dsDNA molecules) used as starting materials in step (a) above can also be obtained from synthesis from single-stranded molecules, such as ssDNA or cDNA. A population of ds nucleic acid molecules (e.g., dsDNA molecules) can be obtained from cDNA. ds nucleic acid molecules (e.g., dsDNA molecules) can also be obtained from mRNA (e.g., viral RNA) by a method well known in the art, which includes the isolation of mRNA, reverse transcription of RNA to obtain single-stranded cDNA, and treatment of single-stranded DNA to obtain double-stranded DNA.

[0139] Samples used to obtain multiple ds nucleic acid molecules (e.g., dsDNA molecules) can be obtained from biological or environmental sources. Biological samples include, but are not limited to, animal or human samples, and liquid and solid food and feed products (dairy products, vegetables, meat, etc.). The sample to be analyzed may originate from a single source (e.g., a single organism, tissue, or cell) or may be a pool of nucleic acids from multiple organisms, tissues, or cells. Samples used to obtain multiple ds nucleic acid molecules (e.g., dsDNA molecules) may also be synthetic nucleic acid molecules.

[0140] After Step a, you can proceed with Step b1 and / or Step b2: b1. Covalently link the forward and reverse single-stranded nucleic acid molecules provided in step a. And, In step b, the covalent bond is linked by a nucleotide sequence to which the primer can bind, obtaining a nucleic acid molecule containing the 5' and 3' regions. The 5' and 3' regions are covalently linked by nucleotide sequences to which primers can bind. Base identification in either the 5' or 3' region, along with base identification in the other region, independently provides information about the base identification of the corresponding gene locus in the original nucleic acid molecule.

[0141] The definition of "nucleotide sequence to which a primer can bind" has been described above in this specification, and this definition also applies equally to this embodiment.

[0142] For example, the forward and reverse single-stranded nucleic acid molecules provided in step a. may be covalently linked by a "hairpin" or "hairpin loop". For example, the forward and reverse single-stranded nucleic acid molecules provided in step a. may be the Watson strand and the Crick strand of a dsDNA molecule.

[0143] The generated molecules can be treated with reagents that allow for the conversion of unmethylated cytosine to a base (preferably uracil) that is detectably different from cytosine in terms of hybridization properties, in order to analyze the methylation pattern of the sample, as will be described in detail below.

[0144] The nucleic acid molecule of step (i) of the method of the present invention can also be produced as described in International Publication No. 2015 / 104302. In this case, step a is as described in step a above.

[0145] Step b2 of this specific embodiment involves, at least in part, a double-strand (ds) adapter, At least one end, preferably both ends, of the chains of multiple double-stranded nucleic acid molecules This includes ligation. Ligation is preferably carried out under conditions sufficient to ligate the ds adapter to both ends of a ds nucleic acid molecule (e.g., a dsDNA molecule), thereby obtaining multiple adapter-containing nucleic acid molecules (also referred to in this document as "adapter-modified nucleic acid molecules"). See, for example, Figure 5.

[0146] In a preferred embodiment, at least a portion of the double-stranded adapters have an arrangement common to all double-stranded adapters used in step (b).

[0147] In one embodiment, the adapter is a so-called "Y adapter." The "Y adapter" has the shape of a "Y." The terms "Y adapter" and "Y adapter" are used interchangeably, and in the context of the present invention, it refers to an adapter formed by two nucleic acid (preferably DNA) strands, wherein the 3' region of the first nucleic acid strand and the 5' region of the second nucleic acid strand form a double-stranded region by sequence complementarity, and the ends of the double-stranded region formed by the 3' region of the first nucleic acid strand and the 5' region of the second nucleic acid strand of the Y adapter are compatible with the ends of a double-stranded nucleic acid molecule. The expression "3' region," as used in this document, refers to a region of a nucleotide strand, including the 3' end of the strand.

[0148] The term "3' end," as used in this invention, refers to the end of a nucleotide chain having the hydroxyl group of the third carbon in the deoxyribose sugar ring as its terminus. The expression "5' region," as used in this invention, refers to a region of a nucleotide chain located toward the 5' end of the chain. The term "5' end," as used in this document, refers to the end of a nucleotide chain having the fifth carbon in the deoxyribose sugar ring as its terminus. The expression "3' region," as used in this invention, refers to a region of a nucleotide chain located toward the 3' end of the nucleotide chain. The term "3' end," as used in this document, refers to the end of a nucleotide chain having the third carbon in the deoxyribose sugar ring as its terminus.

[0149] In one embodiment, the 3' region of the second nucleic acid (e.g., DNA) strand of the Y-adapter forms a hairpin loop by hybridization between a first segment and a second segment within the 3' region, with the first segment located at the 3' end of the 3' region of the second DNA strand and the second segment located near the 5' region of the second DNA strand. The term “hairpin loop,” as used in this document, refers to a region of DNA formed by unpaired bases that occurs when a DNA strand folds and forms base pairs with another section or segment of the same strand.

[0150] Optionally, the 3' region of the second DNA strand of the Y adapter does not form a hairpin loop through hybridization between the first and second segments within the 3' region.

[0151] In the context of the present invention, a ds adapter (which may or may not be a Y adapter, for example, a DNA adapter) includes at least one barcode sequence in the ds region of the adapter. This allows for pairing between each nucleic acid (such as DNA) strand of at least the original ds nucleic acid molecule, thereby eliminating read duplication and distinguishing between reads originating from the same original sequence or reads that are independent but begin and end at the same locus, which is particularly important for enriching low-input / low-diversity libraries or high-depth whole-genome sequencing. This allows for continued tracking of both strands of each ds nucleic acid fragment originally used in step (a) above.

[0152] As used in this document, the term "sequence complementarity" refers to a shared property between two nucleic acid sequences such that, when the two nucleic acid sequences are aligned antiparallel to each other, the nucleotide bases at each position are complementary.

[0153] Therefore, multiple adapter-containing nucleic acid molecules are obtained. Since the adapter is a double-stranded adapter, at least the 3' region of the first strand and the 5' region of the second strand of the ds adapter form a double-stranded region by sequence complementarity, and the ends of the double-stranded region formed by the 3' region of the first strand and the 5' region of the second strand of the adapter are compatible with the ends of a double-stranded nucleic acid molecule.

[0154] If the adapter is a Y-adapter, the Y-adapter may contain one or more barcode sequences in the 5' region of the first nucleic acid (DNA) strand and / or the 3' region (and / or double-stranded region) of the second nucleic acid (DNA) strand of the Y-adapter, which is formed by two nucleic acid (DNA) strands. Thus, the barcode sequences can be located in the single-stranded region of the Y-adapter molecule and / or the double-stranded region of the Y-adapter. In this case, each original nucleic acid (DNA) strand and its synthetic complementary strand (see step (c)) then pair up.

[0155] In a preferred embodiment, prior to step (c), multiple paired adapter-modified nucleic acid (e.g., DNA) molecules are separated to generate a library of paired adapter-modified nucleic acid (e.g., DNA) molecules.

[0156] c. For each strand of nucleic acid molecule obtained in step (b), Using each strand of the nucleic acid molecule obtained in step (b2) as a template, polymerase elongation is performed from the 3' end of the second nucleic acid strand in the adapter molecule. The process of synthesizing complementary chains ("synthetic complementary chains"). Therefore, a ds nucleic acid (e.g., DNA) molecule is obtained, and the original sense strand and antisense strand of the nucleic acid (e.g., DNA) molecule are connected to each other by a hairpin region (a nucleotide sequence to which a primer can be bound, covalently linking the 5' and 3' regions of the molecule). Physically They are bound. In this specific embodiment, each original strand of a nucleic acid (e.g., DNA) molecule is physically bound to a complementary strand obtained by synthetic extension. The original and complementary strands correspond to the 5' and 3' regions of the nucleic acid molecule provided in step (i) of the method of the present invention. See, for example, Figure 4.

[0157] In one embodiment, the adapter-containing nucleic acid (e.g., DNA) molecule obtained in step (b2) is treated prior to step (c) under conditions sufficient to separate the strands of the adapter-containing molecule. Conditions sufficient to separate the strands of the adapter-containing molecule are, but are not limited, conditions under which denaturation of both strands is achieved, such as by heating the molecule to 94-98°C for 20 seconds to 2 minutes, causing the hydrogen bonds between complementary bases to break and resulting in a single-stranded nucleic acid (e.g., DNA) molecule. Strand separation can also be achieved without heating the molecule by using isothermal techniques, for example, by using a large fragment of a strand-displacement DNA polymerase, such as, but not limited to, Phi29 DNA polymerase or Bacillus stearothermophilus DNA polymerase.

[0158] After ligation of the ds adapter in step (b2), each strand of the nucleic acid (e.g., DNA) molecule obtained in step (b2) is converted into a paired double-stranded nucleic acid (e.g., DNA) molecule by polymerase elongation from the 3' end of the second strand in the adapter molecule, using each strand of the nucleic acid (e.g., DNA) molecule obtained in step (b2) as a template (step (c) above).

[0159] In the context of the present invention, the expression "converting each strand into a paired double-stranded nucleic acid (e.g., DNA) molecule" refers to the synthesis of a strand complementary to each strand, where both strands pair up. Thus, the two strands are physically linked (by covalent bonds) (the 3' region of the second strand of an adapter, e.g., a Y-adapter, forms a hairpin loop by hybridization between a first segment and a second segment within the 3' region, with the first segment located at the 3' end of the 3' region of the second DNA strand, and the second segment located near the 5' region of the second strand), resulting in a double-stranded conformation in which, as described above, a single nucleic acid (e.g., DNA) strand folds on top of itself, thus producing the nucleic acid molecule provided in step (i) of the method of the present invention.

[0160] Therefore, the original nucleic acid strand of a nucleic acid (e.g., DNA) molecule and its synthetic complementary strand are physically linked by one of their ends by a loop, which is a nucleotide sequence to which a primer can at least partially bind, as defined above. Each of these molecules may also become linear conformational if the complementarity between the two chains is partially or completely lost. See Figure 5. Thus, the nucleic acid molecule provided in step (i) of the method of the present invention is generated. In this case, the nucleic acid molecule is: - The 5' and 3' regions corresponding to the original chain and its synthetic complementary chain; - A nucleotide sequence that corresponds at least partially to a hairpin loop, covalently linking the 5' and 3' regions, and to which a primer can bind; - One adapter at the 5' end of the molecule and one adapter at the 3' end of the molecule, corresponding to the adapter chain that does not contain a hairpin loop. Includes.

[0161] For further details, please refer to Figure 5.

[0162] In the context of this embodiment, the terms “hairpin adapter,” “hairpin sequence,” and / or “hairpin molecule” refer to a double helix formed by single-stranded nucleic acids that double back on itself and form a double-stranded region maintained by base pairing between complementary base sequences of the same strand, a hairpin loop region formed by unpaired bases, and a terminal that is compatible with the terminal in the molecule provided in step (a). The hairpin adapter may include a blunt end or an overhanging end, preferably an overhanging end.

[0163] As used in this document, the term "polymerase elongation" refers to the synthesis of a complementary strand by DNA polymerase, which adds a free nucleotide to the 3' end of the second DNA strand in the adapter molecule. The adapter molecule can function as a primer for the elongation step, as described above. During this step, the temperature is selected according to the optimal temperature of the specific DNA polymerase used.

[0164] After step (c), two double-stranded nucleic acid (e.g., DNA) molecules are obtained from each adapter-containing nucleic acid (e.g., NA) molecule, each of which is formed by the original nucleic acid (e.g., DNA) strand ("5' region or 3' region") of the nucleic acid (e.g., DNA) molecule and its synthetic complementary strand ("3' region or 5' region"), which are at least physically paired by a nucleotide sequence to which the primer can bind, as described above. They can also be paired by a barcode sequence, as described above.

[0165] The complementary strands of the multiple paired adapter-modified DNA molecules obtained in step (c) can be provided using primers whose sequences are complementary to at least a portion of the double-stranded adapter. The complementary strands of the multiple paired adapter-modified DNA molecules obtained in step (c) can be provided using nucleotides A, G, C, T, and their modifications, such as modified nucleotides, such as modified cytosine, mC (e.g., 5mC, 5hmC, or 5fC).

[0166] Selectively, the paired double-stranded nucleic acid molecules obtained in step (c) Amplify It provides amplified, paired double-stranded nucleic acid molecules.

[0167] The pairing of the two strands of the original double-stranded nucleic acid (e.g., DNA) molecule allows both strands of each double-stranded nucleic acid (e.g., DNA) fragment originally used to be tracked.

[0168] Therefore, each adapter may include a unique, combinatorial barcode that enables the identification, multiplexing, and quantitative analysis of the sample. In a preferred embodiment, the adapter, for example, the Y adapter, is provided as a library of adapters, where each member of the library can be distinguished from others by a combinatorial sequence located within a double-stranded region formed by the 3' region of the first strand and the 5' region of the second strand of the adapter.

[0169] Optionally, the Y adapter incorporates a base labeled with the second member of the binding pair, which allows for the recovery of the original nucleic acid (e.g., DNA) template after the extension or amplification step. This provides the advantage of identifying the sample used as the nucleic acid (e.g., DNA) template, storing it during the process, recovering and storing it, and subjecting it to multiple amplifications under different conditions and sequencing without exhausting the sample.

[0170] Optionally, the adapter (e.g., a Y-adapter) may include a "cutting section" as described above.

[0171] In one embodiment, if the ds adapter does not contain a hairpin loop, the molecule generated in step (b2) is brought into contact with the hairpin adapter under conditions sufficient for ligation of the hairpin adapter to the molecule generated in step (b2), as detailed below. If the adapter ligated in step (b2) does not contain a hairpin loop, the hairpin adapter can be ligated to the adapter ligated in step (b2). For example, the hairpin adapter can be incorporated as described in International Publication No. 2015 / 104302.

[0172] Optionally, this embodiment of the method for generating nucleic acid molecules provided in step (i) of the present invention includes the following step (c1) after step (c): - Each strand of the adapter-containing nucleic acid molecule is brought into contact with a complex of an extension primer and a hairpin adapter under conditions sufficient for hybridization of the extension primer to the second strand of the adapter, wherein the extension primer is complementary to the second strand of the adapter molecule and includes a 3' region that forms an overhang end after hybridization with the second strand of the adapter molecule, and the hairpin adapter includes a hairpin loop region and an overhang end that is compatible with the overhang end of the extension primer formed after hybridization with the second strand of the adapter molecule.

[0173] Optionally, the method for generating the nucleic acid molecule provided in step (i) of the present invention is further advanced after step (c) to the following steps (c21 and / or c22, respectively): - Converting unmodified nucleotides (e.g., unmethylated cytosines) in a paired adapter-modified nucleic acid molecule to bases that, if present, can be read separately from unmodified nucleotides (e.g., uracil) in a paired adapter-modified DNA molecule (c21); and / or - Converting modified nucleotides (e.g., methylated cytosines) in the paired adapter-modified nucleic acid molecule into bases that, if present, can be read separately from modified nucleotides (e.g., cytosines) in the paired adapter-modified DNA molecule (c22) Includes.

[0174] For example, optionally, the method for generating the nucleic acid molecule provided in step (i) of the present invention may be further performed after step (c) in the following steps (c21 and / or c22, respectively): - Converting the unmethylated cytosine(s) in the paired adapter-modified nucleic acid molecule to a base (e.g., uracil) that, if present, can be read separately from the cytosine in the paired adapter-modified DNA molecule (c21); and / or - Converting methylated cytosines (or multiple methylated cytosines) in the paired adapter-modified nucleic acid molecule into bases that, if present, can be read distinctly from the cytosines in the paired adapter-modified DNA molecule (c22). Includes.

[0175] Therefore, step c2 (in any of its other forms) provides the conversion of an unmodified (e.g., unmethylated) (c21) or modified (e.g., methylated) (c22) nucleotide (e.g., cytosine) in a nucleic acid molecule provided in step (i) of the method of the present invention.

[0176] For example, in this optional embodiment, to analyze the epigenetic modification state (e.g., methylation pattern) of the sample, the nucleic acid molecule obtained in step (c) is treated with a reagent capable of converting (or transforming) the nucleotide into another nucleotide that can be read in a way that distinguishes it from the original nucleotide (e.g., a reagent capable of converting unmethylated cytosine into a base (preferably uracil) that is detectably different from cytosine in terms of hybridization properties). (Optional step c21).

[0177] In the context of this invention, the term “epigenetic modification” refers to any chemical modification that may be present in one or more nucleotides but does not alter the nucleotide sequence (i.e., does not alter the genetic code sequence of the nucleotide). Epigenetic modifications, or “tags,” such as DNA methylation, alter the contactability and chromatin structure of DNA, thereby altering the regulatory pattern of gene expression. Therefore, in the context of this invention, epigenetic modifications may be present in any nucleotide in a DNA or RNA sequence, i.e., A, T, U, G, and / or C. In the context of this invention, “epigenetic modification” may also be used interchangeably with the term “chemical modification” of a nucleotide. Therefore, “modified nucleotide,” “chemically modified nucleotide,” or “epigenetic modified nucleotide” refers to a nucleotide that has been chemically modified by an epigenetic modification or “tag.” Therefore, in the context of this invention, a “modified nucleotide” is a nucleotide that has a different structure from a primary nucleotide (guanine, cytosine, thymine, uracil, or adenine) because it contains, for example, an epigenetically modified base. Therefore, as described above, a modified nucleotide can be a nucleotide that carries "epigenetic information," that is, a nucleotide that carries "epigenetic modification." Preferably, the epigenetic modified base is a methylated base, a hydroxymethylated base, a formylated base, an acetylated base, or a carboxylic acid-containing base. Preferably, the epigenetic modification is methylation, and the modified base is a modified (e.g., methylated) cytosine.

[0178] Epigenetic modification sometimes refers to cytosine methylation. In this case, the nucleotide sequence (C) does not change, but the cytosine of the nucleotide is chemically modified by the incorporation of a methyl group. Therefore, cytosine is chemically modified by methylation. In differentiated mammalian cells, the main epigenetic tag found in DNA is the covalent bonding of a methyl group to the C5 position of the cytosine residue in the CpG dinucleotide sequence (see, for example, Handy DE. et al. "Epigenetic modifications: basic mechanisms and role in cardiovascular disease", Circulation, 2011, 123(19):2145-56). Along with histone modifications, DNA methylation modifies chromatin structure and influences the expression of congeneral genes by maintaining various expression patterns across cell types. The presence of DNA methylation in the promoter region is directly linked to transcriptional repression. In contrast, DNA methylation in the gene itself shows a positive correlation with gene expression. Of particular note is that epigenetic modifications may also be associated with the presence of disease. For example, 5mC oxidized derivatives are sometimes used as markers for the diagnosis and prognosis of cancer (see, e.g., Chen K., Zhao BS. and He C., "Nucleic acid modifications in regulation of gene expression", Cell Chem Biol., 2016;23(1):74-85). In differentiated mammalian cells, the main epigenetic modification found in DNA is the covalent bonding of a methyl group to the C5 position of cytosine residues in CpG dinucleotide sequences (called CpG), but cytosines other than CpG can also be similarly methylated. See, e.g., Handy DE. et al., "Epigenetic modifications: basic mechanisms and role in cardiovascular disease", Circulation, 2011;123(19):2145-5.Chemical or epigenetic modifications that occur on nucleotides include, for example, 5-methylcytosine (5mC) and its oxidized derivatives (e.g., 5-hydroxymethylcytosine (5hmC), 5-formylcytosine (5fC) and 5-carboxylcytosine (5caC)) in DNA and N. 6 -methyladenine (6mA); N in messenger RNA and long non-coding RNA 6 ​​​​​​​​A) is also present in mRNA. See, for example, Chen K., Zhao BS. and He C., "Nucleic acid modifications in regulation of gene expression", Cell Chem Biol., 2016;23(1):74-85. Furthermore, pseudouridine is a ubiquitous component of structural RNA (transfer RNA, ribosomal RNA, micronuclear RNA, and micronucleolar RNA). See, for example, Charette M and Gray MW, "Pseudouridine in RNA: what, where, how, and why", IUBMB Life, 2000;49(5):341-51. Cytosine can also be methylated in RNA to form 5mC. tRNA modifications are known to affect translation and influence various physiological processes. For example, in S. cerevisiae, there are 74 genes involved in the introduction of approximately 25 chemically distinct modifications presented at 36 positions on yeast cytoplasmic tRNA. See Chen K., Zhao BS. and He C., "Nucleic acid modifications in regulation of gene expression", Cell Chem Biol., 2016;23(1):74-85.

[0179] As used in this document, the term "modified cytosine" refers to a cytosine base that has been modified by the substitution or addition of one or more atoms or chemical groups, such as a methyl group.

[0180] Modified cytosines, such as methylated cytosines, are resistant to treatment with reagents, such as bisulfite and A3A. This is because the cytosines remain unchanged after treatment with these reagents (e.g., they remain cytosines), or, during treatment or after copying (e.g., after amplification by PCR), they are converted to bases complementary to guanine and read as (unmodified, e.g., unmethylated) cytosines in polymerase base amplification and sequencing (e.g., 5-hydroxymethylcytosine, which is converted to cytosine-5-methylsulfonate, or 5mC / 5fC, which is converted to 5hmC / 5caC, respectively, after treatment with TET-methylcytosine dioxygenase 2). Preferably, the modified cytosines are 5-methylcytosine (5mC), 5-hydroxymethylcytosine (5hmC), or 5-formylcytosine (5fC).

[0181] In the context of the present invention, "easily modifiable nucleotide" refers to any nucleotide that may carry epigenetic modifications, as described above. For example, cytosine is a preferred easily modifiable nucleotide. Cytosine may or may not be methylated (mC) or unmethylated (uC), meaning that cytosine can carry at least one epigenetic modification (e.g., methylation). Preferably, the easily modifiable nucleotide is cytosine.

[0182] In the context of this invention, the term “agent capable of converting (or transforming) a nucleotide into another nucleotide that can be read indistinguishable from the original nucleotide” means any agent (e.g., a reagent, reagent, or enzyme) or method or process capable of converting or transforming (i.e., changing the chemical structure of) a particular nucleotide, the converted or transformed nucleotide being read (recognized) by an enzyme (e.g., polymerase) responsible for copying nucleic acid molecules as a different nucleotide from the original nucleotide. Therefore, once an agent converts / transforms a modified / unmodified nucleotide, the enzyme (e.g., polymerase) responsible for copying nucleic acid molecules introduces a nucleotide at that position that is different from the nucleotide the enzyme would have introduced if the original nucleotide had not been modified. For example, an agent capable of converting a nucleotide base into another base that can be read indistinguishable from the original base may be bisulfite. Bisulfite can convert uC to U by deamination of C. U is read differently by polymerase than C. That is, when polymerase reads U, it introduces A instead of G, which it would introduce if it read C. Other agents exist that can convert or transform nucleotides into other nucleotides that can be read distinctly from the original nucleotide. Examples of other agents include, but are not limited to, deamination agents, metasulfites, or cytidine deaminases, such as activation-inducible cytidine deaminase (AID). For example, the enzyme β-glycosyltransferase can glycosylate 5hmCs, and the enzyme APOBEC3A cytosine deaminase (A3A) can deaminate uC to u. The enzyme 10-11 translocation (TET) methylcytosine dioxygenase 2 can oxidize 5mC to 5hmC, 5fC, or 5caC.For example, AID / APOBEC family enzymes can convert a nucleotide base into another base that is distinguishable from the original base. See, for example, Berney, M. and McGouran, JF, "Methods for detection of cytosine and thymine modifications in DNA", Nat Rev Chem, 2018, 2, 332-348. AID / APOBEC family enzymes can deaminate mC to T (Nabel CS. et al., "AID / APOBEC deaminases disfavor modified cytosines implicated in DNA demethylation", Nat Chem Biol., 2012, 8(9):751-8).

[0183] In one embodiment, a bisulfite is a chemical agent capable of converting a nucleotide base to another base that is read distinctly from the original base. As described above, sodium bisulfite (commonly known as "bisulfite") selectively converts unmethylated cytosine to uracil by deamination without altering methylated cytosine (both 5-methylcytosine and 5-hydroxymethylcytosine). As used in this text, the bisulfite ion has the conventional meaning of HSO3-. Typically, bisulfite is used as an aqueous solution of a bisulfite, such as sodium bisulfite having the formula NaHSO3, or magnesium bisulfite having the formula Mg(HSO3)2. Suitable counterions for bisulfite compounds are monovalent or divalent. Examples of monovalent cations include, but are not limited to, sodium, lithium, potassium, ammonium, and tetraalkylammonium. Suitable divalent cations include, but are not limited to, magnesium, manganese, and calcium. When DNA is treated with bisulfite, unmethylated cytosine bases are converted to uracil, while 5-methylcytosine bases remain unaffected. This conversion is performed using standard methods (Frommer et al. 1992, Proc Natl Acad Sci USA, 89:1827-31; Olek, 1996, Nucleic Acid Res. 24:5064-6; EP 1394172). Methods for obtaining samples include those used in reduced representation bisulfite sequencing (RRBS).

[0184] In another embodiment, the agent capable of converting a nucleotide base to another base that can be read in distinction from the original base is A3A. In yet another embodiment, the agent capable of converting a nucleotide base to another base that can be read in distinction from the original base is the enzyme β-glycosyltransferase.

[0185] When used in this document, the phrase "a base that is detectably different from a particular nucleotide in terms of hybridization properties" refers to a base in the complementary strand that is otherwise complementary to it and cannot hybridize (i.e., does not have a hydrogen bridge) (for example, adenine if the original base is guanine; another example is uracil if the original base is cytosine).

[0186] The phrase "a base that is detectably distinct from cytosine in terms of hybridization properties," as used in this document, means a base that cannot hybridize with guanine (such as uracil) in the complementary chain (i.e., a hydrogen bridge is not present). Preferably, the base that is detectably distinct from cytosine is thymine or uracil, more preferably uracil. The reagent used in this step may be a reagent that can convert unmethylated cytosine to a base that is detectably distinct from cytosine in terms of hybridization properties, but cannot act on methylated cytosine. As discussed above, examples of such agents are, but are not limited to, deamination agents, bisulfites, metabisulfites, or cytidine deaminases, such as activation-inducible cytidine deaminase (AID). In a preferred embodiment, the reagent is a bisulfite.

[0187] Preferably, the conversion of (unmethylated) cytosine in the paired DNA molecule to uracil is carried out using a deamination agent, such as bisulfite; however, as described above, any other agent or enzymatic treatment (e.g., TET oxidation of modified cytosine followed by APOBEC deamination of unmodified cytosine) can be used.

[0188] In one specific embodiment of step (c2), the presence of modified cytosine (e.g., 5C-modified cytosine, e.g., methylated cytosine) at a given position is determined by whether the cytosine, if present, appears on one of the strands of the paired double-stranded nucleic acid molecule obtained in step (c2) or step (d or e), and whether guanine appears at the corresponding position on the other strand of the paired double-stranded nucleic acid molecule, and / or the presence of unmethylated cytosine at a given position is determined by whether uracil or thymine, if present, appears on one of the strands of the paired double-stranded nucleic acid molecule obtained in step (c2) or step (d or e), and whether guanine appears at the corresponding position on the other strand of the paired double-stranded nucleic acid molecule.

[0189] International Publication No. 2015 / 104302 describes exemplary embodiments of the molecule provided in step i. of the present invention.

[0190] The present invention kit The present invention further provides a kit comprising at least two different primers, wherein the at least two different primers can at least partially bind to at least three, preferably at least four, different regions in a nucleic acid molecule provided in step (i) of the method of the present invention.

[0191] In one embodiment, at least two, for example three or preferably four different primers can at least partially bind to at least three, preferably at least four different regions in the nucleic acid molecule provided in step (i) of the method of the present invention: 1. At least one primer can at least partially bind to at least a portion of the adapter at the 5' end of the molecule to sequence at least a portion of the 5' region of the nucleic acid molecule provided in (i); 2. At least one of the primers can at least partially bind to a region of a nucleotide sequence that covalently links the 5' and 3' regions of the nucleic acid molecule provided in (i), thereby sequencing at least a portion of the 3' region of the nucleic acid molecule provided in (i); 3. At least one primer can at least partially bind to at least a portion of the adapter at the 3' end of the molecule to sequence at least a portion of the 3' region of the nucleic acid molecule provided in (i); and / or 4. At least one primer can at least partially bind to a region of a nucleotide sequence that covalently links the 5' and 3' regions of the nucleic acid molecule provided in a., thereby sequencing at least a portion of the 5' region of the nucleic acid molecule provided in (i).

[0192] In one embodiment, at least one of the primers (e.g., a first primer) can at least partially bind (hybridize) to at least a portion of the adapter at the 5' end of the nucleic acid molecule provided in (i) to sequence at least a portion of the 5' region of the nucleic acid molecule provided in (i). Furthermore, at least one of the primers (e.g., a second primer) can at least partially bind to at least a portion of the adapter at the 3' end of the molecule to sequence at least a portion of the 3' region of the nucleic acid molecule provided in (i). Finally, at least one of the primers (e.g., a third primer) can at least partially: - To sequence the 3' region of the nucleic acid molecule provided in (i), to the region of the nucleotide sequence that covalently links the 5' and 3' regions of the nucleic acid molecule provided in (i); or - (i) To sequence the 5' region of the nucleic acid molecule provided in (i), a region of the nucleotide sequence that covalently links the 5' and 3' regions of the nucleic acid molecule provided in a. It can be combined (hybridized) with any of the following.

[0193] In a preferred embodiment, the present invention provides a kit comprising at least two different primers, for example, at least three different primers, preferably four different primers, the at least two different primers, for example, at least three different primers, preferably four different primers, capable of at least partially binding to at least four different regions in the nucleic acid molecule provided in step (i) of the method of the present invention: 1. At least one of the primers can at least partially bind (hybridize) to at least a portion of the adapter at the 5' end of the nucleic acid molecule provided in (i) so that at least a portion of the 5' region of the nucleic acid molecule provided in (i) can be sequenced; 2. At least one of the primers can at least partially bind (hybridize) to a region of a nucleotide sequence that covalently links the 5' and 3' regions of the nucleic acid molecule provided in (i), thereby sequencing at least a portion of the 3' region of the nucleic acid molecule provided in (i); 3. At least one of the primers can at least partially bind (hybridize) to at least a portion of the adapter at the 3' end of the nucleic acid molecule provided in (i) so that at least a portion of the 3' region of the nucleic acid molecule provided in (i) can be sequenced; 4. At least one of the primers can at least partially bind (hybridize) to a region of a nucleotide sequence that covalently links the 5' and 3' regions of the nucleic acid molecule provided in (i), thereby sequencing at least a portion of the 5' region of the nucleic acid molecule provided in (i).

[0194] In a preferred embodiment, the kit of the present invention further includes instructions for its use. In another preferred embodiment, the kit of the present invention further includes a double-stranded adapter for use in a method for generating nucleic acid molecules provided in step (i) of the method of the present invention, the adapter comprising a first nucleic acid strand and a second nucleic acid strand, wherein the 3' region of the first nucleic acid strand and the 5' region of the second nucleic acid strand form a double-stranded region by sequence complementarity, the ends of the double-stranded region formed by the 3' region of the first nucleic acid strand and the 5' region of the second nucleic acid strand of the adapter are compatible with the ends of a double-stranded nucleic acid molecule, the double-stranded region of the adapter comprises one or more barcode sequences, and the 3' region of the second strand of the adapter is Hybridization between a first segment and a second segment within the 3' region forms a hairpin loop, the first segment is located at the 3' end of the 3' region of the second DNA strand, the second segment is located near the 5' region of the second strand, and / or the adapter includes at least one barcode sequence in the single-stranded region of the adapter, the barcode sequence consisting of a unique identifier that enables the identification of a specific construct and its amplification product, and compatibility means that the ends of the double-stranded region of the adapter molecule are ligable to one or both ends of a double-stranded nucleic acid molecule.

[0195] Preferably, the adapter has a restriction region in the 5' region of the first strand of the adapter. In another preferred embodiment, the adapter includes at least one barcode sequence in the single-strand region of the adapter, and the 3' region of the second strand of the adapter forms a hairpin loop by hybridization between a first segment and a second segment within the 3' region, with the first segment located at the 3' end of the 3' region and the second segment located near the 5' region of the second strand.

[0196] In another preferred embodiment, the kit is as follows: (i) A library of double-stranded adapters, wherein the adapter comprises a first strand and a second strand, the 3' region of the first strand and the 5' region of the second strand form a double-stranded region by sequence complementarity, and the ends of the double-stranded region are compatible with the ends of a double-stranded nucleic acid molecule; (ii) Multiple extension primers, each extension primer being complementary to the second chain of the adapter molecule as defined in (i), and including a 3' region that forms an overhang end after hybridization with the second chain of the adapter molecule; and (iii) Multiple hairpin adapters, each hairpin adapter comprising a hairpin loop region and an overhang end that is compatible with the overhang end formed after the hybridization of the extension primer defined in (ii) and the second chain of the Y adapter defined in (i). It further includes, (ii) the extension primer and (iii) the hairpin adapter may be provided as a composite; The adapter (i), the extension primer (ii), and the hairpin adapter (iii) are suitable for obtaining a library of adapters for use in a method for generating nucleic acid molecules provided in step (i) of the method of the present invention.

[0197] In the context of the present invention, the term "sequence recognition" refers to the relationship between two nucleotide sequences or two amino acid sequences. For the purposes of the present invention, the degree of sequence recognition between two nucleotide sequences or two amino acid sequences is determined using the multiple sequence alignment tool Clustal Omega (https: / / www.ebi.ac.uk / Tools / msa / clustalo / ; Sievers, F. et al., 2011, "Fast, scalable generation of high-quality protein multiple sequence alignments using Clustal Omega", Mol.Syst.Biol., 7:539) with a standard parameter t.

[0198] When used in this document, the term “approximately” (or “about”) means the indicated value ± 1% of that value, or the term “approximately” means the indicated value ± 2% of that value, or the term “approximately” means the indicated value ± 5% of that value, or the term “approximately” means the indicated value ± 10% of that value, or the term “approximately” means the indicated value ± 20% of that value, or the term “approximately” means the indicated value ± 30% of that value; preferably, the term “approximately” means exactly the indicated value (±0%).

[0199] Throughout this specification and the claims, the word “comprise” and its variations (e.g., “comprising,” “having,” “including,” “containing”) are typically not restrictive and therefore do not preclude other features, which may be, for example, technical features, materials, substrates, components, or steps. However, whenever the word “comprise” is used in this document, it also includes special embodiments in which the word is understood to be restrictive; in these special embodiments, the word “comprise” has the meaning of the term “consisting of.”

[0200] In the context describing the present invention (particularly in the context of the following claims), the use of the terms “a,” “an,” and “the,” and similar references, shall be interpreted as covering both singular and plural forms unless otherwise indicated herein or unless the context clearly contradicts this interpretation. Unless otherwise indicated herein, the descriptions of value ranges are intended merely as abbreviations for individually referring to each distinct value falling within the range, and each distinct value is incorporated into this document as if it were individually memorized. All methods described herein may be performed in any appropriate order unless otherwise indicated herein or unless the context clearly contradicts this interpretation.

[0201] Any and all embodiments or exemplary language (e.g., “for example”) provided herein are intended solely to better illustrate the invention and, unless otherwise asserted, do not constitute a limitation on the scope of the invention. Nothing in this document shall be construed as indicating that any unclaimed element is essential to the implementation of the invention.

[0202] While the aforementioned invention is described in some detail for clarity and understanding purposes, it will be understood by those skilled in the art that various modifications in form and detail can be made without departing from the true scope of the invention and the appended claims.

[0203] The present invention will be described below with reference to the following examples, but these are merely illustrative and should not be considered to limit the scope of the invention in any way.

[0204] [Examples] Step i of the method of the present invention: This section provides examples of a method for providing the molecule of step i of the claimed method. In particular, the molecule illustrated in this document is the GEUS molecule described in International Publication No. 2015 / 104302. As stated above, this molecule is an example of the molecule defined in step i of the method of the present invention. The advantages and effects described in this document for the GEUS molecule are equally applicable to any molecule defined in step i of the method of the present invention.

[0205] The nucleic acid molecule in step i can be produced by: Step a. Apply the double-stranded nucleic acid molecule: For example, see Figure 5A: ATCGAAMGMT TAGCTTGMGA ("M" indicates methylated cytosine, and "C" indicates unmethylated cytosine.)

[0206] Step b2. Ligate at least partially a double-strand (ds) adapter to at least one end, preferably both ends, wherein one end of the adapter includes a hairpin (hereinafter referred to as "hairpin"): For example, please refer to Figure 5B. Adapter - ATCGAAMGMT - Hairpin Hairpin-TAGCTTGMGA-Adapter

[0207] Step c. For each strand of nucleic acid molecule obtained in step (b), Using each strand of the nucleic acid molecule obtained in step (b2) as a template, polymerase elongation from the 3' end of the "hairpin" is performed. The process of synthesizing complementary chains ("synthetic complementary chains"). For example, please refer to Figure 5D. Adapter - ATCGAAMGMT - Hairpin - AGCGTTCGAT Adapter - AGMGTTCGAT - Hairpin - ATCGAACGCT

[0208] Step c.21: Take the unmethylated cytosine(s) treatment This involves converting cytosine to a base that is read separately (e.g., uracil / thymine): For simplicity, the inventors continue this embodiment using only one of the molecules obtained after step c. However, it should be noted that the method can be continued using both (see Figure 5E). TIFF2026514126000014.tif8160 (Underline the converted unmethylated cytosine)

[0209] The resulting molecule is, for example, a “GEUS molecule” as described in International Publication No. 2015 / 104302, and comprises a 5' region (bold) and a 3' region (italic), both of which are covalently linked by a nucleotide sequence to which a primer can bind (including the adapter hairpin, as described above, also referred to as the “linking region”), and the base identification in either the 5' or 3' region and the base identification in the other region independently provide information regarding the base identification of the corresponding gene locus of the original nucleic acid molecule (in this case, they provide information regarding the base identification of the original strand in step a “ATCGAAMGMT”), and the molecule comprises one adapter at the 5' end and another adapter at the 3' end. The molecule is an example of the molecule provided in step i of the method of the present invention, and is also shown in Figure 6A.

[0210] Step II of the method of the present invention. The base sequence of the molecule provided in step i. is determined using at least two primers that bind to at least three, preferably at least four, different regions. In the examples, four different primers that bind to four different regions in the nucleic acid molecule provided in (i) are used, as described below. However, those skilled in the art will immediately recognize that at least some of the advantages and effects described herein are equally applicable to a method in which at least two different primers that bind to at least three different regions in the nucleic acid molecule provided in (i), for example, three different primers, are used.

[0211] In this embodiment, as described above, four different primers are used: Primer 1 (p1): Binds to a portion of the adapter at the 5' end of the GEUS molecule, enabling sequencing of at least a portion of the 5' region of the provided GEUS molecule (primer corresponding to item 1 of the claimed method, step ii). Primer 2 (p2): Binds to the hairpin region and enables sequencing of the 3' region of the GEUS molecule (primer corresponding to item 2 of the requested method, step ii). Primer 3 (p3): Binds to the adapter at the 3' end of the GEUS molecule, enabling sequencing of at least a portion of the 3' region of the GEUS molecule (primer corresponding to item 3 of the requested method, step ii). Primer 4(p3): Binds to the hairpin region and enables sequencing of the 5' region of the GEUS molecule (primer corresponding to item 4 of the requested method, step ii).

[0212] The four primers described above generate four leads; see also Figure 6B. p1: Lead 1 ATTGAACGCT p3: Lead 3 ATCAAACACT p4: Lead 4 TAACTTGCGA p2: Lead 2 TAGTTTGTGA

[0213] Importantly, considering only reads 1 and 3 derived from p1 and p3, the GEUS molecule is actually sequenced based on single-ended (SE) sequencing, since reads 1 and 3 provide information about bases located at the same locus of the original molecule. Because this GEUS molecule contains two regions (the 5' region highlighted in bold and the 3' region highlighted in italics) that provide information about the base identification of the corresponding locus of the original nucleic acid molecule (the strand ATCGAAMGMT in step a above), SE sequencing is actually performed, as shown in Figure 7A, by using only two conventional primers that hybridize to the molecule's adapter (i.e., "normal PE sequencing").

[0214] However, the requested method yields two additional reads, named read 2 and read 4. These reads originate from primers that hybridize at the hairpin region of the GEUS molecule (p2 and p4, respectively; see Figure 6B). Obtaining these two additional reads yields true paired-end (PE) sequencing for each of the 3' and 5' regions of the GEUS molecule, as the 5' and 3' regions are read from their respective ends, as shown above and in Figure 7B.

[0215] Therefore, the method of the present invention enables PE sequencing for molecules defined in step i of the method, for example, the GEUS molecule described in International Publication No. 2015 / 104302. For the advantages of PE sequencing over SE sequencing as described above, please refer to the background of the present invention.

[0216] In addition, importantly, when the four reads can at least partially, preferably completely cover the 5' and 3' regions of the GEUS molecule, the method claimed in this document provides more than two, for example at least three, preferably at least four information sources (i.e., four reads) per original single-stranded molecule (for example, when the original molecule is a double-stranded molecule as in this example, a total of up to eight sources). The presence of more than two, for example three, preferably four information per original strand is advantageous in that it allows detection of sequencing errors that would not be detected if only two reads were obtained, as described above and below in this document:

[0217] If only p1 and p3 are used, referring to the above, the following reads are obtained. See also Figure 7A:

[0218]

Table 11

[0219] The underlined bases in the reads mean that they correspond to the true bases in the original DNA template. For example, in Read 1, if an "A" is read, it means that the original DNA template has an "A" at that position.

[0220] The non-underlined bases in the reads, when read, mean bases that provide two possible options, and the true identity of the base at that position in the DNA template cannot be inferred from that information source. For example, in Read 1, if a "T" is read, in the original DNA template there could be a "T" or an "unmethylated C".

[0221] The ambiguity or redundancy of the bases in read 1, where the background is white, is resolved by reading read 3. For example, the second base "T" in read 1, which is not underlined, can be inferred if "T" appears in read 3. This is because in read 3, "T" represents the true "T" in the DNA template. Therefore, even if information about the second base of the DNA template is not obtained in read 1, the base sequence of the template DNA can be inferred because read 3 provides true identification of the said base.

[0222] However, if there is a sequencing error in one of the underlined bases in read 1, read 3 cannot overcome the ambiguity, and the true identification of the base is incorrectly inferred. This is the case with the error "G" highlighted in a thick black box in read 1. Read 3 shows "A", but the base "A" in read 3 could mean either the true "G" or the true "A". Therefore, if read 1 contains a base error that cannot be directly inferred in read 3, the true base identification in the DNA template will be misidentified: since the "G" in read 1 means the true "G" in the DNA template, the misidentified "G" highlighted in a thick black box is inferred to be the true "G", leading to an error in sequencing the DNA template.

[0223] These sequencing errors are eliminated by the method of the present invention, resulting in at least one, preferably at least two, more reads, namely read 2 and / or read 4. See also Figure 7B:

[0224] [Table 12]

[0225] In the table above, leads 1 and 3 are the same as those shown above. However, two new leads are provided: leads 4 and 2 are leads derived from primers 2 and 4. See Figure 7.

[0226] In this case, the presence of reads 2 and 4 indicates that read 4 shows an underlined "T," but a base mismatch is obtained, meaning that the "T" in read 4 represents the true "T" in the DNA template. Therefore, true paired-end sequencing yields two additional reads, and a sequencing error due to a base mismatch can be detected in these reads.

[0227] Therefore, the claimed method allows for a very low error rate when sequencing the molecules defined in step i of the method of the present invention, because it yields more than two reads per molecule (e.g., at least three, preferably at least four reads).

[0228] Items of the present invention This invention provides the following: 1. Follow these steps: i. A step of providing a nucleic acid molecule including a 5' region and a 3' region, The 5' and 3' regions are covalently linked by nucleotide sequences to which primers can bind. Base recognition in either the 5' or 3' region, and base recognition in the other region, independently provide information about base recognition at the corresponding locus in the original nucleic acid molecule. The molecules are as follows: - A single adapter at the 5' end of the molecule; - A single adapter at the 3' end of the molecule Further steps include; ii. A step of sequencing the molecule provided in step (i) using at least two different primers, wherein at least two different primers bind to four different regions in the nucleic acid molecule provided in (i): 1. At least one primer can at least partially bind to at least a portion of the adapter at the 5' end of the molecule to sequence at least a portion of the 5' region of the nucleic acid molecule provided in (i); 2. At least one of the primers is capable of sequencing the 3' region of the nucleic acid molecule provided in (i) by at least partially binding to a region of a nucleotide sequence that covalently links the 5' and 3' regions of the nucleic acid molecule provided in (i); 3. At least one of the primers is capable of sequencing at least a portion of the 3' region of the nucleic acid molecule provided in (i) by at least partially binding to at least a portion of an adapter at the 3' end of the molecule; and / or 4. At least one of the primers is capable of sequencing the 5' region of the nucleic acid molecule provided in (i) by at least partially binding to a region of a nucleotide sequence that covalently links the 5' and 3' regions of the nucleic acid molecule provided in (i), comprising a method.

[0229] 2. The method according to item 1, wherein the 3' region of the nucleic acid molecule provided in step (i) is at least partially complementary to the reverse strand of the 5' region.

[0230] 3. The method according to any one of items 1 or 2, wherein the nucleic acid molecule provided in step (i) is a DNA molecule.

[0231] The method according to any one of items 1 to 3, wherein the 5' region and / or the 3' region of the nucleic acid molecule provided in step (i) is a fragment of genomic DNA.

[0232] 5. The method according to any one of items 1 to 4, wherein the 5' region of the nucleic acid molecule provided in step (i) is a fragment of genomic DNA, and the 3' region of the nucleic acid molecule provided in step (i) is the complementary strand of the reverse strand of the 5' region.

[0233] 6. The method according to any one of items 1 to 4, wherein the 3' region of the nucleic acid molecule provided in step (i) is a fragment of genomic DNA, and the 5' region of the nucleic acid molecule provided in step (i) is the complementary strand of the reverse strand of the 3' region.

[0234] 7. The method according to any one of items 1 to 6, wherein the nucleotide sequence covalently linking the 5' region and the 3' region of the nucleic acid molecule provided in step (i) has a length of at least 5 nucleotides, preferably at least 10 nucleotides, and more preferably at least 17 nucleotides.

[0235] 8. The method according to any one of items 1-7, wherein in step (ii), the molecule provided in step (i) is sequenced using four different primers.

[0236] 9. The nucleic acid molecule from step (i) is used in the following steps: a. A step of providing a collection of double-stranded nucleic acid molecules, preferably a plurality of double-stranded DNA molecules being fragments of genomic DNA; b. A step of obtaining multiple adapter-containing nucleic acid molecules by ligating at least one end of the strands of multiple double-stranded nucleic acid molecules with at least partially double-stranded (ds) adapters; c. For each strand of the nucleic acid molecule obtained in step (b), a complementary strand (synthetic complementary strand) is synthesized by polymerase elongation from the 3' end of the second nucleic acid strand in the adapter molecule, using each strand of the nucleic acid molecule obtained in step (b) as a template, and thereafter each strand of the nucleic acid molecule obtained in step (b) is paired with the synthetic complementary strand to provide a plurality of adapter-modified nucleic acid molecules. The original nucleic acid strand and its synthetic complementary strand obtained in step (c) are covalently linked by a nucleotide sequence to which the primer can at least partially bind; d. A step in which complementary strands of multiple adapter-modified DNA molecules obtained in step (c) are optionally provided using primers whose sequences are complementary to at least a portion of the double-stranded adapter; e. A step of optionally amplifying the paired double-stranded nucleic acid molecules obtained in step (d) and providing the amplified paired double-stranded nucleic acid molecules. A method used to generate any one of items 1-9.

[0237] 10. The method according to item 9, wherein at least a portion of the double-stranded adapters have an arrangement common to all double-stranded adapters used in step (b).

[0238] 11. One or more methods of items 9-10, wherein, prior to step (c), multiple paired adapter-modified DNA molecules are separated to produce a library of paired adapter-modified DNA molecules.

[0239] 12. The method according to one or more of items 9-11, wherein the 3' region of the second DNA strand of the adapter forms a hairpin loop by hybridization between a first segment and a second segment within the 3' region, with the first segment located at the 3' end of the 3' region and the second segment located near the 5' region of the second DNA strand.

[0240] 13. The method is as follows, after step (c): The step of bringing each chain of the adapter-containing nucleic acid molecule into contact with the complex of the extension primer and the hairpin adapter under conditions sufficient for hybridization of the extension primer to the second chain of the adapter, wherein the extension primer is complementary to the second chain of the adapter molecule and includes a 3' region that forms an overhang end after hybridization with the second chain of the adapter molecule, and the hairpin adapter includes a hairpin loop region and an overhang end that is compatible with the overhang end formed after hybridization of the extension primer to the second chain of the adapter. The method described in one or more of items 9-12, including further.

[0241] 14. The method according to one or more of items 9 to 13, wherein the adapter has a first barcode sequence in a double-stranded region and / or a second barcode sequence in the 3' region of the second strand of the adapter.

[0242] 15. The method according to one or more of items 9-14, wherein the adapter has a restriction site in the 5' region of the first strand of the adapter.

[0243] 16. The method of one or more of items 9-15, wherein in step (b), multiple adapter-modified nucleic acid molecules are provided and strands of genomic DNA fragments are further paired using barcode sequences.

[0244] 17. The method of item 16, wherein the ligation can be performed before, after, or simultaneously with the ligation.

[0245] 18. The method according to one or more of items 9-17, wherein at least a portion of the double-stranded adapters has an sequence common to all double-stranded adapters used in step (b).

[0246] 19. The method is as follows: step (c2) after step (c): A step to convert unmethylated cytosine(s) in the paired adapter-modified nucleic acid molecule to uracil in the paired adapter-modified DNA molecule. The method described in one or more of items 9-18, including further details.

[0247] 20. The method according to item 19, wherein the presence of 5C-modified cytosine at a given position is determined by the appearance of cytosine on one strand of the paired double-stranded nucleic acid molecule obtained in step (d) or step (e) and the appearance of guanine at the corresponding position on the other strand of the paired double-stranded nucleic acid molecule, and / or the presence of unmodified cytosine at a given position is determined by the appearance of uracil or thymine on one strand of the paired double-stranded nucleic acid molecule obtained in step (d) or step (e) and the appearance of guanine at the corresponding position on the other strand of the paired double-stranded nucleic acid molecule.

[0248] 21. The nucleic acid molecule from step (i) is used in the following steps: a. A step of providing a double-stranded nucleic acid molecule, preferably the double-stranded DNA molecule being a fragment of genomic DNA; b. A step of covalently linking the forward and reverse single-stranded nucleic acid molecules provided in step a. Generated by, In step b, covalent linking is performed by a nucleotide sequence to which the primer can bind, obtaining a nucleic acid molecule containing the 5' and 3' regions. The 5' and 3' regions are covalently linked by a nucleotide sequence to which the primer can bind. The method according to any one of items 1 to 8, wherein the base identification in either the 5' or 3' region and the base identification in the other region independently provide information regarding the base identification of the corresponding gene locus in the original nucleic acid molecule.

[0249] 22. The method according to item 21, wherein the double-stranded nucleic acid molecules provided in step a are provided as a collection of double-stranded nucleic acid molecules, preferably the plurality of double-stranded DNA molecules are fragments of genomic DNA.

[0250] 23. A method of one or more of items 1 to 21, further comprising determining the true identification of a base at a specific gene locus of the original nucleic acid molecule based on the information provided in step (ii).

[0251] 24. The method of item 23, wherein if the first base identification from reads 1 and 3 and the second base identification from reads 4 and 2 do not match any of the following combinations: 1) adenine, adenine, thymine, and thymine corresponding to A; 2) thymine, thymine, adenine, and adenine corresponding to T; 3) thymine, guanine, adenine, and guanine corresponding to unmethylated C; 4) guanine, adenine, guanine, and thymine corresponding to G; and 5) cytosine, cytosine, guanine, and guanine corresponding to methylated C, then the true base identification at the locus of the original nucleic acid molecule is determined to be a misidentification.

[0252] 25. The method of one or more of items 1 to 24, further comprising using a computer, the method comprising a processor, memory, and instructions stored therein, which, when executed, identify a base at a particular position in the original nucleic acid molecule and / or an associated BQ based on the information provided in step (ii).

[0253] 26. A computer program, when executed by a computer, that includes instructions capable of determining the identification of a confirmed / predicted base and / or associated BQ at a specific location in the original nucleic acid molecule, based on the information provided in step (ii) of the method defined in any one of items 1 to 25.

[0254] 27. A kit comprising at least two different primers, preferably four different primers, wherein at least two different primers, preferably four different primers, can at least partially bind to four different regions in a nucleic acid molecule provided in step (i) of a method defined in any one of items 1 to 25: 1. At least one of the primers can at least partially bind (hybridize) to at least a portion of the adapter at the 5' end of the nucleic acid molecule provided in (i) so that at least a portion of the 5' region of the nucleic acid molecule provided in (i) can be sequenced; 2. At least one of the primers can at least partially bind (hybridize) to a region of a nucleotide sequence that covalently links the 5' and 3' regions of the nucleic acid molecule provided in (i), thereby sequencing at least a portion of the 3' region of the nucleic acid molecule provided in (i); 3. At least one of the primers can at least partially bind (hybridize) to at least a portion of the adapter at the 3' end of the nucleic acid molecule provided in (i) so that at least a portion of the 3' region of the nucleic acid molecule provided in (i) can be sequenced; and / or 4. A kit in which at least one of the primers can at least partially bind (hybridize) to a region of a nucleotide sequence that covalently links the 5' and 3' regions of the nucleic acid molecule provided in (i), thereby enabling sequencing of at least a portion of the 5' region of the nucleic acid molecule provided in (i).

[0255] 28. A kit further comprising a double-stranded adapter for use in a method defined in any one of items 9 to 24, wherein the adapter comprises a first nucleic acid strand and a second nucleic acid strand, The 3' region of the first nucleic acid strand and the 5' region of the second nucleic acid strand form a double-stranded region through sequence complementarity. The end of the double-stranded region formed by the 3' region of the first nucleic acid strand of the adapter and the 5' region of the second nucleic acid strand is compatible with the end of the double-stranded nucleic acid molecule. The adapter's double-stranded region contains one or more barcode sequences. The 3' region of the adapter's second chain forms a hairpin loop by hybridization between the first and second segments within the 3' region, with the first segment located at the 3' end of the 3' region and the second segment located near the 5' region of the second chain, and / or The adapter includes at least one barcode sequence in the single-strand region of the adapter, The barcode sequence consists of unique identifiers that enable the identification of a specific construct including the identifier and its amplification product, and The kit described in item 26, where compatibility means that the ends of the double-stranded region of the adapter molecule can be ligated to one or both ends of a double-stranded nucleic acid molecule.

[0256] 29. The kit according to item 27, wherein the adapter has a restriction site in the 5' region of the first strand of the adapter.

[0257] 30. A kit according to any one of items 27-29, wherein the adapter includes at least one barcode sequence in a single-stranded region of the adapter, and the 3' region of the second strand of the adapter forms a hairpin loop by hybridization between a first segment and a second segment within the 3' region, with the first segment located at the 3' end of the 3' region and the second segment located near the 5' region of the second strand.

[0258] 31. The following: (i) A library of double-stranded adapters, wherein the adapter comprises a first strand and a second strand, the 3' region of the first strand and the 5' region of the second strand form a double-stranded region by sequence complementarity, and the ends of the double-stranded region are compatible with the ends of a double-stranded nucleic acid molecule; (ii) Multiple extension primers, each extension primer being complementary to the second chain of the adapter molecule as defined in (i), and including a 3' region that forms an overhang end after hybridization with the second chain of the adapter molecule; and (iii) A plurality of hairpin adapters, each hairpin adapter comprising a hairpin loop region and an overhang end that is compatible with the overhang end formed after hybridization of the extension primer defined in (ii) and the second chain of the Y adapter defined in (i); A kit that further includes, (ii) the extension primer and (iii) the hairpin adapter may be provided as a composite; The kit described in any one of items 27-30 is suitable for obtaining a library of adapters for use in any one of items 9-24, the adapter of (i), the extension primer of (ii), and the hairpin adapter of (iii).

Claims

1. The following steps: i. A step of providing a nucleic acid molecule including a 5' region and a 3' region, The 5' and 3' regions are covalently linked by nucleotide sequences to which primers can bind. Base recognition in either the 5' or 3' region, and base recognition in the other region, independently provide information about base recognition at the corresponding gene locus in the original nucleic acid molecule. The molecules are as follows: - A single adapter at the 5' end of the molecule; - A single adapter at the 3' end of the molecule Further steps include; ii. A step of sequencing the molecule provided in step (i) using at least two different primers, for example, at least three different primers, preferably at least four different primers, wherein at least two different primers, for example, at least three different primers, preferably at least four different primers, bind to at least three different regions, preferably at least four different regions in the nucleic acid molecule provided in (i), 1. At least one primer can bind at least partially to at least a portion of the adapter at the 5' end of the molecule to sequence at least a portion of the 5' region of the nucleic acid molecule provided in (i); 2. At least one of the primers can at least partially bind to a region of a nucleotide sequence that covalently links the 5' and 3' regions of the nucleic acid molecule provided in (i), thereby sequencing at least a portion of the 3' region of the nucleic acid molecule provided in (i); 3. At least one primer can at least partially bind to at least a portion of the adapter at the 3' end of the molecule to sequence at least a portion of the 3' region of the nucleic acid molecule provided in (i); and / or 4. A method comprising the step of having at least one primer at least partially bind to a region of a nucleotide sequence covalently linking the 5' and 3' regions of the nucleic acid molecule provided in a., thereby enabling sequencing of at least a portion of the 5' region of the nucleic acid molecule provided in (i).

2. The method according to claim 1, wherein the 3' region of the nucleic acid molecule provided in step (i) is at least partially complementary to the reverse strand of the 5' region.

3. The method according to any one of claims 1 to 2, wherein the nucleic acid molecule provided in step (i) is a DNA molecule.

4. The method according to any one of claims 1 to 3, wherein the 5' region and / or 3' region of the nucleic acid molecule provided in step (i) is a fragment of genomic DNA.

5. The method according to any one of claims 1 to 4, wherein the 5' region of the nucleic acid molecule provided in step (i) is a fragment of genomic DNA, and the 3' region of the nucleic acid molecule provided in step (i) is the complementary strand to the reverse strand of the 5' region.

6. The method according to any one of claims 1 to 4, wherein the 3' region of the nucleic acid molecule provided in step (i) is a fragment of genomic DNA, and the 5' region of the nucleic acid molecule provided in step (i) is the complementary strand to the reverse strand of the 3' region.

7. The method according to any one of claims 1 to 6, wherein the nucleotide sequence covalently linking the 5' region and the 3' region of the nucleic acid molecule provided in step (i) has a length of at least 5 nucleotides, preferably at least 10 nucleotides, and more preferably at least 17 nucleotides.

8. The method according to any one of claims 1 to 7, wherein in step (ii), the molecule provided in step (i) is sequenced using four different primers.

9. The nucleic acid molecule in step (i) is: a. A step of providing a collection of double-stranded nucleic acid molecules, preferably a plurality of double-stranded DNA molecules being fragments of genomic DNA; b. A step of obtaining multiple adapter-containing nucleic acid molecules by ligating at least one end of the strands of multiple double-stranded nucleic acid molecules with at least partially double-stranded (ds) adapters; c. For each strand of the nucleic acid molecule obtained in step (b), a complementary strand ("synthetic complementary strand") is synthesized by polymerase elongation from the 3' end of the second nucleic acid strand in the adapter molecule, using each strand of the nucleic acid molecule obtained in step (b) as a template, and thereafter each strand of the nucleic acid molecule obtained in step (b) is paired with the synthetic complementary strand to provide a plurality of adapter-modified nucleic acid molecules. The original nucleic acid strand and its synthetic complementary strand obtained in step (c) are covalently linked by a nucleotide sequence to which the primer can at least partially bind; d. A step in which the complementary strands of the multiple adapter-modified DNA molecules obtained in step (c) are optionally provided using a primer whose sequence is complementary to at least a portion of the double-stranded adapter; e. A step of optionally amplifying the paired double-stranded nucleic acid molecules obtained in step (d) and providing the amplified paired double-stranded nucleic acid molecules. The method according to any one of claims 1 to 8, which is produced by...

10. The method according to claim 9, wherein at least a portion of the double-stranded adapters has an arrangement common to all double-stranded adapters used in step (b).

11. The method according to any one of claims 9 to 10, wherein, prior to step (c), a plurality of paired adapter-modified DNA molecules are separated to produce a library of paired adapter-modified DNA molecules.

12. The method according to any one of claims 9 to 11, wherein the 3' region of the second DNA strand of the adapter forms a hairpin loop by hybridization between a first segment and a second segment within the 3' region, the first segment being located at the 3' end of the 3' region of the second DNA strand, and the second segment being located near the 5' region of the second DNA strand.

13. After step (c), follow step (c1): The step of bringing each chain of the adapter-containing nucleic acid molecule into contact with the complex of the extension primer and the hairpin adapter under conditions sufficient for hybridization of the extension primer to the second chain of the adapter, wherein the extension primer includes a 3' region that is complementary to the second chain of the adapter molecule and forms an overhang end after hybridization with the second chain of the adapter molecule, and the hairpin adapter includes a hairpin loop region and an overhang end that is compatible with the overhang end formed after hybridization of the extension primer to the second chain of the adapter. The method according to any one of claims 9 to 12, further comprising:

14. The method according to any one of claims 9 to 13, wherein the adapter has a first barcode sequence in a double-stranded region and / or a second barcode sequence in the 3' region of the second strand of the adapter.

15. The method according to any one of claims 9 to 14, wherein the adapter has a restricting portion in the 5' region of the first chain of the adapter.

16. The method according to any one of claims 9 to 15, wherein in step (b), a plurality of adapter-modified nucleic acid molecules are provided and strands of genomic DNA fragments are further paired using barcode sequences.

17. The method according to claim 16, wherein the pairing can be performed before, after, or simultaneously with the ligation.

18. The method according to any one of claims 9 to 17, wherein at least a portion of the double-stranded adapter has an arrangement common to all double-stranded adapters used in step (b).

19. After step (c), proceed to steps c21 and / or c22: c21) A step of converting an unmodified nucleotide in the paired adapter-containing nucleic acid molecule, if present, into another nucleotide in the paired adapter-containing nucleic acid molecule that can be read separately from the said nucleotide; and / or c22) A step of converting a modified nucleotide in the paired adapter-containing nucleic acid molecule, if present, into another nucleotide in the paired adapter-containing nucleic acid molecule that can be read separately from the aforementioned nucleotide. The method according to any one of claims 9 to 18, further comprising:

20. The nucleic acid molecule in step (i) is used in the following steps: a. A step of providing a double-stranded nucleic acid molecule, preferably the double-stranded DNA molecule being a fragment of genomic DNA; and b. A step of covalently linking the forward and reverse single-stranded nucleic acid molecules provided in step a. Generated by, In step b, covalent linking is performed by a nucleotide sequence to which the primer can bind, obtaining a nucleic acid molecule containing the 5' and 3' regions. The 5' and 3' regions are covalently linked by a nucleotide sequence to which the primer can bind. The method according to any one of claims 1 to 8, wherein the base identification in one of the 5' region or the 3' region and the base identification in the other region both independently provide information regarding the base identification of the corresponding gene locus in the original nucleic acid molecule.

21. The method according to claim 20, wherein the double-stranded nucleic acid molecules provided in step a are provided as a collection of double-stranded nucleic acid molecules, preferably a plurality of double-stranded DNA molecules are fragments of genomic DNA.

22. The method according to any one of claims 1 to 21, further comprising determining the true identification of a base at a specific gene locus of the original nucleic acid molecule based on the information provided in step (ii).

23. The method according to claim 22, wherein if the identification of the first base from reads 1 and 3 and the identification of the second base from reads 4 and 2 do not match any of the following combinations: 1) adenine, adenine, thymine and thymine corresponding to A; 2) thymine, thymine, adenine and adenine corresponding to T; 3) thymine, guanine, adenine and guanine corresponding to unmethylated C; 4) guanine, adenine, guanine and thymine corresponding to G; and 5) cytosine, cytosine, guanine and guanine corresponding to methylated C, then the identification of the true base at the gene locus of the original nucleic acid molecule is determined to be a misidentification.

24. The method according to any one of claims 1 to 23, further comprising using a computer including a processor, memory, and instructions stored therein, which, when executed, identify a base at a particular position in the original nucleic acid molecule and / or an associated BQ based on the information provided in step (ii).

25. A computer program, when executed by a computer, capable of identifying a confirmed / predicted base at a specific location in the original nucleic acid molecule and / or determining the associated BQ based on the information provided in step (ii) as defined in any one of claims 1 to 24.

26. A kit comprising at least two different primers, for example, at least three different primers, preferably four different primers, wherein at least two different primers, for example, at least three different primers, preferably four different primers, can at least partially bind to at least three different regions, preferably at least four different regions, in a nucleic acid molecule provided in step (i) of a method defined in any one of claims 1 to 25:

1. At least one of the primers can at least partially bind (hybridize) to at least a portion of the adapter at the 5' end of the nucleic acid molecule provided in (i) so that at least a portion of the 5' region of the nucleic acid molecule provided in (i) can be sequenced; 2. At least one of the primers can at least partially bind (hybridize) to a region of a nucleotide sequence that covalently links the 5' and 3' regions of the nucleic acid molecule provided in (i), thereby sequencing at least a portion of the 3' region of the nucleic acid molecule provided in (i); 3. At least one of the primers can at least partially bind (hybridize) to at least a portion of the adapter at the 3' end of the nucleic acid molecule provided in (i) to sequence at least a portion of the 3' region of the nucleic acid molecule provided in (i); and / or 4. A kit in which at least one of the primers can at least partially bind (hybridize) to a region of a nucleotide sequence that covalently links the 5' and 3' regions of the nucleic acid molecule provided in (i), thereby enabling sequencing of at least a portion of the 5' region of the nucleic acid molecule provided in (i).

27. A kit further comprising a double-stranded adapter for use in a method defined in any one of claims 9 to 24, wherein the adapter comprises a first nucleic acid strand and a second nucleic acid strand, The 3' region of the first nucleic acid strand and the 5' region of the second nucleic acid strand form a double-stranded region through sequence complementarity. The end of the double-stranded region formed by the 3' region of the first nucleic acid strand of the adapter and the 5' region of the second nucleic acid strand is compatible with the end of the double-stranded nucleic acid molecule. The adapter's double-stranded region contains one or more barcode sequences. The 3' region of the adapter's second chain forms a hairpin loop by hybridization between the first and second segments within the 3' region, with the first segment located at the 3' end of the 3' region of the second chain and the second segment located near the 5' region of the second chain, and / or The adapter includes at least one barcode sequence in the single-strand region of the adapter, The barcode sequence consists of unique identifiers that enable the identification of a specific construct containing the identifier and its amplification product, and The kit according to claim 26, wherein compatibility means that the ends of the double-stranded region of the adapter molecule can be ligated to one or both ends of a double-stranded nucleic acid molecule.

28. The kit according to any one of claims 26 or 27, wherein the adapter has a restricting portion in the 5' region of the first chain of the adapter.

29. The kit according to any one of claims 26 to 28, wherein the adapter includes at least one barcode sequence in a single-strand region of the adapter, the 3' region of the second strand of the adapter forms a hairpin loop by hybridization between a first segment and a second segment within the 3' region, the first segment is located at the 3' end of the 3' region of the second strand, and the second segment is located near the 5' region of the second strand.

30. below: (i) A library of double-stranded adapters, wherein the adapter comprises a first strand and a second strand, the 3' region of the first strand and the 5' region of the second strand form a double-stranded region by sequence complementarity, and the ends of the double-stranded region are compatible with the ends of a double-stranded nucleic acid molecule; (ii) Multiple extension primers, each extension primer being complementary to the second chain of the adapter molecule as defined in (i), and including a 3' region that forms an overhang end after hybridization with the second chain of the adapter molecule; and (iii) A plurality of hairpin adapters, each hairpin adapter comprising a hairpin loop region and an overhang end that is compatible with the overhang end formed after hybridization of the extension primer defined in (ii) and the second chain of the Y adapter defined in (i), A kit that further includes, The extension primer of (ii) and the hairpin adapter of (iii) may be provided as a composite; The kit according to any one of claims 26 to 29, wherein the adapter (i), the extension primer (ii), and the hairpin adapter (iii) are suitable for obtaining a library of adapters to be used in the method defined in any one of claims 9 to 24.