Compositions and techniques for nucleic acid primer extension

By forming a stable ternary complex through specific contact between the primer-template nucleic acid hybrid and the polymerase and nucleotide mixture, the next base can be detected and distinguished from other base types in the template. This solves the problem of insufficient speed and accuracy of nucleic acid sequencing in existing technologies, and realizes rapid and accurate nucleic acid sequence feature labeling and diagnosis.

CN112074603BActive Publication Date: 2026-03-24PACIFIC BIOSCIENCES OF CALIFORNIA INC
View PDF 61 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-02-01
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing nucleic acid sequencing methods are not fast or accurate enough in clinical applications, leading to diagnostic delays and affecting patients' health and treatment options.

Method used

A method for characterizing nucleic acids is employed, which involves contacting a primer-template nucleic acid hybrid with a mixture of polymerase and nucleotides to generate an extended primer hybrid, which is then contacted with the polymerase and nucleotide mixture to form a stable ternary complex. The method detects the next base and distinguishes it from other base types in the template, thus determining the presence of base multivariates.

Benefits of technology

It improves the accuracy and speed of nucleic acid sequencing, reduces exposure to light and chemicals, increases read length and throughput, provides rapid nucleic acid sequence characterization, and supports rapid diagnostic and treatment decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112074603B_ABST
    Figure CN112074603B_ABST
Patent Text Reader

Abstract

A method of characterizing a nucleic acid, the method comprising the steps of: (a) contacting a primer-template nucleic acid hybrid with a polymerase and a mixture of nucleotides to produce an extended primer hybrid and form a stable ternary complex comprising the extended primer hybrid, the polymerase, and a nucleotide cognate of a next base in the template, wherein the mixture contains nucleotide cognates of four different base types, wherein the nucleotide cognate of a first base type has a reversible terminator, and wherein nucleotide cognates of second, third, and fourth base types are extendable; (b) detecting the stable ternary complex to distinguish the next base from other base types in the template; and (c) determining the presence of a base multiplet in the template nucleic acid, the base multiplet comprising the first base type followed by the next base.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross Reference to Related Applications

[0002] This application claims the benefit of U.S. Provisional Application No. 62 / 626,836, filed February 6, 2018, the entire contents of which are incorporated herein by reference for all purposes.

[0003] Reference to Sequence Listing, Table, or Computer Program Appendix Submitted on Compact Disc, ASCII File

[0004] The sequence list written to the file 053195-505001WO_Sequence_Listing_ST25.txt, created on February 1, 2019, is 1,865 bytes long, is on an IBM-PC machine, and runs on an MS Windows operating system. It is incorporated into this article by reference. Background Technology

[0005] This disclosure generally relates to molecular analysis and diagnostics and has specific applicability to nucleic acid sequencing.

[0006] Over the past decade, the time required to sequence the human genome has decreased dramatically. Procedures that previously took years and millions of dollars can now be completed in just a few thousand dollars and a few hours. While this rate of improvement is impressive, currently available commercial methods still fall short of meeting the needs of many clinical applications.

[0007] For many clinicians, sequencing holds promise for providing crucial information to develop reliable diagnoses of whether a patient has a life-threatening disease and to guide the selection of expensive or life-altering treatment options. For instance, sequencing can play a critical role in confirming an initial cancer diagnosis and helping patients decide on treatment options. Even a delay of a few days in receiving this confirmation can have a seriously adverse impact on a patient's emotional and psychological state.

[0008] In some cases, clinical outcomes depend heavily on rapid diagnosis. For example, neonatal intensive care units have used sequencing to identify mysterious illnesses in newborns and guide doctors toward life-saving treatment options that would otherwise have gone unrecognized. However, far too many newborns die each year due to a lack of timely diagnosis.

[0009] Therefore, there is a need to improve the accuracy, speed, and cost of nucleic acid sequencing. This invention addresses these needs and provides related advantages. Summary of the Invention

[0010] This disclosure provides a method for characterizing nucleic acids. The method may include the following steps: (a) contacting a primer-template nucleic acid hybrid with a polymerase and a nucleotide mixture under certain conditions to generate an extended primer hybrid, wherein the mixture comprises nucleotide homologs of first, second, third, and fourth different base types, wherein the nucleotide homolog of the first base type has a reversible terminator, and wherein at least one of the nucleotide homologs of the second, third, or fourth base types is extendable; (b) contacting the extended primer hybrid with the nucleotide homolog of at least one of the different base types and the polymerase to form a stable ternary complex comprising the extended primer hybrid, the polymerase, and a nucleotide homolog of the next base in the template; (c) detecting the stable ternary complex to distinguish the next base from other base types in the template; and (d) determining the presence of a base multivariate in the template nucleic acid, the base multivariate comprising the first base type followed by the next base. Optionally, the method further comprises the following steps: (e) repeating steps (a) to (c) and (f) using the extended primer hybrid as the primer-template nucleic acid hybrid to determine the presence of a series having at least two base multiple states in the template nucleic acid. In some embodiments, the nucleotide homolog of the first base type has a reversible terminator and the nucleotide homologs of the second, third, and fourth base types are extended. Alternatively, the nucleotide homologs of the first and second base types have reversible terminators and the nucleotide homologs of the third and fourth base types are extended.

[0011] In another embodiment, a method for characterizing nucleic acids may include the following steps: (a) contacting a primer-template nucleic acid hybrid with a polymerase and a nucleotide mixture under certain conditions to produce an extended primer hybrid, wherein the mixture comprises nucleotide homologs of first, second, third, and fourth different base types, wherein the nucleotide homolog of the first base type has a reversible terminator, and wherein the nucleotide homologs of the second, third, and fourth base types are extendable; (b) contacting the extended primer hybrid with at least one nucleotide homolog of the different base types and a polymerase to form a stable ternary complex comprising the extended primer hybrid, the polymerase, and a nucleotide homolog of the next base in the template; (c) detecting the stable ternary complex to determine the next base in relation to the template. (d) Differentiating other base types in the template; (e) determining the presence of a base multimorphism in the template nucleic acid, the base multimorphism comprising the first base type followed by the next base; (f) repeating steps (a) to (c) using the extended primer hybrid as the primer-template nucleic acid hybrid; (g) determining the presence of a series having at least two base multimorphisms in the template nucleic acid; and (c) repeating steps (a) to (f), wherein the nucleotide homolog of the second base type comprises a reversible terminator, and wherein the nucleotide homologs of the first, third, and fourth base types in the mixture are extended, and wherein the base multimorphism determined in step (c) comprises the second base type followed by the next base, thereby determining the presence of a series of two base multimorphisms in the template nucleic acid. Optionally, the method may further comprise (h) repeating steps (a) to (f), wherein the nucleotide homolog of the third base type contains a reversible terminator, and wherein the nucleotide homologs of the first, second, and fourth base types in the mixture are extendable, and wherein the base multistate determined in step (c) contains the third base type, followed by the next base, thereby determining the presence of a three-base multistate series in the template nucleic acid. Alternatively, the method may comprise (i) repeating steps (a) to (f), wherein the nucleotide homolog of the fourth base type contains a reversible terminator, and wherein the nucleotide homologs of the first, second, and third base types in the mixture are extendable, and wherein the base multistate determined in step (c) contains the fourth base type, followed by the next base, thereby determining the presence of a four-base multistate series in the template nucleic acid.

[0012] In another embodiment, a method for characterizing nucleic acids may include the following steps: (a) (i) performing a sequencing process, the sequencing process including polymerase extension of primers hybridized to a template to generate a primer-template nucleic acid hybrid and detecting a signal indicating a sequence located in a first region of the template at the 3' position of the series having at least two base multiples; and (ii) contacting the primer-template nucleic acid hybrid with a polymerase and a nucleotide mixture under certain conditions to generate an extended primer hybrid, wherein the mixture comprises nucleotide homologs of first, second, third, and fourth different base types, wherein the nucleotide homolog of the first base type has a reversible terminator, and wherein the nucleotide homologs of the second, third, and fourth base types are extendable; and (b) contacting the extended primer hybrid with a polymerase and a template. (a) A nucleotide homolog of at least one base type from different base types is contacted with a polymerase to form a stable ternary complex comprising the extended primer hybrid, the polymerase, and a nucleotide homolog of the next base in the template; (b) The stable ternary complex is detected to distinguish the next base from other base types in the template; (c) The presence of a base multivariate in the template nucleic acid is determined, the base multivariate comprising the first base type followed by the next base; (e) Steps (a) to (c) and (f) are repeated using the extended primer hybrid as the primer-template nucleic acid hybrid to determine the presence of a sequence having at least two base multivariates in the template nucleic acid, wherein the first region of the template is located at the 3' position of the sequence having at least two base multivariates. Optionally, the method further comprises the step of: (g) Polymerase extension of the extended primer during a second sequencing process, wherein the second sequencing process detects a signal indicating a sequence of the template located at the 5' position of the sequence having at least two base multivariates.

[0013] This disclosure provides a method for characterizing nucleic acids, the method comprising the steps of: (a) contacting a primer-template nucleic acid hybrid with a polymerase and a nucleotide mixture under certain conditions to generate an extended primer hybrid and form a stable ternary complex comprising the extended primer hybrid, the polymerase, and a nucleotide homolog of the next base in the template, wherein the mixture comprises nucleotide homologs of first, second, third, and fourth different base types, wherein the nucleotide homolog of the first base type includes a reversible terminator, and wherein at least one of the nucleotide homologs of the second, third, or fourth base types is extended; (b) detecting the stable ternary complex to distinguish the next base from other base types in the template; and (c) determining the presence of a base multivariate in the template nucleic acid, the base multivariate comprising the first base type followed by the next base. Optionally, the method comprises the steps of: (d) repeating steps (a) and (b) using the extended primer hybrid as the primer-template nucleic acid hybrid; and (e) determining the presence of a series having at least two base multiple states in the template nucleic acid. In some embodiments, the nucleotide homolog of the first base type has a reversible terminator and the nucleotide homologs of the second, third, and fourth base types are extendable. Alternatively, the nucleotide homologs of the first and second base types have reversible terminators and the nucleotide homologs of the third and fourth base types are extendable.

[0014] In another embodiment, a method for characterizing nucleic acids may include the following steps: (a) contacting a primer-template nucleic acid hybrid with a polymerase and a nucleotide mixture under certain conditions to produce an extended primer hybrid and form a stable ternary complex comprising the extended primer hybrid, the polymerase, and a nucleotide homolog of the next base in the template, wherein the mixture comprises nucleotide homologs of first, second, third, and fourth different base types, wherein the nucleotide homolog of the first base type includes a reversible terminator, and wherein the nucleotide homologs of the second, third, and fourth base types are extended; (b) detecting the stable ternary complex to distinguish the next base from other base types in the template; and (c) determining base multiplicity. The presence of the base multiplex in the template nucleic acid, wherein the base multiplex comprises the first base type followed by the next base; (d) repeating steps (a) and (b) using the extended primer hybrid as the primer-template nucleic acid hybrid; (e) determining the presence of a series having at least two base multiplexes in the template nucleic acid; and (f) repeating steps (a) to (e), wherein the nucleotide homolog of the second base type has a reversible terminator, and wherein the nucleotide homologs of the first, third, and fourth base types in the mixture are extended, and wherein the base multiplex determined in step (c) comprises the second base type followed by the next base, thereby determining the presence of a two-base multiplex series in the template nucleic acid. Optionally, the method may further comprise the following steps: (g) repeating steps (a) to (e), wherein the nucleotide homolog of the third base type has a reversible terminator, and wherein the nucleotide homologs of the first, second, and fourth base types in the mixture are extendable, and wherein the base multistate determined in step (c) comprises the third base type, followed by the next base, thereby determining the presence of a three-base multistate series in the template nucleic acid. Alternatively, the method comprises the following step (h) repeating steps (a) to (e), wherein the nucleotide homolog of the fourth base type has a reversible terminator, and wherein the nucleotide homologs of the first, second, and third base types in the mixture are extendable, and wherein the base multistate determined in step (c) comprises the fourth base type, followed by the next base, thereby determining the presence of a four-base multistate series in the template nucleic acid.

[0015] In another embodiment, a method for characterizing nucleic acids may include the following steps: (a) (i) performing a sequencing process, the sequencing process including polymerase extension of primers hybridized to a template to generate a primer-template nucleic acid hybrid and detecting a signal indicating a sequence located in a first region of the template at the 3' position of the series having at least two base multiple states; and (ii) contacting the primer-template nucleic acid hybrid with a polymerase and a nucleotide mixture under certain conditions to generate an extended primer hybrid and form a stable ternary complex comprising the extended primer hybrid, the polymerase, and a nucleotide homolog of the next base in the template, wherein the mixture comprises nucleotide homologs of first, second, third, and fourth different base types, wherein the The method includes: (a) a nucleotide homolog of the first base type comprising a reversible terminator, and wherein the nucleotide homologs of the second, third, and fourth base types are extendable; (b) detecting the stable ternary complex to distinguish the next base from other base types in the template; (c) determining the presence of a base multivariate in the template nucleic acid, the base multivariate comprising the first base type followed by the next base; and (d) repeating steps (a) and (b) using the extended primer hybrid as the primer-template nucleic acid hybrid; and (e) determining the presence of a sequence having at least two base multivariates in the template nucleic acid, wherein the first region of the template is located at the 3' position of the sequence having at least two base multivariates. Optionally, the method may include the step of: (f) polymerase extending the extended primer during a second sequencing process, wherein the second sequencing process detects a signal indicating a sequence of the template located at the 5' position of the sequence having at least two base multivariates.

[0016] This disclosure provides a method for characterizing nucleic acids, the method comprising the following steps: (a) contacting a primer-template nucleic acid hybrid with a polymerase and a nucleotide mixture under certain conditions to produce an extended primer hybrid, wherein the mixture contains nucleotide homologs of no more than three of four different base types; (b) further extending the extended primer hybrid with a nucleotide homolog of the fourth of the four different base types in the absence of the nucleotide homologs in (a), thereby producing a further extended primer hybrid; (c) forming a stable ternary complex comprising the further extended primer hybrid, the polymerase, and a nucleotide homolog of the next base in the template; (d) detecting the stable ternary complex to distinguish the next base from other base types in the template; and (e) determining the presence of a base multivariate in the template nucleic acid, the base multivariate comprising the fourth of the four different base types, followed by the next base. Optionally, the method further comprises the following steps: (f) repeating steps (a) to (d) and (g) using the further extended primer hybrid as the primer-template nucleic acid hybrid to determine the presence of a series having at least two base multiple states in the template nucleic acid.

[0017] In another embodiment, a method for characterizing nucleic acids may include the following steps: (a) contacting a primer-template nucleic acid hybrid with a polymerase and a nucleotide mixture under certain conditions to produce an extended primer hybrid, wherein the mixture contains nucleotide homologs of no more than three of four different base types; (b) further extending the extended primer hybrid with a nucleotide homolog of the fourth of the four different base types in the absence of the nucleotide homologs in (a), thereby producing a further extended primer hybrid; (c) forming a stable ternary complex comprising the further extended primer hybrid, the polymerase, and a nucleotide homolog of the next base in the template; and (d) detecting the stable ternary complex to identify the next base. (e) Distinguishing the base from other base types in the template; (f) Determining the presence of a base multimorphism in the template nucleic acid, the base multimorphism comprising the fourth base type of the four different base types, followed by the next base; (g) Repeating steps (a) to (d) using the further extended primer hybrid as the primer-template nucleic acid hybrid; (h) Determining the presence of a series having at least two base multimorphisms in the template nucleic acid; and (h) Repeating steps (a) to (g), wherein the nucleotide homolog of the fourth base type is replaced by a nucleotide homolog of the first base type and wherein the nucleotide homolog of the first base type is not used in step (a), thereby determining the presence of a two-base multimorphism series in the template nucleic acid. Optionally, the method may further comprise (i) Repeating steps (a) to (g), wherein the nucleotide homolog of the fourth base type is replaced by a nucleotide homolog of the second base type and wherein the nucleotide homolog of the second base type is not used in step (a), thereby determining the presence of a three-base multimorphism series in the template nucleic acid. As an alternative, the method may include (j) repeating steps (a) to (g), wherein the nucleotide homolog of the fourth base type is replaced by a nucleotide homolog of the third base type and wherein the nucleotide homolog of the third base type is not used in step (a), thereby determining the presence of a four-base multivariate series in the template nucleic acid.

[0018] In another embodiment, a method for characterizing nucleic acids may include the following steps: (a) (i) performing a sequencing process, the sequencing process including polymerase extension of primers hybridized to a template to generate a primer-template nucleic acid hybrid and detecting a signal indicating a sequence located in a first region of the template at the 3' position of the series having at least two base multiples; and (ii) contacting the primer-template nucleic acid hybrid with a mixture of polymerase and nucleotides under certain conditions to generate an extended primer hybrid, wherein the mixture contains nucleotide homologs of no more than three of four different base types; and (b) further extending the extended primer hybrid with a nucleotide homolog of the fourth of the four different base types in the absence of the nucleotide homologs in (a), thereby generating a sequence of primers hybridized to a template. (a) extending the primer hybrid; (c) forming a stable ternary complex comprising the further extended primer hybrid, polymerase, and a nucleotide homolog of the next base in the template; (d) detecting the stable ternary complex to distinguish the next base from other base types in the template; (e) determining the presence of a base multivariate in the template nucleic acid, the base multivariate comprising the fourth base type of the four different base types, followed by the next base; (f) repeating steps (a) to (d) using the extended primer hybrid as the primer-template nucleic acid hybrid; and (g) determining the presence of a sequence having at least two base multivariates in the template nucleic acid, wherein the first region of the template is located at the 3' position of the sequence having at least two base multivariates. Optionally, the method may further comprise the step of: (h) polymerase extension of the extended primer during a second sequencing process, wherein the second sequencing process detects a signal indicating a sequence of the template located at the 5' position of the sequence having at least two base multivariates.

[0019] Details of one or more embodiments are set forth in the following drawings and description. Other features, objectives, and advantages will become apparent from the description, the drawings, and the claims. Attached Figure Description

[0020] Figure 1 The following capillary electrophoresis traces are shown: non-extended primers from the primer-template hybrid of Tough-22 (SEQ ID NO:1). Figure 1 A) Ten SBBs with four reversible terminator nucleotides were used during the extension step. TM Primers after the cycle ( Figure 1 B); During the extension step, 10 SBBs of dT-aminoknoxy, reversible terminator nucleotide, and 3 natural nucleotides are used. TM Primers after the cycle ( Figure 1C); and 10 SBBs consisting of 4 natural nucleotides during the extension step. TM Primers after the cycle ( Figure 1 D).

[0021] Figure 2 The following capillary electrophoresis traces are shown: non-extended primers from the primer-template hybrid of Tough-29 (SEQ ID NO:2). Figure 2 A) Ten SBBs with four reversible terminator nucleotides were used during the extension step. TM Primers after the cycle ( Figure 2 B); During the extension step, 10 SBBs of dT-aminoknoxy, reversible terminator nucleotide, and 3 natural nucleotides are used. TM Primers after the cycle ( Figure 2 C); and 10 SBBs consisting of 4 natural nucleotides during the extension step. TM Primers after the cycle ( Figure 2 D). Detailed Implementation

[0022] Many nucleic acid sequencing technologies (such as Binding) TM (SBB TM Fluorescence-based sequencing technologies utilize light radiation during detection. An undesirable consequence of light-based detection is that repeated or prolonged periods of light radiation can introduce byproducts or artifacts that interfere with the reaction. For example, photooxidation may prevent nucleic acids from proceeding in the sequencing reaction due to cross-linking with other components or through fragmentation. The cumulative effects of radiation can adversely limit sequencing read length and throughput.

[0023] This disclosure provides a method for increasing read length by reducing the number of reaction cycles in which nucleic acids are exposed. Fewer cycles mean less exposure of nucleic acids to light and potentially damaging chemicals. Compared to other sequencing methods, SBB... TM The technology has the unique property of separating primer extension and nucleotide detection into two steps. Typically, SBB... TM The technique incrementally detects each position in the template because a nucleotide is added to the primer during the extension step of each cycle, and then the next correct nucleotide is detected before the next cycle begins. In this case, the number of cycles equals the length of the read.

[0024] In the specific configuration of the method described herein, the primer extension step is performed such that the number of nucleotides added to the primer is not necessarily known or determined; however, the type of nucleotide at the 3' end of the primer or the type of nucleotide adjacent to the 3' end of the primer is known or can be determined. The nucleotide type can be determined based on the composition of the nucleotide mixture used during extension, without needing to detect the extension product. For example, the nucleotide mixture may contain four different base types, where the first base type has a reversible terminator and the other three nucleotide types are extendable. In this case, the extension product will have a nucleotide of the first base type at its 3' end. In another example, the nucleotide mixture may contain only three of the four different base types that are expected to have complementary bases in the template. The extension product will not extend beyond the template position of the omitted base type. In both examples, the length of the extension product produced from multiple cycles will be variable and, in many cases, unknown. However, the number of nucleotides of a specific type terminating in the extension product can be determined based on the number of cycles performed. Nucleotide counting can provide: useful information about the composition of nucleic acid templates; and information that can be used, for example, as a characterizing marker of the template or in combination with other information to determine the sequence of the template.

[0025] In a specific configuration of this method, the extension step of the sequencing procedure is modified such that a variable number of nucleotides are added during each cycle and then the next correct nucleotide is detected. As explained below, extension can be performed in a manner where the type of nucleotide at the 3' end of the primer is known. The end result is that detecting the next correct nucleotide during each cycle allows the identification of dinucleotides containing the nucleotide at the 3' end of the primer and the detected nucleotide. Repeated steps identify dinucleotide sequences where the number and type of nucleotides in the template that separate each dinucleotide are not necessarily known. Because some extension steps will contain more than one nucleotide, the dinucleotide sequence can span a template length that is relatively long compared to the number of cycles performed. The dinucleotide sequence provides a low-resolution characteristic marker sequence of the template.

[0026] Dinucleotides are a type of base multiplex that can be identified and used in the methods of this disclosure. Longer base multiplexes of a particular template nucleic acid can be identified by following the variable-length extension steps described above, which involve more than one cycle of checking and single-nucleotide extension. For example, a trinucleotide can be identified by first performing a variable-length extension step with the known type of nucleotide at the 3' end of the extended primer (i.e., the first nucleotide in the trinucleotide), followed by a cycle involving checking the next correct nucleotide (i.e., the second nucleotide in the trinucleotide) and extending the primer by single nucleotides, and then again performing a checking step for the next correct nucleotide in the extended primer (i.e., the third nucleotide in the trinucleotide). Thus, base multiplexes of any length can form the basis for characteristic markers characterizing template nucleic acids.

[0027] In certain embodiments, higher-resolution signature sequences can be obtained by combining low-resolution signatures obtained from the same template nucleic acid. Specifically, extended primers generated in previous sequencing runs can be removed from the template and variable extension SBBs can be repeated. TM Technology. However, in repeated runs, the variable extension step will be configured to cap the primers with a different type of nucleotide than that used in the previous sequencing run. As further elaborated below, three or four base multimorphic series (each generated with primer-capped nucleotides of a different type) can be combined to obtain the sequence of the template determined at single-base resolution.

[0028] In certain embodiments, the method of this disclosure employs an extension step using a mixture of nucleotides that will stop at a specific type of base in the template. For example, the mixture used for extension may contain extendable nucleotides that pair with three base types in the template and reversibly terminated nucleotides that pair with a fourth base type in the template. The extension step will produce an extension primer having a homolog of the fourth base type at its 3' end. Thus, the identity of the nucleotide at the 3' end of the extension primer can be deduced from the nucleotide composition of the extension mixture. The next correct nucleotide in the template can be detected by forming a stable ternary complex at the end of the extension primer. The extension primer can then be deblocked and cycled repeatedly to determine the dinucleotide series present in the template.

[0029] Alternatively, the nucleotide mixture used in the extension step may contain homologous nucleotides of only three base types from the template (i.e., omitting the fourth base type homologous nucleotide from the template). The type of nucleotide at the 3' end of the extended primer does not necessarily have to be known. However, the identity of the next correct nucleotide can be deduced from the nucleotide composition of the extension mixture. In this example, the next template base is of the fourth base type. Subsequent extension using only homologous nucleotides of the fourth base type will produce an extension primer known to have a homolog with the fourth base type at its 3' end. The next correct nucleotide or the next template base can be detected by forming a stable ternary complex at the end of the extended primer. Termination is not required for the extended primer. Instead, stability can be provided to the ternary complex by the absence of the catalytic metal required for the polymerase used for extension, the presence of a polymerase inhibitor (such as a non-catalytic metal), or the use of a polymerase variant that forms the ternary complex but cannot catalyze primer extension. Alternatively, the homologous nucleotide at the 3' end of the primer may contain a reversible terminator that provides stability to the ternary complex. In this case, the reversibly terminated primer can then be deblocked to extend the dinucleotide to a longer multistate and / or allow cyclic repeating.

[0030] Specific embodiments of the methods described herein identify one or more base multiplexes that can be used as characteristic markers to identify nucleic acid sequences. For example, a series of characteristic dinucleotides of a target nucleic acid can be identified, and the characteristic markers can be compared with characteristic markers of known nucleic acid sequences to confirm or refute that the target nucleic acid has a sequence identical to that of a known nucleic acid. An advantage of such characteristic markers is that they can be obtained much faster than the full sequence, thus allowing for rapid identification of target nucleic acids. Another advantage is that the length of the genome spanned by the characteristic marker will typically be longer than that spanned by a standard run, and the information conferred by the characteristic marker can provide long-range genomic structural information that is not provided by shorter reads generated from the same number of cycles from a standard run.

[0031] In some embodiments, a single-base resolution sequence of a first region of the target nucleic acid can be determined and a signature of a second region of the target nucleic acid can be obtained. The signature can be, for example, a series of base multiples or a count of specific types of nucleotides. The target can then be aligned to a longer sequence (e.g., a genomic fragment can be aligned to a genome) by aligning not only the single-base resolution region of the aligned sequence but also the signature region. The added information from the signature region provides additional information for the aligned read. The signature can be, for example, obtained by comparing previously sequenced sequences with standard SBBs. TM Technically extended primers perform variable extension SBB TM The technology (or vice versa) is appended to the end of the high-resolution read segment. Similarly, variable-length extended SBBs can be performed. TMTechniques are used to generate signature linkers between two single-base resolution reads. The resulting sequences can be aligned as “paired reads,” where the linker signatures can be used in conjunction with the single-base resolution reads for alignment with a reference sequence. Linker signatures can add alternative or additional information to standard paired read alignment techniques (also known in the art as “paired-end” alignment techniques), in which only the distance between the linkers is considered in the alignment determination. Specifically, paired reads can be aligned to a reference genome not only based on the proximity of the two reads in the reference but also based on a comparison of the observed signature region with the reference region where the intercalation aligns with the two reads. For example, the count of a specific type of nucleotide in the observed signature can be compared with the count in the relevant region of the reference sequence, or the series of base multiples in the observed signature can be compared with the series in the relevant region of the reference sequence.

[0032] In another paired read configuration, the number of nucleotides in the adapter region can be determined empirically and used as a characteristic marker of the adapter region. More specifically, the first region of the template nucleic acid can be sequenced by using single-base resolution technology to extend the primer along the template, then the extended primer can be further extended along the second region of the template using "dark" extension technology, and then the further extended primer can be extended even further along the third region of the template using single-base resolution technology. Dark extension can be achieved by repeating cycles of primer extension, where each cycle does not contain a detection step that would otherwise be used to observe the primer extension product. For example, each dark extension cycle can employ the method described herein for sequencing SBB-based sequences. TM The technique implements one or more steps of primer extension, but each cycle may omit detection steps set forth herein or otherwise known in the art. Each dark extension cycle may use nucleotide homologs of all base types expected in the template (e.g., homologs of all four base types expected in natural genomic DNA). Since they will not be detected, nucleotide homologs do not need to be labeled. However, nucleotide homologs may contain reversible terminators, in which case each cycle may include a deblocking step. In this exemplary configuration, the number of dark extension cycles performed will be related to the length of the second region of the template (i.e., the number of bases in the second region), and this empirically determined length will, in turn, provide a characterization marker of the linker region between the first and third regions sequenced by a single-base resolution technique. The single-base resolution techniques for the first and third regions may be the same or different. Optionally, one or both of the single-base resolution techniques may be SBB. TM technology.

[0033] Unless otherwise stated, the terms used herein should be understood to have their common meaning in the relevant fields. The following explains several terms used herein and their meanings.

[0034] As used herein, when referring to two nucleotides in a nucleic acid molecule or sequence, the term "adjacent" means one of the nucleotides immediately following other nucleotides in the molecule or sequence. Thus, adjacent nucleotides in a nucleic acid molecule are covalently linked to each other. In contrast, two nucleotides that are close to each other may optionally be separated by one or more intercalation nucleotides in the nucleic acid molecule or sequence.

[0035] As used herein, when referring to nucleic acid sequences, the term "alignment" means comparing two or more sequences to identify regions of similarity. Aligned sequences can be represented as individual rows in a matrix. A particular sequence (or row in a matrix) may have one or more gaps when aligned with another sequence (or row in a matrix). For example, gaps can represent one or more positions in an unknown, ambiguous, or non-existent gapped sequence. For instance, when a first sequence is aligned with a reference sequence, the first sequence can be represented as a series of base multiplexes in which gaps can separate two or more base multiplexes. In this example, when the first sequence is aligned with the reference sequence, each gap can separate one base multiplexe from another.

[0036] As used herein, the term "array" refers to a group of molecules attached to one or more solid substrates such that molecules at one feature are distinguished from molecules at other features. An array may comprise different molecules located at different addressable features on a solid substrate. Alternatively, an array may comprise separate solid substrates acting as features with different molecules, wherein the different molecules can be identified by the location of the solid substrate on the surface to which it is attached or by the location of the solid substrate in a liquid (such as a fluid flow). The molecules in an array may be, for example, nucleotides, nucleic acid primers, nucleic acid templates, or nucleases (such as polymerases, ligases, exonucleases, or combinations thereof).

[0037] As used herein, the term "base multistate" refers to at least two adjacent bases or at least two nucleotides attached to each other by a 5' to 3' phosphodiester bond in a nucleic acid sequence. Typically, the nucleotides in a base multistate are listed in the 5' to 3' direction (e.g., 5'GT3' is often referred to as a GT dinucleotide unless otherwise specified). Exemplary base multistates include, for example, a dinucleotide containing two adjacent bases or nucleotides; a trinucleotide containing three adjacent bases or nucleotides; a tetranucleotide containing four adjacent nucleotides or bases, etc. Each base in a multistate can be explicitly identified as a specific base type (e.g., A, C, T, or G). Alternatively, the bases in the multistate can be degenerate, for example identified as R (i.e., A or G), M (i.e., A or C), W (i.e., A or T), S (i.e., C or G), Y (i.e., C or T), K (i.e., G or T), B (i.e., C or G or T), D (i.e., A or G or T), H (i.e., A or C or T), V (i.e., A or C or G) or N (i.e., A or C or G or T).

[0038] As used herein, when referring to nucleotides, the term "closing portion" means a portion of a nucleotide that inhibits or prevents the 3' oxygen of the nucleotide from forming a covalent bond to the next correct nucleotide during nucleic acid polymerization. The closing portion of a "reversible terminator" nucleotide may be removed from a nucleotide analog or otherwise modified to allow the 3' oxygen of the nucleotide to covalently attach to the next correct nucleotide. Such a closing portion is referred to herein as a "reversible terminator portion." Exemplary reversible terminator portions are described in U.S. Patent Nos. 7,427,673; 7,414,116; 7,057,026; 7,544,794 or 8,034,923; or PCT Publications WO 91 / 06678 or WO 07 / 123744, each of which is incorporated herein by reference. Nucleotides having a closing portion or a reversible terminator portion may be located at the 3' end of a nucleic acid (such as a primer) or may be monomers without covalent attachment to the nucleic acid.

[0039] The term "includes" is intended to be open-ended in this document, encompassing not only the elements described but also any additional elements.

[0040] As used herein, when referring to a set of items, the term "each" is intended to identify an individual item in the set, but not necessarily every single item in the set. Exceptions may occur if explicitly stated otherwise by public notice or context.

[0041] As used herein, when referring to parts of a molecule, the term "exogenous" means a chemical part that is not present in the natural analogue of the molecule. For example, an exogenous label on a nucleotide is a label that is not present on a naturally occurring nucleotide. Similarly, an exogenous label present on a polymerase is not present on the polymerase in its natural environment.

[0042] As used herein, when referring to nucleic acids, “extension” means the process of adding at least one nucleotide to the 3' end of a nucleic acid. When referring to nucleic acids, “polymerase extension” refers to the polymerase-catalyzed process of adding at least one nucleotide to the 3' end of a nucleic acid. The nucleotide or oligonucleotide added to a nucleic acid through extension is referred to as incorporation into the nucleic acid. Therefore, the term “incorporation” can be used to refer to the process of attaching a nucleotide or oligonucleotide to the 3' end of a nucleic acid by forming a phosphodiester bond.

[0043] As used herein, when referring to nucleotides, the term "extendable" means a nucleotide that has an oxygen or hydroxyl moiety at the 3' position and is capable of forming a covalent bond to the next correct nucleotide if and when incorporated into a nucleic acid. An extendable nucleotide can be at the 3' position of a primer or it can be a monomeric nucleotide. An extendable nucleotide will lack a closing moiety, such as a reversible terminator moiety.

[0044] As used herein, the term "extended primer hybrid" refers to a primer-template nucleic acid hybrid formed after the incorporation of at least one nucleotide into the primer. Incorporation events can be, for example, the catalytic addition of one or more nucleotide polymerases to the 3' end of the primer.

[0045] As used herein, when referring to arrays, the term "feature" means the location of a particular molecule within an array. A feature may contain only a single molecule or may contain a group of several molecules of the same species (i.e., all molecules). Alternatively, a feature may contain a group of molecules of different species (e.g., a group of ternary complexes with different template sequences). Array features are typically discrete. Discrete features may be continuous or they may have spatial spacing from each other. Arrays as used herein may have features spaced apart, for example, less than 100 micrometers, 50 micrometers, 10 micrometers, 5 micrometers, 1 micrometer, or 0.5 micrometers. Alternatively or additionally, arrays may have features spaced apart greater than 0.5 micrometers, 1 micrometer, 5 micrometers, 10 micrometers, 50 micrometers, or 100 micrometers. Features may have areas of less than 1 square millimeter, 500 square micrometers, 100 square micrometers, 25 square micrometers, 1 square micrometer, or less, respectively.

[0046] As used herein, the term "label" refers to a molecule or part thereof that provides a detectable property. For example, a detectable property can be an optical signal such as radiation absorptivity, fluorescence emission, luminescence emission, fluorescence lifetime, fluorescence polarization, etc.; Rayleigh and / or Mie scattering; binding affinity for a ligand or acceptor; magnetic properties; electrical properties; charge; mass; radioactivity, etc. Exemplary labels include, but are not limited to, fluorophores, luminophores, chromophores, nanoparticles (e.g., gold, silver, carbon nanotubes), heavy atoms, radioactive isotopes, mass labels, charge labels, spin labels, acceptors, ligands, etc.

[0047] As used herein, the term "next correct nucleotide" refers to the type of nucleotide that will bind to and / or be incorporated at the 3' end of the primer to complement the bases in the template strand that the primer hybridizes to. The base in the template strand is referred to as the "next base" and immediately follows the 5' of the base in the template that hybridizes to the 3' end of the primer. The next correct nucleotide can be referred to as a "homolog" of the next base and vice versa. Homologous nucleotides that interact with each other in a ternary complex or in a double-stranded nucleic acid are referred to as "paired" with each other. Nucleotides with bases that are not complementary to the next template base are referred to as "incorrect," "mismatched," or "non-homologous" nucleotides.

[0048] As used herein, the term "nucleotide" may be used to refer to natural nucleotides or their analogues. Examples include, but are not limited to, nucleotide triphosphates (NTPs) (such as ribonucleoside triphosphates (rNTPs), deoxyribonucleoside triphosphates (dNTPs)) or their natural analogues (such as dideoxyribonucleoside triphosphates (ddNTPs) or reversibly terminated nucleotide triphosphates (rtNTPs)).

[0049] As used herein, the term "polymerase" can refer to a nucleic acid synthase, including but not limited to DNA polymerase, RNA polymerase, reverse transcriptase, primase, and transferase. Typically, a polymerase has one or more active sites where the catalysis of nucleotide binding and / or nucleotide polymerization can occur. A polymerase can catalyze the polymerization of a nucleotide to the 3' end of the first strand of a double-stranded nucleic acid molecule. For example, a polymerase catalyzes the addition of the next correct nucleotide to the 3' oxygen group of the first strand of a double-stranded nucleic acid molecule via a phosphodiester bond, thereby covalently incorporating a nucleotide into the first strand of the double-stranded nucleic acid molecule. Optionally, the polymerase does not need to be able to perform nucleotide incorporation under one or more of the conditions used in the methods described herein. For example, a mutant polymerase may be able to form a ternary complex but cannot catalyze nucleotide incorporation.

[0050] As used herein, the term "primer-template hybrid" or "primer-template hybrid" refers to a nucleic acid hybrid having a double-stranded region such that one of the strands has a 3' end that can be extended by polymerase. The two strands can be part of a continuous nucleic acid molecule (e.g., a hairpin structure) or the two strands can be separate molecules that are not covalently attached to each other.

[0051] As used herein, the term "primer" refers to a nucleic acid having a sequence that binds to a nucleic acid at or near a template sequence. Typically, primers bind in a configuration that allows replication of the template, for example, by extension of a primer polymerase. A primer can be a first part of a nucleic acid molecule that binds to a second part of the nucleic acid molecule, the first part being the primer sequence and the second part being the primer-binding sequence (e.g., a hairpin primer). Alternatively, a primer can be a first nucleic acid molecule that binds to a second nucleic acid molecule having a template sequence. Primers can consist of DNA, RNA, or analogues thereof.

[0052] As used herein, the term "base multistate series" refers to the representation of the relative order of specific base multistates in a sequence. Base multistates in a series may be identical or different from each other. For example, dinucleotides in a series may differ from each other due to variability in the type of nucleotide present at the 3' position, the 5' position, or both. In some embodiments, dinucleotides in a series may differ from each other due to variability in the type of nucleotide present at the 3' position, while the type of nucleotide at the 5' position is uniform. Alternatively, dinucleotides in a series may have a uniform type of nucleotide at the 3' position while the type of nucleotide at the 5' position is variable. Similarly, a longer base multistate may be uniform at one position and variable at others. One of the positions in a multistate may be reduced to two or more nucleotide types. Taking IUPAC symbols as an example, one of the positions can be identified as R (i.e., A or G), M (i.e., A or C), W (i.e., A or T), S (i.e., C or G), Y (i.e., C or T), K (i.e., G or T), B (i.e., C, G, or T), D (i.e., A, G, or T), H (i.e., A, C, or G), V (i.e., A, C, G, or T). The base multiple states in the series do not need to be consecutive when aligned with the reference sequence. Therefore, two base multiple states in the series can be separated by gaps when aligned with the reference sequence.

[0053] As used herein, the term "ternary complex" refers to the intermolecular association between a polymerase, a double-stranded nucleic acid, and a nucleotide. Typically, the polymerase facilitates the interaction between the next correct nucleotide and the template strand of the initiating nucleic acid. The next correct nucleotide can interact with the template strand via Watson-Crick hydrogen bonding. The term "stable ternary complex" refers to a ternary complex with enhanced or prolonged presence or whose disruption has been inhibited. Typically, the stability of a ternary complex prevents the covalent incorporation of its nucleotide components into the initiating nucleic acid component.

[0054] As used herein, the term "type" is used to identify molecules that share the same chemical structure. For example, a mixture of nucleotides may contain several dCTP molecules. dCTP molecules should be understood as being of the same type as each other, but different types compared to dATP, dGTP, dTTP, etc. Similarly, individual DNA molecules with the same nucleotide sequence are of the same type, while DNA molecules with different sequences are of different types. The term "type" can also identify parts that share the same chemical structure. For example, cytosine bases in a template nucleic acid should be understood as bases of the same type as each other, independent of their position in the template sequence.

[0055] The embodiments described below and those recounted in the claims can be understood in accordance with the definitions above.

[0056] This disclosure provides a method for characterizing nucleic acids. The method may include the following steps: (a) contacting a primer-template nucleic acid hybrid with a polymerase and a nucleotide mixture under certain conditions to generate an extended primer hybrid, wherein the mixture comprises nucleotide homologs of first, second, third, and fourth different base types, wherein the nucleotide homolog of the first base type has a reversible terminator, and wherein at least one of the nucleotide homologs of the second, third, or fourth base types is extendable; (b) contacting the extended primer hybrid with the nucleotide homolog of at least one of the different base types and the polymerase to form a stable ternary complex comprising the extended primer hybrid, the polymerase, and a nucleotide homolog of the next base in the template; (c) detecting the stable ternary complex to distinguish the next base from other base types in the template; and (d) determining the presence of a base multivariate in the template nucleic acid, the base multivariate comprising the first base type followed by the next base. Optionally, the method further comprises the steps of: (e) repeating steps (a) to (c) and (f) using the extended primer hybrid as the primer-template nucleic acid hybrid to determine the presence of a series having at least two base multiple states in the template nucleic acid. In an embodiment, prior to step (e), the method further comprises removing a reversible terminator from the extended primer.

[0057] A method for characterizing nucleic acids is also provided, the method comprising the steps of: (a) contacting a primer-template nucleic acid hybrid with a polymerase and a nucleotide mixture under certain conditions to produce an extended primer hybrid and form a stable ternary complex comprising the extended primer hybrid, the polymerase and a nucleotide homolog of the next base in the template, wherein the mixture comprises nucleotide homologs of first, second, third and fourth different base types, wherein the nucleotide homolog of the first base type includes a reversible terminator, and wherein at least one of the nucleotide homologs of the second, third or fourth base types is extended; (b) detecting the stable ternary complex to distinguish the next base from other base types in the template; and (c) determining the presence of a base multivariate in the template nucleic acid, the base multivariate comprising the first base type followed by the next base. Optionally, the method comprises the following steps: (d) repeating steps (a) and (b) using the extended primer hybrid as the primer-template nucleic acid hybrid; and (e) determining the presence of a series having at least two base multiple states in the template nucleic acid.

[0058] This disclosure further provides a method for characterizing nucleic acids, the method comprising the steps of: (a) contacting a primer-template nucleic acid hybrid with a polymerase and a nucleotide mixture under certain conditions to produce an extended primer hybrid, wherein the mixture contains nucleotide homologs of no more than three of four different base types; (b) further extending the extended primer hybrid with a nucleotide homolog of the fourth of the four different base types in the absence of the nucleotide homologs in (a), thereby producing a further extended primer hybrid; (c) forming a stable ternary complex comprising the further extended primer hybrid, the polymerase, and a nucleotide homolog of the next base in the template; (d) detecting the stable ternary complex to distinguish the next base from other base types in the template; and (e) determining the presence of a base multivariate in the template nucleic acid, the base multivariate comprising the fourth of the four different base types, followed by the next base. Optionally, the method further comprises the following steps: (f) repeating steps (a) to (d) and (g) using the further extended primer hybrid as the primer-template nucleic acid hybrid to determine the presence of a series having at least two base multiple states in the template nucleic acid.

[0059] The nucleic acids used in the methods or compositions described herein can be DNA, such as genomic DNA, synthetic DNA, amplified DNA, copy DNA (cDNA), etc. RNA, such as mRNA, ribosomal RNA, tRNA, etc., can also be used. Nucleic acid analogs can also be used as templates in this document. Therefore, the template nucleic acids used herein can be derived from biological sources, synthetic sources, or amplification products. The primers used herein can be DNA, RNA, or analogs thereof.

[0060] A particularly useful nucleic acid template is a genomic fragment containing a sequence identical to a portion of the genome. A population of genomic fragments can contain at least 5%, 10%, 20%, 30%, or 40%, 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% of the genome. Genomic fragments can have a sequence substantially identical to, for example, at least about 25, 50, 70, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 or more nucleotides of the genome. Alternatively or additionally, genomic fragments can have a sequence not exceeding 1 × 10⁻⁶ nucleotides of the genome. 5 1×10 4 1×10 3Genome fragments can be essentially identical sequences of 1, 800, 600, 400, 200, 100, 75, 50, or 25 nucleotides. Genome fragments can be DNA, RNA, or analogues.

[0061] Exemplary organisms from which nucleic acids can be derived include, for example, those derived from mammals such as rodents, mice, rats, rabbits, guinea pigs, ungulates, horses, sheep, pigs, goats, cattle, cats, dogs, primates, humans, or non-human primates; plants such as Arabidopsis thaliana, maize, sorghum, oats, wheat, rice, rapeseed, or soybeans; algae such as Chlamydomonas reinhardtii; nematodes such as Caenorhabditis elegans; insects such as Drosophila melanogaster, mosquitoes, fruit flies, bees, or spiders; fish such as zebrafish; reptiles; amphibians such as frogs or Xenopus laevis; dictyostelium discoideum; and fungi such as Pneumocystis carinii. Nucleic acids can be derived from prokaryotes such as *Carinii*, *Takifugu rubripes*, yeast, *Saccharomyces cerevisiae*, or *Schizosaccharomyces pombe*; or *Plasmodium falciparum*. Nucleic acids can also be derived from prokaryotes such as bacteria, *Escherichia coli*, *Staphylococcus*, or *Mycoplasma pneumoniae*; archaea; viruses such as hepatitis C virus or human immunodeficiency virus; or viroids. Nucleic acids can be derived from homogeneous cultures or groups of organisms, or alternatively from collections of several different organisms (e.g., communities or ecosystems). Nucleic acids can be isolated using methods known in the art, including, for example, those described in Sambrook et al., Molecular Cloning: A Laboratory Manual, 3rd Edition, Cold Spring Harbor Laboratory, New York (2001) or in Ausubel et al., Current Protocols in Molecular Biology, Wiley and Sons, Baltimore, Maryland (1998), each of which is incorporated herein by reference.

[0062] Template nucleic acids can be obtained from preparation methods such as genome isolation, genome fragmentation, gene cloning, and / or amplification. Templates can be obtained from amplification techniques such as polymerase chain reaction (PCR), rolling circle amplification (RCA), and multiple displacement amplification (MDA). Exemplary methods for isolating, amplifying, and fragmenting nucleic acids to produce templates for array analysis are described in U.S. Patent Nos. 6,355,431 and 9,045,796, each of which is incorporated herein by reference. Amplification can also be performed using methods described in Sambrook et al., Molecular Cloning: A Laboratory Manual, 3rd Edition, Cold Spring Harbor Laboratory, New York (2001) or Ausubel et al., Methods in Contemporary Molecular Biology, John Willie & Son, Baltimore, Maryland (1998), each of which is incorporated herein by reference.

[0063] The method disclosed herein may include the following steps: contacting a primer-template nucleic acid hybrid with a polymerase and a nucleotide mixture under certain conditions to produce an extended primer hybrid. One or more nucleotides in the mixture may be reversibly terminated. For example, at least one, two, three, four, or more nucleotide types in the mixture may be reversibly terminated. Alternatively or additionally, up to four, three, two, or one nucleotide type in the mixture may be reversibly terminated. Similarly, one or more reversibly terminated nucleotide types in the mixture may be complementary to at least one, two, three, or four base types in the template nucleic acid. Alternatively or additionally, reversibly terminated nucleotide types in the mixture may be complementary to up to four, three, two, or one base type in the template nucleic acid. Reversibly terminated nucleotides and non-termining nucleotides may coexist in the extension reaction. For example, some or all nucleotide types may be delivered simultaneously in a single extension reaction. Alternatively, different nucleotide types may be delivered sequentially (alone or in subsets) such that they are combined into a single extension reaction. Nucleotide types may have base portions selected from those base portions described herein (such as those present in natural DNA or RNA or their analogues).

[0064] Adding a reversibly terminating nucleotide to the 3' end of a primer provides a means to prevent subsequent nucleotides from being added to the primer during the extension step and to further prevent unwanted extension of the primer in subsequent check steps. Therefore, positions in the template adjacent to a specific type of nucleotide can be checked. In such embodiments, a stable ternary complex can form at said position and be checked to detect the next correct nucleotide in the template hybridized to the extended, reversibly terminating primer. The method can be repeated stepwise by then removing or modifying the reversible termination portion from the extended, reversibly terminating primer to produce an extendable primer.

[0065] Typically, the reversibly terminated nucleotides added to the primers in the methods described herein do not have exogenous labels. This is because the extended primers do not need to be detected in the methods described herein. However, if desired, one or more types of reversibly terminated nucleotides used in the methods described herein can be detected, for example, by an exogenous label attached to the nucleotide. Exemplary reversible terminator portions, methods for incorporating said reversible terminator portions into primers, and methods for modifying primers for further extension (often referred to as “deblocking”) are described in U.S. Patent Nos. 7,544,794; 7,956,171; 8,034,923; 8,071,755; 8,808,989; or 9,399,798. Other examples are illustrated in Bentley et al., Nature 456:53-59 (2008), WO 04 / 018497; U.S. Patent No. 7,057,026; WO 91 / 06678; WO 07 / 123744; U.S. Patent No. 7,329,492; U.S. Patent No. 7,211,414; U.S. Patent No. 7,315,019; U.S. Patent No. 7,405,281 and US 2008 / 0108082, each of which is incorporated herein by reference.

[0066] As explained above, reversibly terminated nucleotides provide a means of stopping elongation at specific base types in the template nucleic acid. Another means of achieving this type of control over elongation is to use a mixture of nucleotides lacking nucleotide types that can pair with one or more types of nucleotides expected to be present in the template nucleic acid. For example, the nucleotide mixture used for elongation may lack at least one, two, or three types of nucleotides expected to pair with bases in the template. Alternatively or additionally, up to three, two, or one nucleotide types may be absent from the mixture. Similarly, the mixture may contain nucleotides complementary to no more than three, two, or one base types expected to be present in the template nucleic acid. Alternatively or additionally, the mixture may contain nucleotides complementary to at least one, two, or three base types expected to be present in the template nucleic acid. Different nucleotide types may be present simultaneously in the elongation reaction, or they may participate in the elongation reaction sequentially. For example, some or all nucleotide types may be delivered simultaneously in a single elongation reaction. Alternatively, different nucleotide types can be delivered sequentially (alone or in subsets) such that they are combined into a single elongation reaction or that sequential elongation reactions occur. Nucleotide types can have base motifs selected from those described herein (such as those present in natural DNA or RNA or their analogues).

[0067] The primer extension step can be performed by contacting the primer-template nucleic acid hybrid with the extension reaction mixture. Typically, the fluid present in the previous check step is removed and replaced by the extension reaction mixture. Alternatively, the extension reaction mixture can be formed by adding one or more reagents to the fluid present in the check step. Optionally, the extension reaction mixture comprises a composition of nucleotides different from those in the check step. For example, the check step may contain one or more nucleotide types not present in the extension reaction, and vice versa. By more specific examples, at least one type of nucleotide may be omitted in the extension step, and the check step may employ at least four types of nucleotides. Optionally, one or more nucleotide types are added to the check mixture for the primer extension step.

[0068] If nucleotides present in the inspection step are carried into the extension step, these nucleotides can cause unwanted nucleotide incorporation. Therefore, a washing step can be employed before the primer extension step to remove nucleotides. Optionally, free nucleotides can be removed by: enzymes such as phosphatases, adenosine triphosphate diphosphatases, or hexokinases; chemical modifications; or physical separation techniques such as liquid-phase extraction, solid-phase extraction, or electrophoretic separation.

[0069] In some embodiments, reagents for extension and detection may be present simultaneously. For example, the reaction mixture may contain a polymerase, a primer-template hybrid, a reversibly terminating nucleotide with one type of base homolog expected in the template, and an extendable nucleotide with three other types of base homologs expected in the template. In this case, primer extension will occur until the reversibly terminating nucleotide is incorporated and a ternary complex will form at the termination of the extended primer. The ternary complex will contain one of the four nucleotides that is the appropriate next correct nucleotide. The ternary complex can be detected by a label on the nucleotide, a label on the polymerase, or both. For example, a label may be present at the 5' position of the nucleotide such that the label is removed from any nucleotide incorporated into the primer during extension, but is retained for any nucleotide involved in the formation of the ternary complex. Thus, the ternary complex can be detected while avoiding unwanted signals from the incorporated nucleotide.

[0070] The primer extension step does not require the use of a labeled polymerase. For example, the polymerase used in the extension step does not need to be attached to a foreign label (e.g., covalently or otherwise). Alternatively, the polymerase used for primer extension may contain a foreign label, such as the label used in the previous examination step.

[0071] The methods disclosed herein may include a check step of forming and detecting a ternary complex. Examples of the methods utilize the specificity of a polymerase to form a stable ternary complex having a primer-template nucleic acid hybrid and the next correct nucleotide. The next correct nucleotide can bind nonvalently to the stable ternary complex, thereby interacting with the other members of the complex only through nonvalent interactions. Useful methods and compositions for forming stable ternary complexes are further described in detail below and in commonly owned U.S. Patent Application Publication No. 2017 / 0022553 A1 or U.S. Patent Application Serial No. 15 / 677,870 (published as U.S. 2018 / 0044727 A1); U.S. 2018 / 0187245 A1 (which claims priority to U.S. Patent Application Serial No. 62 / 440,624); or U.S. 2018 / 0208983 A1 (which claims priority to U.S. Patent Application Serial No. 62 / 450,397), each of which is incorporated herein by reference.

[0072] Typically, the inspection step is separated from the extension step, for example, due to the exchange of reagents between steps. However, as explained above, in some embodiments, the extension step and the inspection step may occur in the same mixture.

[0073] Although ternary complexes can exist in the absence of certain catalytic metal ions (e.g., Mg), 2+ In the presence of polymerase, primer-template nucleic acid hybrids, and the next correct nucleotide, formation occurs; however, chemically added nucleotides are inhibited in the absence of catalytic metal ions. Low or insufficient levels of catalytic metal ions induce non-covalent chelation of the next correct nucleotide in a stable ternary complex. Other methods disclosed herein can also be used to generate stable ternary complexes.

[0074] Optionally, a stable ternary complex can be formed when the primer of the primer-template nucleic acid hybrid contains a blocking portion (e.g., a reversible terminator portion) that prevents the enzymatic incorporation of the entering nucleotide into the primer. The interaction can occur in the presence of a stabilizer, thereby stabilizing the polymerase-nucleic acid interaction in the presence of the next correct nucleotide. Optionally, the primer of the primer-template nucleic acid hybrid can be an extendable primer or a primer that is blocked and cannot be extended at its 3' end (e.g., blocking can be achieved by the presence of a reversible terminator portion at the 3' end of the primer). The primer-template nucleic acid hybrid, polymerase, and homologous nucleotide can form a stable ternary complex when the base of the homologous nucleotide is complementary to the next base of the primer-template nucleic acid hybrid.

[0075] As described above, conditions favorable to or stabilizing ternary complexes can be provided by the presence of blocking groups that hinder enzymatic incorporation of nucleotides into primers (e.g., the reversible terminator portion of the 3' nucleotide of the primer) or the absence of catalytic metal ions. Other useful conditions include the presence of ternary complex stabilizers that inhibit nucleotide incorporation or polymerization (e.g., divalent or trivalent noncatalytic metal ions). Noncatalytic metal ions include, but are not limited to, calcium, strontium, scandium, titanium, vanadium, chromium, iron, cobalt, nickel, copper, zinc, gallium, germanium, arsenic, selenium, rhodium, europium, and terbium ions. Optionally, conditions unfavorable to binary complexes (i.e., complexes between the polymerase and the initiated nucleic acid but lacking homologous nucleotides) or destabilizing said binary complexes are provided by the presence of one or more monovalent cations and / or glutamate anions. As an additional option, polymerases engineered to have reduced catalytic activity or a reduced tendency to form binary complexes can be used.

[0076] As explained above, ternary complex stabilization conditions can, for example, emphasize the difference in polymerase affinity for primer-template nucleic acid hybrids in the presence of different nucleotides by destabilizing binary complexes. Optionally, the conditions result in different polymerase affinities for the primer-template in the presence of different nucleotides. For example, conditions include, but are not limited to, high salt concentrations and glutamate ions. For example, the salt can be dissolved in an aqueous solution to produce a monovalent cation, such as a monovalent metal cation (e.g., sodium or potassium ions). Optionally, the salt providing the monovalent cation (e.g., a monovalent metal cation) further provides glutamate ions. Optionally, the source of glutamate ions can be potassium glutamate. In some instances, the concentration of potassium glutamate used to alter the polymerase affinity of the primer-template hybrid can range from 10 mM to 1.6 M or any amount between 10 mM and 1.6 M. As indicated above, high salt concentration refers to a salt concentration ranging from 50 mM to 1,500 mM.

[0077] It should be understood that the options for stabilizing ternary complexes described herein need not be mutually exclusive and, on the contrary, can be used in various combinations. For example, ternary complexes can be stabilized by one or a combination of methods including, but not limited to, cross-linking of polymerase domains, cross-linking of polymerase with nucleic acids, polymerase mutations stabilizing ternary complexes, ectopic inhibition by small molecules, non-competitive inhibitors, competitive inhibitors, non-competitive inhibitors, the absence of catalytic metal ions, the presence of blocking moieties on primers, and other methods described herein.

[0078] Stable ternary complexes may comprise native nucleotides, nucleotide analogs, or modified nucleotides suited to a specific application or configuration desired for the method. Optionally, nucleotide analogs have a nitrogenous base, a pentose sugar, and a phosphate group, wherein any portion of the nucleotide may be modified, removed, and / or substituted compared to the native nucleotide. Nucleotide analogs may be non-incorporable nucleotides (i.e., nucleotides that cannot react with the 3' oxygen of a primer to form a covalent bond). Such non-incorporable nucleotides include, for example, monophosphate and diphosphate nucleotides. In another instance, the nucleotide may contain one or more modifications to the triphosphate group that makes the nucleotide non-incorporable. Examples of non-incorporable nucleotides can be found in U.S. Patent No. 7,482,120, which is incorporated herein by reference. In some embodiments, non-incorporable nucleotides may be subsequently modified to become incorporable. Non-incorporable nucleotide analogs include, but are not limited to, α-phosphate-modified nucleotides, α-β nucleotide analogs, β-phosphate-modified nucleotides, β-γ nucleotide analogs, γ-phosphate-modified nucleotides, or caged nucleotides. Examples of nucleotide analogs are described in U.S. Patent No. 8,071,755, which is incorporated herein by reference.

[0079] Nucleotide analogs involved in stabilizing ternary complexes may include terminators that reversibly prevent subsequent nucleotide incorporation at the 3' end of the primer after the analog has been incorporated into the primer. For example, U.S. Patents 7,544,794 and 8,034,923 (the disclosures of which are incorporated herein by reference) describe reversible terminators in which the 3'-OH group is partially substituted by 3'-ONH2. Another type of reversible terminator is attached to a nitrogenous base of a nucleotide, such as that described in U.S. Patent 8,808,989 (the disclosure of which is incorporated herein by reference). Other reversible terminators that can be similarly used in conjunction with the methods described herein are those described in references cited elsewhere herein or in U.S. Patents 7,956,171, 8,071,755, and 9,399,798 (the disclosures of which are incorporated herein by reference). In some embodiments, the reversible terminator portion may be removed from the primer in a process known as "deblocking," thereby allowing subsequent nucleotide incorporation. In the context of reversible terminators, compositions and methods for deblocking are described in the references cited herein. Other exemplary terminator moieties that may attach to the 3' oxygen include -CH2N3 (azidomethane), -CH2CH=CH2, photoreactive moieties, or other moieties described in Chen et al., Genomics, Proteomics & Bioinformatics 11:34-40 (2013) (the literature is incorporated herein by reference) or other references incorporated herein by reference.

[0080] Alternatively, nucleotide analogs reversibly prevent the incorporation of nucleotides at the 3' end of primers into which the nucleotide analog has already been incorporated. Irreversible nucleotide analogs comprise 2',3'-dideoxynucleotides (ddNTPs, such as ddGTP, ddATP, ddTTP, ddCTP). These dideoxynucleotides lack the 3'-OH group of the dNTP that would otherwise participate in polymerase-mediated primer extension. Therefore, the 3' position has a hydrogen moiety instead of the native hydroxyl moiety. Irreversible terminating nucleotides are particularly useful for genotyping applications or other applications where primer extension or sequential detection along template nucleic acids is undesirable.

[0081] In some embodiments, the nucleotides involved in forming the ternary complex may contain an exogenous label. For example, the exogenously labeled nucleotide may contain a reversible or irreversible terminator motif, the exogenously labeled nucleotide may be non-incorporable, the exogenously labeled nucleotide may lack a terminator motif, the exogenously labeled nucleotide may be incorporable, or the exogenously labeled nucleotide may be both incorporable and non-terminator. The exogenously labeled nucleotide can be particularly useful for forming a stable ternary complex with an unlabeled polymerase. Alternatively, the exogenous label on the nucleotide may provide one of the couplers in a fluorescence resonance energy transfer (FRET) pair, and the exogenous label on the polymerase may provide a second coupler in the pair. Thus, FRET detection can be used to identify stable ternary complexes containing two couplers. Alternatively, the nucleotides involved in forming the ternary complex may lack an exogenous label (i.e., the nucleotide may be "unlabeled"). For example, unlabeled nucleotides may contain reversible or irreversible terminator motifs, may be non-incorporatable, may lack terminator motifs, may be incorporatable, or may be both incorporatable and non-terminator. Unlabeled nucleotides can be useful when labeling on polymerases is used to detect stable ternary complexes. Unlabeled nucleotides can also be useful in extension steps of the methods described herein. It should be understood that the absence of a portion or function of a nucleotide means a nucleotide that does not have that function or portion. However, it should also be understood that one or more functions or portions of the nucleotide or its analogue described herein or otherwise known in the art may be explicitly omitted in the methods or compositions described herein.

[0082] Optionally, nucleotides (e.g., natural nucleotides or nucleotide analogs) are present in the mixture during the formation of a stable ternary complex. For example, at least one, two, three, four, or more nucleotide types may be present. Alternatively or additionally, up to four, three, two, or one nucleotide type may be present. Similarly, one or more nucleotide types present may be complementary to at least one, two, three, or four base types in the template nucleic acid. Alternatively or additionally, one or more nucleotide types present may be complementary to up to four, three, two, or one base type in the template nucleic acid.

[0083] Any nucleotide modification of the polymerase in the stable ternary complex can be used in the methods described herein. The nucleotide can be permanently or temporarily bound to the polymerase. Optionally, a nucleotide analogue is fused to the polymerase, for example, via a covalent linker. Optionally, multiple nucleotide analogues are fused to multiple polymerases, wherein each nucleotide analogue is fused to a different polymerase. Optionally, the nucleotides present in the stable ternary complex are not the means by which the ternary complex is stabilized. Therefore, any of the various other methods for stabilizing ternary complexes can be combined in a reaction utilizing nucleotide analogues.

[0084] In certain embodiments, the primer chains of the primer-template hybrid molecules present in the stable ternary complex do not undergo chemical changes during one or more steps of the method described herein, provided that a polymerase is present. For example, the primers do not need to be extended by forming new phosphodiester bonds or shortened by nucleolytic degradation during the steps for forming the stable ternary complex or for detecting the stable ternary complex.

[0085] Any polymerase of various kinds can be used to form a stable ternary complex in the methods described herein. Polymerases that can be used include naturally occurring polymerases and their modified variants, including, but not limited to, mutants, recombinants, fusions, genetically modified, chemically modified, synthetics, and analogs. Naturally occurring polymerases and their modified variants are not limited to polymerases capable of catalyzing polymerization reactions. Optionally, naturally occurring and / or their modified variants have the ability to catalyze polymerization reactions under conditions not used during the formation or examination of a stable ternary complex. Optionally, the naturally occurring and / or modified variants involved in the stable ternary complex have modified properties, such as enhanced affinity for nucleic acids, decreased affinity for nucleic acids, enhanced affinity for nucleotides, decreased affinity for nucleotides, enhanced specificity for the next correct nucleotide, decreased specificity for the next correct nucleotide, decreased catalytic rate, no catalytic activity, etc. Mutant polymerases include, for example, polymerases in which one or more amino acids are substituted by other amino acids or by the insertion or deletion of one or more amino acids. Exemplary polymerase variants that can be used to form stable ternary complexes include, for example, those polymerase variants set forth in U.S. Patent Application Serial No. 15 / 866,353 (now published as U.S. 2018 / 0155698 A1) or U.S. Patent Application Publication No. 2017 / 0314072 A1, each of which is incorporated herein by reference.

[0086] The modified polymerase comprises a polymerase containing an exogenous label motif (e.g., an exogenous fluorophore) that can be used to detect the polymerase. Optionally, the label motif can be attached after the polymerase has been at least partially purified using protein separation techniques. For example, the exogenous label motif can be chemically linked to the polymerase using the free thiol or free amino motif of the polymerase. This can involve chemically linking to the polymerase via a side chain of a cysteine ​​residue or via a free amino group at the N-terminus. The exogenous label motif can also be attached to the polymerase via protein fusion. Exemplary label motifs that can be attached via protein fusion include, for example, green fluorescent protein (GFP), phycobiliproteins (e.g., phycocyanin or phycoerythrin), or wavelength-shifted variants of GFP or phycobiliproteins. In some embodiments, the exogenous label on the polymerase can act as a member of a FRET pair. The other member of the FRET pair can be an exogenous label of a nucleotide attached to the polymerase bound to a stable ternary complex. Thus, a stable ternary complex can be detected or identified by FRET.

[0087] Alternatively, polymerases involved in stabilizing ternary complexes do not require attachment to an exogenous label. For example, polymerases do not require covalent attachment to an exogenous label. Instead, polymerases may lack any label until they are associated with labeled nucleotides and / or labeled nucleic acids (e.g., labeled primers and / or labeled templates).

[0088] Ternary complexes manufactured or used according to this disclosure may optionally contain one or more exogenous markers. The markers may be attached to components of the ternary complex (e.g., to polymerases, template nucleic acids, primers, and / or homologous nucleotides) prior to the formation of the ternary complex. Exemplary attachments include covalent or non-covalent attachments, such as those set forth herein, cited in the references herein, or known in the art. In some embodiments, the labeled component is delivered in solution to a solid support to which an unlabeled component is attached, thereby recruiting the marker to the immobilized support by means of the formation of a stable ternary complex. Thus, the component attached to the support can be detected or identified based on observation of the recruited markers. Whether used in solution or on a solid support, exogenous markers can be useful for detecting a stable ternary complex or its individual components during an inspection step. The exogenous markers may remain attached to the component after it has dissociated from other components that have formed a stable ternary complex. Exemplary markings, methods for affixing markings, and methods for using marked components are set forth in commonly owned U.S. Patent Application Publication No. 2017 / 0022553 A1 or U.S. Patent Application Serial No. 15 / 677,870 (published as U.S. 2018 / 0044727 A1); 15 / 581,383 (published as U.S. 2018 / 0187245 A1); 15 / 873,343 (published as U.S. 2018 / 0208983 A1); or U.S. Patent Application Publication No. 2018 / 0208983 A1 (which claims priority to U.S. Patent Application Serial Nos. 62 / 450,397 and 62 / 506,759), each of which is incorporated herein by reference.

[0089] Examples of useful exogenous labels include, but are not limited to, any of the following signal-generating motifs: radiolabeled motifs, luminescent motifs, fluorophore motifs, quantum dot motifs, chromophore motifs, enzyme motifs, electromagnetic spin-labeled motifs, nanoparticle light-scattering motifs, and various other signal-generating motifs known in the art. Suitable enzyme motifs include, for example, horseradish peroxidase, alkaline phosphatase, β-galactosidase, or acetylcholinesterase. Exemplary fluorophore motifs include, but are not limited to, umbelliferone, fluorescein, isothiocyanate, rhodamine, tetramethyl-rhodamine, eosin, green fluorescent protein, erythrosine, coumarin, methylcoumarin, pyrene, malachite green, violet, and lucifer yellow. TM Cascade Blue TM Texas Red (Texas Red) TM Dansyl chloride, phycoerythrin, phycocyanin, fluorescent lanthanide complexes (such as those containing europium and terbium), Cy3, Cy5, and other fluorophore moieties known in the art (such as those containing europium and terbium). Fluorescence Microscopy PrinciplesJoseph R. Lakowicz (ed.), Plenum Pub Corp., 2nd edition (July 1999) and Richard P. Hoagland Molecular Probes Handbook Those fluorophores described in the 6th edition).

[0090] Secondary markers may be used in the methods of this disclosure. A secondary marker is a binding portion that specifically binds to a ligand portion. For example, the ligand portion may be attached to a polymerase, nucleic acid, or nucleotide to allow detection by specific affinity for the labeled receptor. Exemplary pairs of binding portions that may be used include, but are not limited to, antigens and immunoglobulins or their active fragments, such as FAb; immunoglobulins and immunoglobulins (or their respective fragments); avidin and biotin or analogs thereof specific to avidin; streptavidin and biotin or analogs thereof specific to streptavidin; carbohydrates and lectins.

[0091] In some embodiments, the secondary labeling can be a chemically modifiable portion. In this embodiment, a label having a reactive functional group can be incorporated into a stable ternary complex. Subsequently, the functional group can covalently react with the primary labeling portion. Suitable functional groups include, but are not limited to, amino, carboxyl, maleimide, oxo, and thiol groups.

[0092] In alternative embodiments, the ternary complex may lack an exogenous label. For example, the ternary complex and all components involved in the ternary complex (e.g., polymerase, template nucleic acid, primers, and / or homologous nucleotides) may lack one, several, or all of the exogenous labels described herein or in the references incorporated above. In such embodiments, the ternary complex can be detected based on the inherent properties of the stable ternary complex, such as mass, charge, inherent optical properties, etc. Methods for detecting unlabeled ternary complexes are set forth in commonly owned U.S. Patent Application Publication No. 2017 / 0022553 A1, PCT Application Serial No. PCT / US16 / 68916 (published as WO 2017 / 117243), or U.S. Patent Application Serial Nos. 62 / 375,379 and 15 / 677,870 (published as U.S. 2018 / 0044727 A1), each of which is incorporated herein by reference.

[0093] Typically, detection can be performed in the inspection step by means of a method that can sense the properties of the ternary complex or the labeled portion attached thereto. Detection can be based on exemplary properties including, but not limited to, mass, conductivity, energy absorbance, fluorescence, etc. Detection of luminescence can be performed using methods known in the art related to nucleic acid arrays. The luminescent group can be detected based on any of a variety of luminescence properties, including, for example, emission wavelength, excitation wavelength, fluorescence resonance energy transfer (FRET) intensity, quenching, anisotropy, or lifetime. Other detection techniques that can be used in the methods described herein include, for example, mass spectrometry for sensing mass; surface plasmon resonance for sensing binding to a surface; absorbance at a wavelength for sensing the energy absorbed by the label; calorimetry for sensing temperature changes due to the presence of the label; conductivity or impedance for sensing the electrical properties of the label; or other known analytical techniques. Examples of reagents and conditions that can be used to create, manipulate, and detect stable ternary complexes include, for example, those set forth in commonly owned U.S. Patent Application Publication No. 2017 / 0022553 A1; PCT Application Serial No. PCT / US16 / 68916 (published as WO2017 / 117243); or U.S. Patent Application Serial Nos. 15 / 677,870 (published as U.S. 2018 / 0044727A1); 15 / 851,383 (published as U.S. 2018 / 0187245 A1); 15 / 873,343 (published as U.S. 2018 / 0208983 A1); or U.S. Patent Application Publication No. 2018 / 0208983 A1 (which claims priority to U.S. Patent Application Serial Nos. 62 / 450,397 and 62 / 506,759), each of which is incorporated herein by reference.

[0094] Specific embodiments of the methods described herein include the step of forming a mixture comprising several components. For example, the mixture may be formed between a primer-template nucleic acid hybrid, a polymerase, and one or more nucleotide types. The components of the mixture may be delivered in any desired order or they may be delivered simultaneously. Furthermore, some of the components may be mixed with each other to form a first mixture, which is then contacted with other components to form a more complex mixture. Taking the step of forming a mixture comprising a primer-template nucleic acid hybrid, a polymerase, and multiple different nucleotide types as an example, it should be understood that different nucleotide types among the multiple different nucleotide types may be contacted with each other before being contacted with the primer-template nucleic acid hybrid. Alternatively, two or more nucleotide types may be delivered individually to the primer-template hybrid and / or the polymerase. Thus, a first nucleotide type may be contacted with the primer-template hybrid before being contacted with a second nucleotide type. Alternatively or additionally, a first nucleotide type may be contacted with the polymerase before being contacted with a second nucleotide type.

[0095] Some embodiments of the methods described herein utilize two or more distinguishable signals to differentiate stable ternary complexes from each other and / or to differentiate one base type in a template nucleic acid from another. For example, two or more luminescent groups can be distinguished from each other based on unique optical properties, such as a unique wavelength of excitation or emission. In particular embodiments, the method can differentiate different stable ternary complexes based on differences in emission intensity. For example, the first ternary complex can be detected if the intensity emitted by the first ternary complex is lower than that of the second ternary complex. This intensity scale (sometimes referred to as a “gray scale”) can utilize any distinguishable intensity difference. Exemplary differences include specific stable ternary complexes whose intensity is 10%, 25%, 33%, 50%, 66%, or 75% compared to the intensity of another stable ternary complex to be detected.

[0096] Intensity differences can be achieved using different luminescent groups, each with different extinction coefficients (i.e., producing different excitation properties) and / or different emission quantum yields (i.e., producing different emission properties). Alternatively, the same type of luminescent group can be used, but in different amounts. For example, all members of a first group of the ternary complex can be labeled with a specific luminescent group, while only half of the members of a second group are labeled with a luminescent group. In this example, the second group is expected to produce half the signal of the first group. The second group can be generated, for example, by using a mixture of labeled and unlabeled nucleotides (compared to the first group which mainly contains labeled nucleotides). Similarly, the second group can be generated, for example, by using a mixture of labeled and unlabeled polymerases (compared to the first group which mainly contains labeled polymerases). In alternative labeling schemes, the first group of the ternary complex can contain multiple labeled polymerase molecules that produce a specific emission signal, and the second group of the ternary complex can contain polymerase molecules, each containing only one labeled polymerase molecule that produces the emission signal.

[0097] In some embodiments, the examination step is performed in a manner that estimating the identity of at least one nucleotide type, as set forth in, for example, commonly owned U.S. Patent Application Serial No. 15 / 712,632 (issued as U.S. Patent No. 9,951,385), each of which is incorporated herein by reference. For example, the examination step may comprise the following steps: (a) forming a mixture under ternary complex stabilization conditions, wherein the mixture comprises an initiating template nucleic acid, a polymerase, and nucleotide homologs of the first, second, and third base types in the template; (b) examining the mixture to determine whether a ternary complex has formed; and (c) identifying the next correct nucleotide of the initiating template nucleic acid molecule, wherein if a ternary complex is detected in step (b), the next correct nucleotide is identified as a homolog of the first, second, or third base type, and wherein the next correct nucleotide is estimated as a nucleotide homolog of a fourth base type based on the absence of the ternary complex in step (b).

[0098] Alternatively or additionally, for the purpose of using estimation, the examination step may use disambiguation to identify one or more nucleotide types, as set forth in, for example, the commonly owned U.S. Patent Application Serial No. 15 / 712,632 (issued as U.S. Patent No. 9,951,385), each of which is incorporated herein by reference. For example, the inspection step may include the following steps: (a) sequentially contacting the initiated template nucleic acid with at least two separate mixtures under ternary complex stability conditions, wherein the at least two separate mixtures each contain a polymerase and a nucleotide, such that the sequential contact in the initiated template nucleic acid results in contact with nucleotide homologs of the first, second, and third base types in the template under the ternary complex stability conditions; (b) inspecting the at least two separate mixtures to determine whether the ternary complex has formed; and (c) identifying the next correct nucleotide of the initiated template nucleic acid molecule, wherein if the ternary complex is detected in step (b), the next correct nucleotide is identified as a homolog of the first, second, or third base type, and wherein the next correct nucleotide is estimated as a nucleotide homolog of the fourth base type based on the absence of the ternary complex in step (b).

[0099] In a particular embodiment, the inspection may be performed by: (a) contacting the initiated template nucleic acid with a first mixture of polymerase and nucleotides under ternary complex stabilization conditions, wherein the first mixture comprises nucleotide homologs of a first base type and a second base type; (b) contacting the initiated template nucleic acid with a second mixture of polymerase and nucleotides under ternary complex stabilization conditions, wherein the second mixture comprises nucleotide homologs of the first base type and a third base type; (c) inspecting the product of steps (a) and (b) for a signal generated by the ternary complex comprising the initiated template nucleic acid, polymerase, and the next correct nucleotide, wherein the signal obtained for the product of step (a) is ambiguous for the first and second base types, and wherein the signal obtained for the product of step (b) is ambiguous for the first and third base types; and (d) disambiguating the signal obtained in step (c) to identify the base type binding the next correct nucleotide. Optionally, in order to achieve disambiguation, (i) the first base type is associated with the presence of a signal for the product of step (a) and the presence of a signal for the product of step (b), (ii) the second base type is associated with the presence of a signal for the product of step (a) and the absence of a signal for the product of step (b), and (iii) the third base type is associated with the absence of a signal for the product of step (a) and the presence of a signal for the product of step (b).

[0100] Provided is a method in which the examination comprises the following steps: (a) contacting an initiated template nucleic acid with a first mixture comprising a polymerase, a nucleotide homolog of a first base type in the template, and a nucleotide homolog of a second base type in the template, wherein the contact occurs in a binding reaction that (i) stabilizes a ternary complex comprising the initiated template nucleic acid, the polymerase, and the next correct nucleotide and (ii) prevents the incorporation of the next correct nucleotide into the primer; (b) examining the binding reaction to determine whether a ternary complex has formed; and (c) subjecting the initiated template nucleic acid to a repetition of steps (a) and (b), wherein the first mixture is replaced by a second mixture comprising a polymerase, a nucleotide homolog of the first base type in the template, and the next correct nucleotide. (b) a nucleotide homolog of the third base type in the template; and (d) using the test for the binding reaction or its product to identify the next correct nucleotide of the initiated template nucleic acid, wherein (i) if the ternary complex is detected in step (b) and in a repeat of step (b), the next correct nucleotide is identified as a homolog of the first base type; (ii) if the ternary complex is detected in step (b) and not in a repeat of step (b), the next correct nucleotide is identified as a homolog of the second base type; and (iii) if the ternary complex is not detected in step (b) and is detected in a repeat of step (b), the next correct nucleotide is identified as a homolog of the third base type.

[0101] As described above, different activities of polymerases can be utilized in the methods described herein. Polymerases can be useful, for example, in extension steps, inspection steps, or both. Different activities can arise from structural differences (e.g., through natural activity, mutants, or chemical modifications). However, polymerases can be obtained from a variety of known sources and applied according to the teachings set forth herein and the known activities of the polymerases. Useful DNA polymerases include, but are not limited to, bacterial DNA polymerases, eukaryotic DNA polymerases, archaea DNA polymerases, viral DNA polymerases, and bacteriophage DNA polymerases. Bacterial DNA polymerases include *Escherichia coli* DNA polymerases I, II, and III, IV, and V, the Klenow fragment of *Escherichia coli* DNA polymerase, *Clostridium stercorarium* (Cst) DNA polymerase, *Clostridium thermocellum* (Cth) DNA polymerase, and *Sulfolobus solfataricus* (Sso) DNA polymerase. Eukaryotic DNA polymerases include DNA polymerases α, β, γ, δ, ε, η, ζ, λ, σ, μ, and k, as well as Rev1 polymerase (terminal deoxycytidine transferase) and terminal deoxynucleotidyl transferase (TdT). Viral DNA polymerases include T4 DNA polymerase, phi-29 DNA polymerase, GA-1, phi-29-like DNA polymerase, PZA DNA polymerase, phi-15 DNA polymerase, Cp1 DNA polymerase, Cp7 DNA polymerase, T7 DNA polymerase, and T4 polymerase. Other useful DNA polymerases include thermostable and / or thermophilic DNA polymerases, such as *Thermus aquaticus* (Taq) DNA polymerase, *Thermus filiformis* (Tfi) DNA polymerase, *Thermococcus zilligi* (Tzi) DNA polymerase, *Thermus thermophilus* (Tth) DNA polymerase, *Thermus flavusu* (Tfl) DNA polymerase, *Pyrococcus woesei* (Pwo) DNA polymerase, *Pyrococcus furiosus* (Pfu) DNA polymerase and Turbo Pfu DNA polymerase, *Thermococcus litoralis* (Tli) DNA polymerase, and *Pyrococcus sp.*GB-D polymerase, *Thermotoga maritima* (Tma) DNA polymerase, *Bacillus stearothermophilus* (Bst) DNA polymerase, *Pyrococcus Kodakaraensis* (KOD) DNA polymerase, Pfx DNA polymerase, *Thermococcus sp.* JDF-3 (JDF-3) DNA polymerase, *Thermococcus gorgonarius* (Tgo) DNA polymerase, *Thermococcus acidophilium* DNA polymerase, *Sulfolobus acidocaldarius* DNA polymerase, *go N-7* DNA polymerase of thermococci, *Pyrodictium occultum* DNA polymerase, *Methanococcus* DNA polymerases including: *Methanococcus thermoautotrophicum* DNA polymerase; *Methanococcus jannaschii* DNA polymerase; *Desulfurococcus* strain TOK DNA polymerase (D. Tok Pol); *Pyrococcus abyssi* DNA polymerase; *Pyrococcus horikoshii* DNA polymerase; *Pyrococcus islandicum* DNA polymerase; *Thermococcus fumicolans* DNA polymerase; *Aeropyrum pernix* DNA polymerase; and heterodimeric DNA polymerases DP1 / DP2. Engineered and modified polymerases are also useful in connection with the disclosed techniques. For example, extremely thermophilic marine archaea such as *Thermococcus species 9°N* (e.g., Therminator DNA polymerase from New England BioLabs Inc., Ipswich, Massachusetts) can be used. Other useful DNA polymerases (including 3PDX polymerases) are disclosed in U.S. Patent No. 8,703,461, the disclosure of which is incorporated herein by reference.

[0102] Useful RNA polymerases include, but are not limited to, viral RNA polymerases such as T7 RNA polymerase, T3 polymerase, SP6 polymerase, and Kll polymerase; eukaryotic RNA polymerases such as RNA polymerase I, RNA polymerase II, RNA polymerase III, RNA polymerase IV, and RNA polymerase V; and archaeal RNA polymerases.

[0103] Another useful type of polymerase is reverse transcriptase. Exemplary reverse transcriptases include, but are not limited to, HIV-1 reverse transcriptase (PDB 1HMV) from human immunodeficiency virus type 1, HIV-2 reverse transcriptase from human immunodeficiency virus type 2, M-MLV reverse transcriptase from Moroni murine leukemia virus, AMV reverse transcriptase from avian myeloblastoma virus, and telomere reverse transcriptase for maintaining telomeres of eukaryotic chromosomes.

[0104] For some embodiments, polymerases having intrinsic 3'-5' proofed exonuclease activity can be useful. In some embodiments, such as most genotyping and sequencing embodiments, polymerases substantially lacking 3'-5' proofed exonuclease activity are also useful. The absence of exonuclease activity can be a wild-type characteristic or a characteristic conferred by a variant or engineered polymerase structure. For example, the exo-Klenow fragment is a mutant version of the Klenow fragment lacking 3'-5' proofed exonuclease activity. Klenow fragments and their exo-variants can be useful in the methods or compositions described herein.

[0105] Examples of reagents and conditions that may be used in polymerase-based primer extension steps include, for example, those set forth in commonly owned U.S. Patent Application Publication No. 2017 / 0022553 A1 or U.S. Patent Application Serial Nos. 15 / 677,870 and 2018 / 0044727 A1; 15 / 581,383 and 2018 / 0187245 A1; or U.S. Patent Application Serial Nos. 2018 / 0208983 A1 and 62 / 450,397 and 62 / 506,759, each of which is incorporated herein by reference. Other useful reagents and conditions for polymerase-based primer extension are described in Bentley et al., Nature, 456:53-59 (2008), WO 04 / 018497; WO 91 / 06678; WO 07 / 123744; U.S. Patent Nos. 7,057,026; 7,329,492; 7,211,414; 7,315,019 or 7,405,281, and U.S. Patent Application Publication No. 2008 / 0108082 A1, each of which is incorporated herein by reference.

[0106] In specific embodiments, the reagents used during the primer extension step are removed from contact with the primer-template hybrid before the step of forming a stable ternary complex with a primer-template hybrid. For example, removal of the nucleotide mixture used in the extension step may be desirable when one or more types of nucleotides in the mixture interfere with the formation or detection of the ternary complex in a subsequent check step. Similarly, removal of the polymerase or cofactor used in the primer extension step may be desirable to prevent unwanted catalytic activity during the check step. Following removal, a washing step may be performed, where an inert liquid is used to remove residual components of the primer-template hybrid from the extension mixture.

[0107] The washing step can be performed between any of the various steps described herein. For example, the washing step can be used to separate the primer-template hybrid from other reagents that have been contacted with the primer-template hybrid under ternary complex stability conditions. Such washing can remove one or more reagents from interference with the examination of the mixture or from contamination of a second mixture to be formed on a substrate (container) that has previously been contacted with the first mixture. For example, the primer-template nucleic acid hybrid can be contacted with a polymerase and at least one type of nucleotide to form a first mixture under ternary complex stability conditions, and the first mixture can be examined. Optionally, washing can be performed before examination to remove reagents that have not participated in the formation of a stable ternary complex. Alternatively or additionally, washing can be performed after the examination step to remove one or more components of the first mixture from the primer-template hybrid. The primer-template nucleic acid hybrid can then be contacted with a polymerase and at least one other nucleotide to form a second mixture under ternary complex stability conditions, and a second mixture can be formed for examination of the ternary complex. As previously mentioned, optional washing can be performed before a second examination to remove reagents that have not participated in the formation of a stable ternary complex.

[0108] The methods disclosed herein may comprise multiple repetitions of the steps described herein. Such repetitions may provide the sequence of the template nucleic acid or a characteristic marker of the template nucleic acid. In specific embodiments, the repetitions generate information that can be used to determine a series of base multiple states that provide a characteristic marker of the template nucleic acid. The check and extension steps may be repeated multiple times, as may optional steps to deblock the primers or wash away unwanted reactants or products between the various steps. Thus, the primer-template nucleic acid hybrid may undergo at least 2, 5, 10, 25, 50, 100, or more steps of the methods described herein. Not all steps need to be repeated, nor do repeated steps need to occur in the same order in each repetition. For example, real-time analysis (i.e., in parallel with the fluid and detection steps of the sequencing method) may be used to identify the next correct nucleotide at each position in the template. However, real-time analysis is not required and may instead be used to identify the next correct nucleotide after some or all of the fluid and detection steps that have already been completed.

[0109] Therefore, this disclosure further provides a method for characterizing nucleic acids. The method may include the following steps: (a) contacting a primer-template nucleic acid hybrid with a polymerase and a nucleotide mixture under certain conditions to generate an extended primer hybrid, wherein the mixture comprises nucleotide homologs of first, second, third, and fourth different base types, wherein the nucleotide homolog of the first base type has a reversible terminator, and wherein the nucleotide homologs of the second, third, and fourth base types are extendable; (b) contacting the extended primer hybrid with at least one nucleotide homolog of the different base types and a polymerase to form a stable ternary complex comprising the extended primer hybrid, the polymerase, and a nucleotide homolog of the next base in the template; (c) detecting the stable ternary complex to determine the next base's affinity for other bases in the template. (a) Differentiating base types; (d) Determining the presence of a base multimorphism in the template nucleic acid, the base multimorphism comprising the first base type followed by the next base; (e) Repeating steps (a) to (c) using the extended primer hybrid as the primer-template nucleic acid hybrid; (f) Determining the presence of a series having at least two base multimorphisms in the template nucleic acid; and (g) Repeating steps (a) to (f), wherein the nucleotide homolog of the second base type comprises a reversible terminator, and wherein the nucleotide homologs of the first, third, and fourth base types in the mixture are extended, and wherein the base multimorphism determined in step (c) comprises the second base type followed by the next base, thereby determining the presence of a series of two base multimorphisms in the template nucleic acid. Optionally, the method may further comprise (h) repeating steps (a) to (f), wherein the nucleotide homolog of the third base type contains a reversible terminator, and wherein the nucleotide homologs of the first, second, and fourth base types in the mixture are extendable, and wherein the base multistate determined in step (c) contains the third base type, followed by the next base, thereby determining the presence of a three-base multistate series in the template nucleic acid. Alternatively, the method may comprise (i) repeating steps (a) to (f), wherein the nucleotide homolog of the fourth base type contains a reversible terminator, and wherein the nucleotide homologs of the first, second, and third base types in the mixture are extendable, and wherein the base multistate determined in step (c) contains the fourth base type, followed by the next base, thereby determining the presence of a four-base multistate series in the template nucleic acid.Those skilled in the art will understand that when repeating steps (a) to (f), steps (a) to (f) can be repeated using the same procedure as previous steps (a) to (f), but the nucleotide homologs containing reversible terminators and the extended nucleotide homologs are different from one or more previous iterations of steps (a) to (f) in the same method.

[0110] This disclosure further provides a method for characterizing nucleic acids, the method comprising the steps of: (a) contacting a primer-template nucleic acid hybrid with a polymerase and a nucleotide mixture under certain conditions to generate an extended primer hybrid and form a stable ternary complex comprising the extended primer hybrid, the polymerase, and a nucleotide homolog of the next base in the template, wherein the mixture comprises nucleotide homologs of first, second, third, and fourth different base types, wherein the first base type of nucleotide homolog includes a reversible terminator, and wherein the second, third, and fourth base types of nucleotide homolog are extended; (b) detecting the stable ternary complex to distinguish the next base from other base types in the template; and (c) determining the base polymorphism. The presence of a multiplet state in the template nucleic acid, wherein the multiplet state comprises the first base type followed by the next base; (d) repeating steps (a) and (b) using the extended primer hybrid as the primer-template nucleic acid hybrid; (e) determining the presence of a series having at least two multiplet states in the template nucleic acid; and (f) repeating steps (a) to (e), wherein the nucleotide homolog of the second base type has a reversible terminator, and wherein the nucleotide homologs of the first, third, and fourth base types in the mixture are extended, and wherein the multiplet state determined in step (c) comprises the second base type followed by the next base, thereby determining the presence of a series of two multiplet states in the template nucleic acid. Optionally, the method may further comprise the following steps: (g) repeating steps (a) to (e), wherein the nucleotide homolog of the third base type has a reversible terminator, and wherein the nucleotide homologs of the first, second, and fourth base types in the mixture are extendable, and wherein the base multistate determined in step (c) comprises the third base type, followed by the next base, thereby determining the presence of a three-base multistate series in the template nucleic acid. Alternatively, the method comprises the following step (h) repeating steps (a) to (e), wherein the nucleotide homolog of the fourth base type has a reversible terminator, and wherein the nucleotide homologs of the first, second, and third base types in the mixture are extendable, and wherein the base multistate determined in step (c) comprises the fourth base type, followed by the next base, thereby determining the presence of a four-base multistate series in the template nucleic acid. Those skilled in the art will understand that when repeating steps (a) to (e), steps (a) to (e) can be repeated using the same procedure as previous steps (a) to (e), but the nucleotide homologs containing reversible terminators and the extended nucleotide homologs are different from one or more previous iterations of steps (a) to (e) in the same method.

[0111] Also provided is a method for characterizing nucleic acids, the method comprising the steps of: (a) contacting a primer-template nucleic acid hybrid with a polymerase and a nucleotide mixture under certain conditions to produce an extended primer hybrid, wherein the mixture comprises nucleotide homologs of first, second, third, and fourth different base types, wherein the first and second base types of the nucleotide homologs have reversible terminators, and wherein the third and fourth base types of the nucleotide homologs are extendable; (b) contacting the extended primer hybrid with at least one base type of the nucleotide homologs and a polymerase to form a nucleotide homolog comprising the extended primer hybrid, the polymerase, and the next base in the template. (a) A stable ternary complex; (c) Detecting the stable ternary complex to distinguish the next base from other base types in the template; (d) Determining the presence of base multivariates in the template nucleic acid; (e) Repeating steps (a) to (c) using the extended primer hybrid as the primer-template nucleic acid hybrid; (f) Determining the presence of a series having at least two base multivariates in the template nucleic acid; and (g) Repeating steps (a) to (f), wherein the nucleotide homologs of the third and fourth base types contain reversible terminators, and wherein the nucleotide homologs of the first and second base types in the mixture are extendable, thereby determining the presence of a two-base multivariate series in the template nucleic acid. Optionally, the method may further comprise (h) Repeating steps (a) to (f), wherein the nucleotide homologs of the first and third base types contain reversible terminators, and wherein the nucleotide homologs of the second and fourth base types in the mixture are extendable, thereby determining the presence of a three-base multivariate series in the template nucleic acid. Alternatively, the method may comprise (i) repeating steps (a) to (f), wherein the nucleotide homologs of the second and fourth base types contain reversible terminators, and wherein the nucleotide homologs of the first and third base types in the mixture are extendable, thereby determining the presence of a four-base multivariate series in the template nucleic acid. Those skilled in the art will understand that when repeating steps (a) to (f), steps (a) to (f) can be repeated using the same procedure as previous steps (a) to (f), but the nucleotide homologs containing reversible terminators and the extendable nucleotide homologs differ from one or more previous iterations of steps (a) to (f) in the same method.

[0112] This disclosure further provides a method for characterizing nucleic acids, the method comprising the steps of: (a) contacting a primer-template nucleic acid hybrid with a polymerase and a nucleotide mixture under certain conditions to generate an extended primer hybrid and form a stable ternary complex comprising the extended primer hybrid, the polymerase, and a nucleotide homolog of the next base in the template, wherein the mixture comprises nucleotide homologs of first, second, third, and fourth different base types, wherein the first and second base types of the nucleotide homologs include a reversible terminator, and wherein the third and fourth base types of the nucleotide homologs are extended; (b) detecting the stable ternary complex. The complex is used to distinguish the next base from other base types in the template; (c) determining the presence of base multimorphism in the template nucleic acid; (d) repeating steps (a) and (b) using the extended primer hybrid as the primer-template nucleic acid hybrid; (e) determining the presence of a series having at least two base multimorphisms in the template nucleic acid; and (f) repeating steps (a) to (e), wherein the nucleotide homologs of the third and fourth base types have reversible terminators, and wherein the nucleotide homologs of the first and second base types in the mixture are extendable, thereby determining the presence of a two-base multimorphic series in the template nucleic acid. Optionally, the method may further comprise the following steps: (g) repeating steps (a) to (e), wherein the nucleotide homologs of the first and third base types have reversible terminators, and wherein the nucleotide homologs of the second and fourth base types in the mixture are extendable, thereby determining the presence of a three-base multimorphic series in the template nucleic acid. Alternatively, the method comprises the following steps: (h) repeating steps (a) to (e), wherein the nucleotide homologs of the second and fourth base types have reversible terminators, and wherein the nucleotide homologs of the first and third base types in the mixture are extendable, thereby determining the presence of a four-base multivariate series in the template nucleic acid. Those skilled in the art will understand that when repeating steps (a) to (e), steps (a) to (e) can be repeated using the same procedure as previous steps (a) to (e), but the nucleotide homologs containing reversible terminators and the extendable nucleotide homologs differ from one or more previous iterations of steps (a) to (e) in the same method.

[0113] This disclosure further provides a method for characterizing nucleic acids, the method comprising the steps of: (a) contacting a primer-template nucleic acid hybrid with a polymerase and a nucleotide mixture under certain conditions to produce an extended primer hybrid, wherein the mixture contains nucleotide homologs of no more than three of four different base types; (b) further extending the extended primer hybrid with a nucleotide homolog of the fourth of the four different base types in the absence of the nucleotide homolog in (a), thereby producing a further extended primer hybrid; (c) forming a stable ternary complex comprising the further extended primer hybrid, the polymerase, and a nucleotide homolog of the next base in the template; and (d) detecting the stable ternary complex to identify the next base in the template. (e) Distinguishing one base from the other base types in the template; (f) Determining the presence of a base multimorphism in the template nucleic acid, the base multimorphism comprising the fourth base type of the four different base types, followed by the next base; (g) Repeating steps (a) to (d) using the further extended primer hybrid as the primer-template nucleic acid hybrid; (h) Determining the presence of a series having at least two base multimorphisms in the template nucleic acid; and (h) Repeating steps (a) to (g), wherein the nucleotide homolog of the fourth base type is replaced by a nucleotide homolog of the first base type and wherein the nucleotide homolog of the first base type is not used in step (a), thereby determining the presence of a two-base multimorphic series in the template nucleic acid. Optionally, the method may further comprise (i) Repeating steps (a) to (g), wherein the nucleotide homolog of the fourth base type is replaced by a nucleotide homolog of the second base type and wherein the nucleotide homolog of the second base type is not used in step (a), thereby determining the presence of a three-base multimorphic series in the template nucleic acid. Alternatively, the method may include (j) repeating steps (a) to (g), wherein the nucleotide homolog of the fourth base type is replaced by a nucleotide homolog of the third base type and wherein the nucleotide homolog of the third base type is not used in step (a), thereby determining the presence of a four-base multivariate series in the template nucleic acid. Those skilled in the art will understand that when repeating steps (a) to (g), steps (a) to (g) can be repeated using the same procedure as previous steps (a) to (g), but the nucleotide homologs containing reversible terminators, the extended nucleotide homologs, and the unused nucleotide homologs differ from one or more previous iterations of steps (a) to (g) in the same method.

[0114] In specific embodiments of the above methods, several steps are repeated using a template having the same sequence. In some cases, the same template molecule is used. For example, the steps can be performed to generate extended primers that hybridize to the template; the extended primers can be removed from the template; and new primers can hybridize to the template for repeated steps. The extended primers can be removed using methods known in the art for denaturing double-stranded nucleic acids (e.g., thermal denaturants or chemical denaturants). Exemplary methods for denaturing nucleic acids are described in Sambrook et al., Molecular Cloning: A Laboratory Manual, 3rd Edition, Cold Spring Harbor Laboratory, New York (2001) or in Ausubel et al., Contemporary Methods in Molecular Biology, John Willie & Son, Baltimore, Maryland (1998), each of which is incorporated herein by reference. Typically, the template nucleic acid will be attached to the surface by covalent or strong non-covalent bonds. Therefore, denaturing conditions that effectively separate the nucleic acid strands but are sufficiently mild to maintain template-to-surface attachment can be used. The surface can then be washed or rinsed to remove the extended primers while leaving the template on the surface. In some embodiments, the extended strands can be degraded, for example, by chemicals, nucleases, or physical shearing. Alternatively, the template molecule is not reused, and instead a second template molecule is used in the second iteration of the method, wherein the second template molecule has a template region whose sequence is the same as that of the template used in the first iteration of the method.

[0115] In some embodiments, the template nucleic acid can be characterized to define a sequence of a region (at single-base resolution) and to identify one or more base multimorphic sequences that characterize a second region providing the template. For example, the two regions can be contiguous, such that the sequenced region is adjacent to a tail region with a characterized base multimorphic sequence. This tail characterization can be located upstream or downstream of the sequenced region. Information from both regions can be used to help, for example, align the template sequence with a reference sequence during resequencing applications.

[0116] Therefore, this disclosure provides a method for characterizing nucleic acids. The method may include the following steps: (a) (i) performing a sequencing process, the sequencing process including polymerase extension of primers hybridized to a template to generate a primer-template nucleic acid hybrid and detecting a signal indicating a sequence located at the 3' position of the template having at least two base multiple states; and (ii) contacting the primer-template nucleic acid hybrid with a polymerase and a nucleotide mixture under certain conditions to generate an extended primer hybrid, wherein the mixture contains nucleotide homologs of first, second, third, and fourth different base types, wherein the nucleotide homolog of the first base type has a reversible terminator, and wherein at least one of the nucleotide homologs of the second, third, or fourth base types is extendable; and (b) contacting the extended primer hybrid with the... (a) A nucleotide homolog of at least one base type from different base types is contacted with a polymerase to form a stable ternary complex comprising the extended primer hybrid, the polymerase, and a nucleotide homolog of the next base in the template; (b) The stable ternary complex is detected to distinguish the next base from other base types in the template; (c) The presence of a base multivariate in the template nucleic acid is determined, the base multivariate comprising the first base type followed by the next base; (e) Steps (a) to (c) and (f) are repeated using the extended primer hybrid as the primer-template nucleic acid hybrid to determine the presence of a series having at least two base multivariates in the template nucleic acid, wherein the first region of the template is located at the 3' position of the series having at least two base multivariates.

[0117] In another embodiment, a method for characterizing nucleic acids may include the following steps: (a) (i) performing a sequencing process, the sequencing process including polymerase extension of primers hybridized to a template to generate a primer-template nucleic acid hybrid and detecting a signal indicating a sequence located in a first region of the template at the 3' position of the series having at least two base multiple states; and (ii) contacting the primer-template nucleic acid hybrid with a polymerase and a nucleotide mixture under certain conditions to generate an extended primer hybrid and form a stable ternary complex comprising the extended primer hybrid, the polymerase, and a nucleotide homolog of the next base in the template, wherein the mixture comprises nucleotide homologs of first, second, third, and fourth different base types, wherein the first base... The nucleotide homologs of the base type include a reversible terminator, and at least one of the nucleotide homologs of the second, third, or fourth base type is extendable; (b) detect the stable ternary complex to distinguish the next base from other base types in the template; (c) determine the presence of a base multivariate in the template nucleic acid, the base multivariate comprising the first base type followed by the next base; (d) repeat steps (a) and (b) using the extended primer hybrid as the primer-template nucleic acid hybrid; and I determine the presence of a series having at least two base multivariates in the template nucleic acid, wherein the first region of the template is located at the 3' position of the series having at least two base multivariates.

[0118] This disclosure further provides a method for characterizing nucleic acids, the method comprising the steps of: (a) (i) performing a sequencing process, the sequencing process including polymerase extension of primers hybridized to a template to generate a primer-template nucleic acid hybrid and detecting a signal indicating a sequence located in a first region of the template at the 3' position of a series having at least two base multiplexes; and (ii) contacting the primer-template nucleic acid hybrid with a polymerase and a nucleotide mixture under certain conditions to generate an extended primer hybrid, wherein the mixture contains nucleotide homologs of no more than three of four different base types; and (b) further extending the extended primer hybrid with a nucleotide homolog of the fourth of the four different base types in the absence of the nucleotide homologs in (a), thereby generating a primer-template nucleic acid hybrid. (a) Generate a further extended primer hybrid; (b) Form a stable ternary complex comprising the further extended primer hybrid, polymerase, and a nucleotide homolog of the next base in the template; (c) Detect the stable ternary complex to distinguish the next base from other base types in the template; (e) Determine the presence of a base multivariate in the template nucleic acid, the base multivariate comprising the fourth base type of the four different base types, followed by the next base; (f) Repeat steps (a) to (d) using the extended primer hybrid as the primer-template nucleic acid hybrid; and (g) Determine the presence of a series having at least two base multivariates in the template nucleic acid, wherein the first region of the template is located at the 3' position of the series having at least two base multivariates.

[0119] In some embodiments, the template nucleic acid can be characterized as one or more base multimorphic sequences that identify the sequences of two regions (at single-base resolution) and a third region that provides a characteristic marker for intercalating the two sequenced regions. The two regions can be analyzed as paired reads separated by adapters with characteristic markers such as base multimorphic sequence markers or length markers. Typical paired read alignment methods can be improved by not only aligning the two sequenced regions to a reference genome but also by using adapter characteristic markers as the basis for alignment.

[0120] Therefore, this disclosure provides a method for characterizing paired reads of a template nucleic acid, wherein the template includes a linker region adjacent to and between the first and second regions of the template nucleic acid, wherein the method comprises the steps of: (a) obtaining a single-base resolution sequence of the first region by extending a primer along the linker region of the template nucleic acid, thereby generating an extended primer-template hybrid; (b) obtaining a characterization marker of the linker region by: (i) further extending the extended primer of the extended primer-template hybrid using a nucleotide mixture, wherein the nucleotide mixture includes nucleosides of first, second, third, and fourth different base types. (i) an acid homolog, wherein the nucleotide homolog of the first base type includes a reversible terminator, and wherein at least one of the nucleotide homologs of the second, third, or fourth base types is extendable; and (ii) detecting a stable ternary complex comprising a further extended primer-template hybrid, a polymerase, and a nucleotide homolog of the next base in the template, wherein the signature includes a base multivariate comprising the first base type followed by the next base; and (c) obtaining a single-base resolution sequence of the second region by extending the primer of the further extended primer-template hybrid.

[0121] This disclosure provides a method for characterizing paired reads of nucleic acids, wherein the template includes a linker region adjacent to and between first and second regions of the template nucleic acid, wherein the method comprises the steps of: (a) obtaining a single-base resolution sequence of the first region by extending a primer along the linker region of the template nucleic acid, thereby generating an extended primer-template hybrid; (b) obtaining a characterization marker of the linker region by: (i) further extending the extended primer of the extended primer-template hybrid using a nucleotide mixture, wherein the nucleotide mixture comprises nucleotide homologs of no more than three of four different base types; (ii) in the absence of step... In the case of the nucleotide homolog of step (i), the extended primer-template hybrid of step (i) is further extended with a nucleotide homolog of the fourth base type among the four different base types; (iii) a stable ternary complex is detected, the stable ternary complex comprising the extended primer-template hybrid of step (ii), polymerase, and a nucleotide homolog of the next base in the template, wherein the feature marker comprises a base multivariate comprising the fourth base type among the four different base types, followed by the next base; and (c) a single-base resolution sequence of the second region is obtained by extending the primer of the extended primer-template hybrid of step (iii).

[0122] This disclosure also provides a method for characterizing paired reads of nucleic acids, wherein the template includes a linker region adjacent to first and second regions of the template nucleic acid and between the first and second regions, wherein the method comprises the steps of: (a) obtaining a single-base resolution sequence of the first region by extending a primer along the linker region of the template nucleic acid, thereby generating an extended primer-template hybrid; (b) obtaining a characterizing marker of the linker region by further extending the extended primer without distinguishing the types of nucleotides incorporated into the extended primer of the extended primer-template hybrid, wherein the characterizing marker includes a count of nucleotides in the linker region; and (c) obtaining a single-base resolution sequence of the second region by extending the primer of the extended primer-template hybrid of step (b). Optionally, step (b) is performed without using labeled nucleotides and labeled polymerase. Optionally, step (b) is performed without using a marker that distinguishes the different types of nucleotides incorporated into the extended primer-template hybrid. Optionally, step (b) is performed without detecting the extended primer-template hybrid.

[0123] The paired read method illustrated above allows adjacent regions of a template nucleic acid to be sequenced by extending primers along the single strand of the nucleic acid molecule. This contrasts with the paired-end method, in which a first primer is extended along the first strand of the template, the primer extension product is then removed from the template, the template is replicated to produce a second strand (i.e., the second strand is the complement of the first strand), and a second primer is then extended along the second strand. The specific configuration of the paired read method described herein does not require a process step for denaturing the product of the first sequence read from the template before performing the second paired read. The specific configuration of the paired read method described herein does not require a process step for repeating the first strand of the template after it has been sequenced to produce the second strand of the second primer extension process. The specific configuration of the paired read method described herein does not require primers to be extended in opposite directions along the opposite strands of the template to obtain two reads separately.

[0124] This disclosure also provides a method for characterizing nucleic acids. The method may include the following steps: (a) (i) performing a sequencing process, the sequencing process including polymerase extension of primers hybridized to a template to generate a primer-template nucleic acid hybrid and detecting a signal indicating a sequence located at the 3' position of the template having at least two base multiplexes; and (ii) contacting the primer-template nucleic acid hybrid with a polymerase and a nucleotide mixture under certain conditions to generate an extended primer hybrid, wherein the mixture comprises nucleotide homologs of first, second, third, and fourth different base types, wherein the nucleotide homolog of the first base type has a reversible terminator, and wherein at least one of the nucleotide homologs of the second, third, or fourth base types is extendable; and (b) contacting the extended primer hybrid with at least one of the nucleotide homologs of the different base types and a polymerase to form a hybrid containing the extended primer hybrid. (a) A stable ternary complex of the primer hybrid, the polymerase, and the nucleotide homolog of the next base in the template; (b) Detecting the stable ternary complex to distinguish the next base from other base types in the template; (c) Determining the presence of a base multivariate in the template nucleic acid, the base multivariate comprising the first base type followed by the next base; (e) Repeating steps (a) to (c) using the extended primer hybrid as the primer-template nucleic acid hybrid; (f) Determining the presence of a sequence having at least two base multivariates in the template nucleic acid, wherein the first region of the template is located at the 3' position of the sequence having at least two base multivariates; and (g) Polymerase extension of the extended primer during a second sequencing process, wherein the second sequencing process detects a signal indicating a sequence of the template located at the 5' position of the sequence having at least two base multivariates.

[0125] Also provided is a method for characterizing nucleic acids, the method comprising the steps of: (a) (i) performing a sequencing process, the sequencing process including polymerase extension of primers hybridized to a template to generate a primer-template nucleic acid hybrid and detecting a signal indicating a sequence located at the 3' position of the template having at least two base multiple states; and (ii) contacting the primer-template nucleic acid hybrid with a polymerase and a nucleotide mixture under certain conditions to generate an extended primer hybrid and form a stable ternary complex comprising the extended primer hybrid, the polymerase, and a nucleotide homolog of the next base in the template, wherein the mixture comprises nucleotide homologs of first, second, third, and fourth different base types, wherein the nucleotide homolog of the first base type includes a reversible terminator, and wherein the nucleotide homolog of the second, third, or fourth base type... (a) At least one of the nucleotide homologs is extendable; (b) the stable ternary complex is detected to distinguish the next base from other base types in the template; (c) the presence of a base multivariate in the template nucleic acid is determined, the base multivariate comprising the first base type followed by the next base; (d) steps (a) and (b) are repeated using the extended primer hybrid as the primer-template nucleic acid hybrid; (e) the presence of a series having at least two base multivariates in the template nucleic acid is determined, wherein the first region of the template is located at the 3' position of the series having at least two base multivariates; and (f) polymerase extension of the extended primer is performed in a second sequencing process, wherein the second sequencing process detects a signal indicating a sequence of the template located at the 5' position of the series having at least two base multivariates.

[0126] In another embodiment, a method for characterizing nucleic acids may include the following steps: (a) (i) performing a sequencing process, the sequencing process including polymerase extension of primers hybridized to a template to produce a primer-template nucleic acid hybrid and detecting a signal indicating a sequence located in a first region of the template at the 3' position of the series having at least two base multiplexes; and (ii) contacting the primer-template nucleic acid hybrid with a polymerase and a nucleotide mixture under certain conditions to produce an extended primer hybrid, wherein the mixture contains nucleotide homologs of no more than three of four different base types; (b) further extending the extended primer hybrid with a nucleotide homolog of the fourth of the four different base types in the absence of the nucleotide homologs in (a) to produce a further extended primer hybrid; and (c) forming a primer hybrid comprising the further extended primer hybrid, polymerase, and a polymerase. (a) A stable ternary complex of a nucleotide homolog of the next base in the template; (d) Detecting the stable ternary complex to distinguish the next base from other base types in the template; (e) Determining the presence of a base multivariate in the template nucleic acid, the base multivariate comprising the fourth base type of the four different base types, followed by the next base; (f) Repeating steps (a) to (d) using the extended primer hybrid as the primer-template nucleic acid hybrid; (g) Determining the presence of a sequence having at least two base multivariates in the template nucleic acid, wherein the first region of the template is located at the 3' position of the sequence having at least two base multivariates; and (h) Polymerase extension of the extended primer during a second sequencing process, wherein the second sequencing process detects a signal indicating a sequence of the template located at the 5' position of the sequence having at least two base multivariates.

[0127] The signatures obtained using the methods described herein can have a variety of uses, several of which are illustrated herein. Another use of the signatures obtained using the methods described herein is to provide a count of repetitive sequences present in a particular template nucleic acid. For example, the method for counting single tandem repeats (STRs) or other repetitive sequence elements described in U.S. Patent Application Publication No. 2017 / 0137873 A1 (which is incorporated herein by reference) can be modified to use the methods described herein. Alternatively, the methods or signatures described herein can be used for purposes other than counting repetitive sequence elements. For example, the methods described herein can be used with template sequences that do not contain repetitive sequence units. In such cases, the template or signature derived from the template may lack repetitive sequence units, wherein the missing repetitive units are at least 3, 4, 5, 10 or more nucleotides long. Alternatively or additionally, the template or signature derived from the template may not have repetitive units adjacent to each other. For example, any repetitive units in the template or signature may be spaced apart from each other by regions of non-repetitive sequences at least 3, 4, 5, 10 or more nucleotides long.

[0128] Stable ternary complexes, or components capable of forming (i.e., participating in the formation) of ternary complexes, can be attached to a solid support. The solid support can be made of any of a variety of materials used in analytical biochemistry. Suitable materials may include glass, polymer materials, silicon, quartz (fused silica), borofloat glass, silica, silica-based materials, carbon, metals, optical fibers or fiber bundles, sapphire, or plastic materials. Specific materials can be selected based on properties desired for a particular application. For example, materials transparent to a desired wavelength of radiation are useful for analytical techniques that will utilize said wavelength of radiation. Conversely, materials that do not transmit radiation at a certain wavelength may be desirable (e.g., opaque, absorptive, or reflective). Other properties of the materials that can be utilized may include inertness or reactivity to certain reagents used in downstream processes, as those described herein, ease of manipulation, or low manufacturing cost.

[0129] Particularly useful solid supports are particles such as beads or microspheres. Bead clusters can be used to attach groups of stable ternary complexes or components capable of forming complexes (e.g., polymerases, templates, or nucleotides). In some embodiments, it may be useful to use a configuration where each bead thus has a single type of stable ternary complex or a single type of component capable of forming a complex. For example, a single bead may attach to a single type of ternary complex, a single type of template allele, a single type of allele-specific primer, a single type of locus-specific primer, or a single type of nucleotide. Alternatively, different types of components do not need to be separated on a bead-by-bead basis. Thus, a single bead can carry multiple different types of ternary complexes, template nucleic acids, primers, primer-template nucleic acid hybrids, and / or nucleotides. The composition of the beads can vary depending on, for example, the format to be used, the chemistry, and / or the attachment method. Exemplary bead compositions include a solid support used in protein and nucleic acid capture methods and the chemical functions imparted to it. Such compositions include, for example, plastics, ceramics, glass, polystyrene, melamine, methylstyrene, acrylic polymers, paramagnetic materials, thorium sol, carbon graphite, titanium dioxide, latex, or cross-linked dextran (such as Sepharose). TM Cellulose, nylon, cross-linked micelles and Teflon TM (and other materials described in the "Microsphere Testing Guide" from Bangs Laboratories, Fisher Industries (which is incorporated herein by reference).)

[0130] The geometry of particles, beads, or microspheres can also correspond to a wide variety of different forms and shapes. For example, they can be symmetrical (e.g., spherical or cylindrical) or irregular (controlled porosity glass). Additionally, beads can be porous, thereby increasing the surface area available for trapping ternary composites or their components. Exemplary sizes of beads used herein can range from nanometers to millimeters or from about 10 nm to 1 mm.

[0131] In certain embodiments, the beads may be arranged in an array or otherwise spatially differentiated. Exemplary bead-based arrays that may be used include, but are not limited to, BeadChip, available from Illumina Corporation (San Diego, California). TMArrays, or arrays such as those disclosed in U.S. Patent Applications Nos. 6,266,459, 6,355,431, 6,770,441, 6,859,570, or 7,622,294; or PCT Publication WO 00 / 63437, each of which is incorporated herein by reference. Beads may be positioned at discrete locations (e.g., holes) on a solid support, whereby each location accommodates a single bead. Alternatively, the discrete locations where beads reside may each contain multiple beads, as described, for example, in U.S. Patent Application Publications Nos. 2004 / 0263923 A1, 2004 / 0233485 A1, 2004 / 0132205 A1, or 2004 / 0125424 A1, each of which is incorporated herein by reference.

[0132] As will be appreciated from the above bead array embodiments, the methods of this disclosure can be performed in multiple formats, thereby enabling the parallel detection of multiple different types of nucleic acids in the methods described herein. Although it is also possible to process different types of nucleic acids sequentially using one or more steps of the methods described herein, parallel processing can provide cost savings, time savings, and uniformity of conditions. The apparatus or methods of this disclosure may comprise at least 2, 10, 100, or 1 × 10⁻⁶ beads. 3 1×10 4 1×10 5 1×10 6 1×10 9 One or more distinct nucleic acids. Alternatively or additionally, the apparatus or method of this disclosure may contain up to 1 × 10⁻⁶ different nucleic acids. 9 1×10 6 1×10 5 1×10 4 1×10 3 One, 100, 10, 2, or fewer different nucleic acids. Therefore, the various reagents or products (e.g., primer-template nucleic acid hybrids or stable ternary complexes) described herein can be reused to have different types or species within these ranges.

[0133] Other examples of commercially available arrays that can be used include, for example, Affymetrix's GeneChip. TM Arrays. According to some embodiments, spotting arrays can also be used. An exemplary spotting array is the CodeLink array available from Amersham Biosciences. TMArrays. Another useful array method is using inkjet printing methods (such as SurePrint, available from Agilent Technologies). TM An array manufactured using technology.

[0134] Other useful arrays include those used in nucleic acid sequencing applications. For example, arrays for attaching amplicon fragments (often referred to as clusters) can be particularly useful. Examples of nucleic acid sequencing arrays that may be used herein include those described in Bentley et al., Nature 456:53-59 (2008), PCT Publications WO 91 / 06678; WO04 / 018497 or WO 07 / 123744; U.S. Patents 7,057,026; 7,211,414; 7,315,019; 7,329,492 or 7,405,281; or U.S. Patent Application Publication 2008 / 0108082, each of which is incorporated herein by reference.

[0135] Nucleic acids can be attached to a support in a manner that provides detection at the single-molecule level or at the integrated level. For example, multiple different nucleic acids can be attached to a solid support in such a way that a single, stable ternary complex formed on a nucleic acid molecule on the support can be distinguished from all adjacent ternary complexes formed on nucleic acid molecules on the support. Thus, one or more different templates can be attached to a solid support in a manner in which each individual molecular template is physically separated and can be detected in a way that distinguishes individual molecules from all other molecules on the solid support.

[0136] Alternatively, the methods of this disclosure can target one or more nucleic acid ensembles, where an ensemble is a group of nucleotides having a common template sequence. Clustering methods can be used to attach one or more ensembles to a solid support. Thus, an array can have multiple ensembles, each of which is referred to as a cluster or array feature in the format. Clusters can be formed using methods known in the art, such as bridging amplification or emulsion PCR. Useful bridging amplification methods are described, for example, in U.S. Patent Nos. 5,641,658 or 7,115,400; or U.S. Patent Publication Nos. 2002 / 0055100 A1; 2004 / 0002090 A1; 2004 / 0096853 A1; 2007 / 0128624 A1; or 2008 / 0009420 A1. Emulsion PCR methods include, for example, those described in Dressman et al., Proceedings of the National Academy of Sciences (Proc. Natl. Acad. Sci. USA) 100:8817-8822 (2003), WO 05 / 010145, or U.S. Patent Publication Nos. 2005 / 0130173 A1 or 2005 / 0064460 A1, each of which is incorporated herein by reference in its entirety. Another useful method for amplifying nucleic acids on a surface is rolling circle amplification (RCA), such as that described, for example, in Lizardi et al., Nature Genetics 19:225-232 (1998), or U.S. Patent No. 2007 / 0099208 A1, each of which is incorporated herein by reference.

[0137] In certain embodiments, a stable ternary complex, polymerase, nucleic acid, or nucleotide is attached to a solid support on or within the flow cell surface. The flow cell allows for convenient fluid manipulation by passing a solution in and out of a fluid chamber in contact with the ternary complex bound to the support. The flow cell also provides for the detection of the fluidly manipulated component. For example, a detector may be positioned to detect signals from the solid support, such as signals from tags recruited to the solid support due to the formation of the stable ternary complex. Exemplary flow cells that may be used are described, for example, in U.S. Patent Application Publication No. 2010 / 0111768 A1, WO 05 / 065814, or U.S. Patent Application Publication No. 2012 / 0270305 A1, each of which is incorporated herein by reference.

[0138] This disclosure provides a system for detecting nucleic acids, for example, using the methods set forth herein. For example, the system can be configured to perform a reaction involving examining the interaction between a polymerase and a primer-template nucleic acid hybrid in the presence of nucleotides to identify one or more bases in a template nucleic acid sequence. Optionally, the system includes components and reagents for performing one or more steps set forth herein, said steps including, but not limited to: forming at least one stable ternary complex between the primer-template nucleic acid hybrid, the polymerase, and the next correct nucleotide; detecting one or more stable ternary complexes; extending the primers and / or identifying nucleotides of each primer-template hybrid, the sequence of the nucleotides, or a series of base multiples present in the template.

[0139] The systems disclosed herein may include containers or solid supports for performing nucleic acid detection methods. For example, the system may include arrays, flow cells, multiwell plates, or other convenient devices. The containers or solid supports may be removable, allowing them to be placed into or removed from the system. Thus, the system can be configured to sequentially process multiple containers or solid supports. The system may include a fluid system having a reservoir for one or more reagents containing the reagents described herein (e.g., polymerases, primers, template nucleic acids, one or more nucleotides for ternary complex formation, nucleotides for primer extension, deblocking reagents, or mixtures of such components). The fluid system may be configured to deliver reagents to the containers or solid supports, for example, through channels or droplet transfer devices (e.g., electrowetting devices). Any of a variety of detection devices may be configured to detect containers or solid supports in which reagents interact. Examples include luminescent detectors, surface plasmon resonance detectors, and other detection devices known in the art. Exemplary systems having fluid and detection components that can be readily modified for use in the systems described herein include, but are not limited to, those described in U.S. Patent Application Publication No. 2018 / 0280975 A1 and its priority application, U.S. Patent Application Serial No. 62 / 481,289; U.S. Patent No. 8,241,573; No. 7,329,860 or 8,039,817; or U.S. Patent Application Publication No. 2009 / 0272914 A1 or 2012 / 0270305 A1, each of which is incorporated herein by reference.

[0140] Optionally, the system of this disclosure further includes a computer processing unit (CPU) configured as an operating system component. The same or different CPUs can interact with the system to acquire, store, and process signals (e.g., signals detected in the methods described herein). In a particular embodiment, the CPU can be used to determine the identity of a nucleotide present at a specific location in the template nucleic acid based on a signal. In some cases, the CPU will identify the sequence of nucleotides in the template based on the detected signal. In a particular embodiment, the CPU is programmed to identify base multiplets present in the sequence based on the identity of the nucleotide at the end of the primer hybridized to the template (as determined based on the composition of the nucleotides in the extension mixture) and based on the identity of the next correct nucleotide (as determined based on a specific binding reaction performed using a primer-template hybrid).

[0141] A useful CPU can comprise one or more of the following: personal computer systems, server computer systems, thin clients, fat clients, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, smartphones, and distributed cloud computing environments that include any of the aforementioned systems or devices. The CPU can include one or more processors or processing units and can include a memory architecture of RAM or non-volatile memory. The memory architecture can further include removable / non-removable, volatile / non-volatile computer storage media. Further, the memory architecture can include one or more readers for reading from or writing to a non-removable non-volatile magnetic medium (hard disk drive), a disk drive for reading from or writing to a removable non-volatile disk, and / or an optical disk drive for reading from or writing to a removable non-volatile optical disk (such as a CD-ROM or DVD-ROM). The CPU can also include various computer system readable media. Such media can be any available media that can be accessed by a cloud computing environment, such as volatile and non-volatile media and removable and non-removable media.

[0142] A memory architecture may contain at least one program product having at least one program module having executable instructions implemented to perform one or more steps of the methods described herein. For example, executable instructions may include an operating system, one or more application programs, other program modules, or program data. Typically, a program module may contain routines, programs, objects, components, logic, data structures, etc., that perform the specific task described herein.

[0143] CPU components can be coupled via one or more internal buses that can be implemented as any of several types of bus architectures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, and a processor or local bus using any of the various bus architectures. By way of example, and not limitation, such architectures include the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MCA) bus, the Enhanced ISA (EISA) bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnect (PCI) bus.

[0144] Optionally, the CPU can communicate with one or more external devices (such as a keyboard), pointing devices (e.g., a mouse), displays (such as a graphical user interface (GUI)), or other devices that facilitate interaction with the nucleic acid detection system. Similarly, the CPU can communicate with other devices (e.g., via a network interface card, modem, etc.). This communication can occur through an I / O interface. Furthermore, the CPU of the system described herein can communicate with one or more networks, such as a local area network (LAN), a wide area network (WAN), and / or a public network (e.g., the Internet), via a suitable network adapter.

[0145] This disclosure further provides kits for characterizing nucleic acids. The kits may contain reagents for performing one or more of the methods described herein. For example, the kits may contain reagents for generating stable ternary complexes when mixed with one or more primer-template nucleic acid hybrids. More specifically, the kits may contain one or more mixtures of nucleotides used in the methods described herein (including, for example, those described in the Examples section below). In addition to nucleotide mixtures, the kits may contain polymerases capable of forming stable ternary complexes and / or polymerases for extension steps. Nucleotides, polymerases, or both may contain exogenous markers (e.g., as described herein in the context of the various methods).

[0146] Therefore, any of the components or articles used to perform the methods described herein can be usefully packaged into a kit. For example, a kit may be packaged to contain some, many, or all of the components or articles used to perform the methods described herein. Exemplary components include, for example, labeled nucleotides (e.g., extendable labeled nucleotides), polymerases (labeled or unlabeled), nucleotides with terminator motifs (e.g., unlabeled reversibly terminated nucleotides), deblocking reagents, etc., as described herein and in the references cited herein. Any of such reagents may contain, for example, some, many, or all of the buffers, components, and / or articles used to perform one or more subsequent steps in the analysis of primer-template nucleic acid hybrids. The kit does not need to contain primers or template nucleic acids. Instead, the user of the kit may provide primer-template nucleic acid hybrids to be combined with the components of the kit.

[0147] The kit may also contain one or more adjuvants. Such adjuvants may include any of the reagents exemplified above and / or other types of reagents for performing the methods described herein. The kit may further contain instructions. Instructions may include, for example, procedures for preparing any component or article used in the methods described herein, procedures for performing one or more steps of any embodiment of the methods described herein, and / or instructions for performing any of the analytical steps in subsequent analytical steps using a primer-template nucleic acid hybrid.

[0148] In a particular embodiment, the kit includes a reservoir for containing reagents and further includes a cartridge for transferring the reagents from the reservoir to a detection instrument. For example, the fluid cartridge may be configured to transfer the reagents to a flow cell in which a stable ternary complex is detected. Exemplary fluid cartridges that may be included in kits (or systems) of the present disclosure are described in U.S. Patent Application Publication No. 2018 / 0280975 and its priority application, U.S. Patent Application Serial No. 62 / 481,289, each of which is incorporated herein by reference.

[0149] Other embodiments

[0150] 1. A method for characterizing nucleic acids, the method comprising:

[0151] (a) Under certain conditions, the primer-template nucleic acid hybrid is contacted with a mixture of polymerase and nucleotides to produce an extended primer hybrid.

[0152] The mixture comprises nucleotide homologs of first, second, third and fourth different base types, wherein the nucleotide homolog of the first base type comprises a reversible terminator, and wherein at least one of the nucleotide homologs of the second, third or fourth base types is extensible;

[0153] (b) Contact the extended primer hybrid with a nucleotide homolog of at least one of the different base types and a polymerase to form a stable ternary complex comprising the extended primer hybrid, the polymerase and a nucleotide homolog of the next base in the template.

[0154] (c) Detecting the stable ternary complex to distinguish the next base from other base types in the template; and

[0155] (d) Determine the presence of a base multimorphism in the template nucleic acid, the base multimorphism comprising the first base type followed by the next base.

[0156] 2. The method according to Embodiment 1, further comprising:

[0157] (e) Repeat steps (a) to (c) using the extended primer hybrid as the primer-template nucleic acid hybrid.

[0158] (f) Determine the presence of a series of nucleic acids having at least two base multiple states in the template nucleic acid.

[0159] 3. The method according to Example 2, wherein the nucleotide homologs of the second, third and fourth base types are extensible in step (a).

[0160] 4. The method according to Example 3, further comprising:

[0161] (g) Repeat steps (a) to (f), wherein the nucleotide homolog of the second base type includes a reversible terminator, and wherein the nucleotide homologs of the first, third, and fourth base types in the mixture are extensible, and wherein the base multivariate determined in step (c) includes the second base type, followed by the next base.

[0162] This confirms the presence of two base multiple state series in the template nucleic acid.

[0163] 5. The method according to any one of Examples 3 to 4, further comprising:

[0164] (h) Repeat steps (a) to (f), wherein the nucleotide homolog of the third base type includes a reversible terminator, and wherein the nucleotide homologs of the first, second, and fourth base types in the mixture are extensible, and wherein the base multivariate determined in step (c) includes the third base type, followed by the next base; and

[0165] (i) Repeat steps (a) to (f), wherein the nucleotide homolog of the fourth base type includes a reversible terminator, and wherein the nucleotide homologs of the first, second, and third base types in the mixture are extensible, and wherein the base multivariate determined in step (c) includes the fourth base type, followed by the next base.

[0166] This confirmed the presence of a four-base multivariate sequence in the template nucleic acid.

[0167] 6. The method according to Example 5 further includes a computer-aided process of comparing the four base multiple state series to determine the sequence of the nucleic acid.

[0168] 7. The method according to Example 2, wherein the nucleotide homologs of the first and second base types include reversible terminators, and wherein the nucleotide homologs of the third and fourth base types are extensible in step (a).

[0169] 8. The method according to embodiment 7, further comprising:

[0170] (g) Repeat steps (a) to (f), wherein the nucleotide homologs of the third and fourth base types include reversible terminators, and wherein the nucleotide homologs of the first and second base types in the mixture are extensible.

[0171] 9. The method according to embodiment 8, further comprising:

[0172] (h) Repeat steps (a) to (f), wherein the nucleotide homologs of the first and third base types include reversible terminators, and wherein the nucleotide homologs of the second and fourth base types in the mixture are extensible.

[0173] 10. The method according to any one of Examples 2 to 9, wherein steps (a) to (f) are repeated using the template nucleic acid after the extended primer is removed from the extended primer hybrid and then the second primer is hybridized to the template.

[0174] 11. The method according to any one of Examples 2 to 9, wherein steps (a) to (f) are repeated using a second template nucleic acid molecule whose sequence is identical to that of the template nucleic acid used in step (a).

[0175] 12. The method according to Example 2, wherein the primer-template nucleic acid hybrid of step (a) is the product of polymerase extension during the sequencing process, wherein the sequencing process detects a signal indicating the sequence of the template at the 3' position of the series having at least two base multiple states.

[0176] 13. The method according to Example 12 further includes a computer-aided process of aligning the sequence of the template with a region of a reference genome, wherein the alignment is performed by aligning the region of the reference genome with the sequence of the template at the 3' position of the series having at least two base multiples.

[0177] 14. The method according to Example 13, further comprising (g) polymerase extension of the extended primer during a second sequencing process, wherein the second sequencing process detects a signal indicating a sequence of the template at the 5' position of the series having at least two base multiple states.

[0178] 15. The method according to Example 14, further comprising a computer-aided process of aligning the sequence of the template with a region of the reference genome, wherein the alignment is performed by aligning the region of the reference genome with the sequence of the template at position 3' of the sequence having at least two base multiples, the sequence having at least two base multiples, and the sequence of the template at position 5' of the sequence having at least two base multiples.

[0179] 16. The method according to any one of Examples 1 to 15, wherein the nucleotide homologs of the first, second, third and fourth different base types include the exogenous markers detected in step (c).

[0180] 17. The method according to any one of Examples 1 to 15, wherein the nucleotide homologs of the first, second, third and fourth different base types do not include the exogenous markers detected in step (c).

[0181] 18. The method according to any one of Examples 1 to 17, wherein the polymerase includes the exogenous marker detected in step (c).

[0182] 19. The method according to any one of Examples 1 to 18, wherein step (b) is performed after the mixture has been removed from contact with the extended primer hybrid.

[0183] 20. A method for characterizing nucleic acids, the method comprising:

[0184] (a) Under certain conditions, the primer-template nucleic acid hybrid is contacted with a mixture of polymerase and nucleotides to produce an extended primer hybrid and form a stable ternary complex comprising the extended primer hybrid, the polymerase, and a nucleotide homolog of the next base in the template.

[0185] The mixture comprises nucleotide homologs of first, second, third and fourth different base types, wherein the nucleotide homolog of the first base type comprises a reversible terminator, and wherein at least one of the nucleotide homologs of the second, third or fourth base types is extensible;

[0186] (b) Detecting the stable ternary complex to distinguish the next base from other base types in the template; and

[0187] (c) Determine the presence of a base multivariate in the template nucleic acid, the base multivariate comprising the first base type followed by the next base.

[0188] 21. The method according to embodiment 20, further comprising:

[0189] (d) Repeat steps (a) and (b) using the extended primer hybrid as the primer-template nucleic acid hybrid; and

[0190] (e) Determine the presence of a series of nucleic acids having at least two base multiple states in the template nucleic acid.

[0191] 22. The method according to Example 21, wherein the nucleotide homologs of the second, third and fourth base types are extensible in step (a).

[0192] 23. The method according to embodiment 22, further comprising:

[0193] (f) Repeat steps (a) to (e), wherein the nucleotide homolog of the second base type includes a reversible terminator, and wherein the nucleotide homologs of the first, third, and fourth base types in the mixture are extensible, and wherein the base multivariate determined in step (c) includes the second base type, followed by the next base.

[0194] This confirms the presence of two base multiple state series in the template nucleic acid.

[0195] 24. The method according to embodiment 23, further comprising:

[0196] (g) Repeat steps (a) to (e), wherein the nucleotide homolog of the third base type includes a reversible terminator, and wherein the nucleotide homologs of the first, second, and fourth base types in the mixture are extensible, and wherein the base multivariate determined in step (c) includes the third base type, followed by the next base; and

[0197] (h) Repeat steps (a) to (e), wherein the nucleotide homolog of the fourth base type includes a reversible terminator, and wherein the nucleotide homologs of the first, second, and third base types in the mixture are extensible, and wherein the base multivariate determined in step (c) includes the fourth base type, followed by the next base.

[0198] This confirmed the presence of a four-base multivariate sequence in the template nucleic acid.

[0199] 25. The method according to Example 24, further comprising a computer-aided process of comparing the four base multiple state series to determine the sequence of the nucleic acid.

[0200] 26. The method according to Example 21, wherein the nucleotide homologs of the first and second base types include reversible terminators, and wherein the nucleotide homologs of the third and fourth base types are extensible in step (a).

[0201] 27. The method according to embodiment 26, further comprising:

[0202] (f) Repeat steps (a) to (e), wherein the nucleotide homologs of the third and fourth base types include reversible terminators, and wherein the nucleotide homologs of the first and second base types in the mixture are extensible.

[0203] 28. The method according to embodiment 27, further comprising:

[0204] (g) Repeat steps (a) to (e), wherein the nucleotide homologs of the first and third base types include reversible terminators, and wherein the nucleotide homologs of the second and fourth base types in the mixture are extensible.

[0205] 29. The method according to any one of Examples 21 to 28, wherein steps (a) to (e) are repeated using the template nucleic acid after the extended primer is removed from the extended primer hybrid and then the second primer is hybridized to the template.

[0206] 30. The method according to any one of Examples 21 to 28, wherein steps (a) to (e) are repeated using a second template nucleic acid molecule whose sequence is identical to that of the template nucleic acid used in step (a).

[0207] 31. The method according to Example 21, wherein the primer-template nucleic acid hybrid of step (a) is the product of polymerase extension during the sequencing process, wherein the sequencing process detects a signal indicating the sequence of the template at the 3' position of the series having at least two base multiple states.

[0208] 32. The method according to Example 31 further includes a computer-aided process of aligning the sequence of the template with a region of a reference genome, wherein the alignment is performed by aligning the region of the reference genome with the sequence of the template at the 3' position of the series having at least two base multiples.

[0209] 33. The method according to Example 32, further comprising (f) polymerase extension of the extended primer during a second sequencing process, wherein the second sequencing process detects a signal indicating a sequence of the template at the 5' position of the series having at least two base multiple states.

[0210] 34. The method according to Example 33 further includes a computer-aided process of aligning the sequence of the template with a region of the reference genome, wherein the alignment is performed by aligning the region of the reference genome with the sequence of the template at position 3' of the sequence having at least two base multiples, the sequence having at least two base multiples, and the sequence of the template at position 5' of the sequence having at least two base multiples.

[0211] 35. The method according to any one of Examples 20 to 34, wherein the nucleotide homologs of the first, second, third and fourth different base types include the exogenous markers detected in step (b).

[0212] 36. The method according to any one of Examples 20 to 34, wherein the nucleotide homologs of the first, second, third and fourth different base types do not include the exogenous markers detected in step (b).

[0213] 37. The method according to any one of Examples 20 to 36, wherein the polymerase includes the exogenous marker detected in step (b).

[0214] 38. A method for characterizing nucleic acids, the method comprising:

[0215] (a) Under certain conditions, a primer-template nucleic acid hybrid is contacted with a polymerase and a nucleotide mixture to produce an extended primer hybrid, wherein the mixture comprises nucleotide homologs of no more than three of four different base types.

[0216] (b) In the absence of the nucleotide homolog of (a), the extended primer hybrid is further extended using the nucleotide homolog of the fourth base type among the four different base types, thereby producing a further extended primer hybrid.

[0217] (c) Forming a stable ternary complex comprising the further extended primer hybrid, polymerase, and a nucleotide homolog of the next base in the template;

[0218] (d) Detecting the stable ternary complex to distinguish the next base from other base types in the template; and

[0219] (e) Determine the presence of a base multimorphism in the template nucleic acid, the base multimorphism comprising the fourth base type of the four different base types, followed by the next base.

[0220] 39. The method according to embodiment 38, further comprising:

[0221] (f) Repeat steps (a) to (d) using the further extended primer hybrid as the primer-template nucleic acid hybrid.

[0222] (g) Determine the presence of a series of nucleic acids having at least two base multiple states in the template nucleic acid.

[0223] 40. The method according to embodiment 39, further comprising:

[0224] (h) Repeat steps (a) to (g), wherein the nucleotide homolog of the fourth base type is replaced by a nucleotide homolog of the first base type and wherein, instead of using the nucleotide homolog of the first base type in step (a), three other nucleotide homologs are used.

[0225] This confirms the presence of two base multiple state series in the template nucleic acid.

[0226] 41. The method according to embodiment 40, further comprising:

[0227] (i) Repeating steps (a) to (g), wherein the nucleotide homolog of the first base type is replaced by a nucleotide homolog of the second base type and wherein, instead of using the nucleotide homolog of the second base type in step (a), three other nucleotide homologs are used; and

[0228] (j) Repeat steps (a) to (g), wherein the nucleotide homolog of the second base type is replaced by a nucleotide homolog of the third base type, and wherein the nucleotide homolog of the third base type is not used in step (a) but three other nucleotide homologs are used instead.

[0229] This confirmed the presence of a four-base multivariate sequence in the template nucleic acid.

[0230] 42. The method according to Example 41 further includes a computer-aided process of comparing the four base multimorphic series to determine the sequence of the nucleic acid.

[0231] 43. The method according to any one of Examples 41 to 42, wherein steps (a) to (g) are repeated using a primer-template nucleic acid hybrid generated by removing an extended primer from the extended primer hybrid and hybridizing another primer to the template.

[0232] 44. The method according to any one of Examples 41 to 42, wherein steps (a) to (g) are repeated using a primer-template nucleic acid hybrid molecule whose sequence is identical to that of the primer-template nucleic acid hybrid used in step (a).

[0233] 45. The method according to Example 39, wherein the primer-template nucleic acid hybrid of step (a) is the product of polymerase extension during the sequencing process, wherein the sequencing process detects a signal indicating the sequence of the template at the 3' position of the series having at least two base multiple states.

[0234] 46. ​​The method according to Example 45 further includes a computer-aided process of aligning the sequence of the template with a region of a reference genome, wherein the alignment is performed by aligning the region of the reference genome with the sequence of the template at the 3' position of the series having at least two base multiples.

[0235] 47. The method according to Examples 1 to 46, wherein the sequencing process includes detecting a stable ternary complex at each location of the template, detecting a labeled nucleotide incorporated into a primer at each location of the template, detecting a labeled oligonucleotide conjugated to the primer, detecting pyrophosphate generated at each location of the template due to a nucleotide incorporated into the primer, or detecting a proton generated at each location of the template due to a nucleotide incorporated into the primer.

[0236] 48. The method according to Example 47, further comprising (h) polymerase extension of the extended primer during a second sequencing process, wherein the second sequencing process detects a signal indicating a sequence of the template at the 5' position of the series having at least two base multiple states.

[0237] 49. The method according to Example 48, further comprising a computer-aided process of aligning the sequence of the template with a region of the reference genome, wherein the alignment is performed by aligning the region of the reference genome with the sequence of the template at position 3' of the sequence having at least two base multiples, the sequence having at least two base multiples, and the sequence of the template at position 5' of the sequence having at least two base multiples.

[0238] 50. The method according to any one of Examples 21 to 49, wherein the nucleotide homologs of the first, second, third and fourth different base types include the exogenous markers detected in step (d).

[0239] 51. The method according to any one of Examples 21 to 49, wherein the nucleotide homologs of the first, second, third and fourth different base types do not include the exogenous markers detected in step (d).

[0240] 52. The method according to any one of Examples 21 to 51, wherein the polymerase includes the exogenous marker detected in step (d).

[0241] 53. The method according to any one of Examples 36 to 52, wherein step (b) is performed after the mixture has been removed from contact with the extended primer hybrid.

[0242] 54. The method according to any one of Examples 36 to 53, wherein step (c) is performed after the removal of the nucleotide homolog of the fourth base type among the four different base types.

[0243] 55. A method for characterizing a template nucleic acid, wherein the template includes a linker region adjacent to first and second regions of the template nucleic acid and between the first and second regions, the method comprising the steps of:

[0244] (a) Obtaining a single-base resolution sequence of the first region by extending primers along the first region of the template nucleic acid, thereby generating an extended primer-template hybrid;

[0245] (b) The feature markings of the joint region are obtained by:

[0246] (i) The extended primer-template hybrid is further extended using a nucleotide mixture, wherein the nucleotide mixture comprises nucleotide homologs of first, second, third, and fourth different base types, wherein the first base type of nucleotide homolog comprises a reversible terminator, and wherein at least one of the second, third, or fourth base type of nucleotide homolog is extendable, and

[0247] (ii) Detection of a stable ternary complex comprising a further extended primer-template hybrid, a polymerase, and a nucleotide homolog of the next base in the template, wherein the signature includes a base multiplicity comprising the first base type followed by the next base; and

[0248] (c) Obtain the single-base resolution sequence of the second region by extending the further extended primer-template hybrid.

[0249] 56. A method for characterizing nucleic acids, wherein a template includes a linker region adjacent to first and second regions of the template nucleic acid and between the first and second regions, the method comprising the steps of:

[0250] (a) Obtaining a single-base resolution sequence of the first region by extending primers along the first region of the template nucleic acid, thereby generating an extended primer-template hybrid;

[0251] (b) The feature markings of the joint region are obtained by:

[0252] (i) The extended primer-template hybrid is further extended using a nucleotide mixture, wherein the nucleotide mixture comprises nucleotide homologs of no more than three of four different base types.

[0253] (ii) In the absence of the nucleotide homologs of step (i), further extend the extended primer-template hybrid of step (i) using the nucleotide homologs of the fourth base type among the four different base types.

[0254] (iii) Detection of a stable ternary complex comprising a further extended primer-template hybrid from step (ii), a polymerase, and a nucleotide homolog of the next base in the template, wherein the signature includes a base multiplicity comprising the fourth base type of the four different base types, followed by the next base; and

[0255] (c) Obtain the single-base resolution sequence of the second region by extending the further extended primer-template hybrid of the further extended primer in step (iii).

[0256] 57. A method for characterizing nucleic acids, wherein a template includes a linker region adjacent to first and second regions of the template nucleic acid and between the first and second regions, the method comprising the steps of:

[0257] (a) Obtaining a single-base resolution sequence of the first region by extending primers along the first region of the template nucleic acid, thereby generating an extended primer-template hybrid;

[0258] (b) Obtaining a characteristic marker of the linker region by further extending the extended primer without distinguishing the type of nucleotides incorporated into the extended primer-template hybrid, wherein the characteristic marker includes a count of nucleotides in the linker region; and

[0259] (c) Obtain the single-base resolution sequence of the second region by extending the primer-template hybrid of the extended primer in step (b).

[0260] 58. The method according to Example 57, wherein step (b) is performed without the use of labeled nucleotides and labeled polymerase.

[0261] 59. The method according to Example 57, wherein step (b) is performed without using a label to distinguish the different types of nucleotides incorporated into the extended primer-template hybrid.

[0262] 60. The method according to Example 57, wherein step (b) is performed without detecting the extended primer-template hybrid.

[0263] Example 1

[0264] Mixed extension was performed using three extendable nucleotides and one reversibly terminated nucleotide.

[0265] This example describes a method for identifying four dinucleotide series of a template nucleic acid. One or more dinucleotide series can be used as characteristic markers of the template. The four dinucleotide series can be aligned to determine the template sequence at single nucleotide resolution.

[0266] The template nucleic acid is attached to the inner surface of an optically transparent flow cell. Primers are then hybridized to the template nucleic acid. The primers contain a 3' reversible terminator motif, which has been introduced using a reversibly terminated nucleotide, for example, through chemical synthesis or polymerase-catalyzed elongation. The primers are intended to be elongated after deblocking the reversible terminator motif so that the following nucleotide sequence is added to the primers:

[0267] 5'-TAGCCATCTGACTAACCTACTGTTT-3'(SEQ ID NO:3)

[0268] The resulting nucleic acid template undergoes the following Sequencing By Binding TM program:

[0269] (1) Optional testing of the first nucleotide—Under certain conditions, a test solution consisting of polymerase and four labeled nucleotides (dATP, dCTP, dGTP, and dTTP) is introduced into a flow cell to form a stable ternary complex consisting of a primer-template hybrid, polymerase, and the next correct nucleotide. The ternary complex is then tested under certain conditions to distinguish the signals generated by each nucleotide in the complex during its participation. The next correct nucleotide generates a signal indicating dTTP binding, and thus the first nucleotide of the elongation product is labeled T.

[0270] (2) Deblocking—The test solution is removed from the flow cell and replaced with a deblocking reagent that removes the reversible terminator portion from the primer, thus obtaining an extendable primer.

[0271] (3) First run extension – Remove from the flow cell and use a solution containing polymerase and three extendable nucleotides (dATP, dCTP, and dGTP) and a reversible termination nucleotide. rt The extension mixture of dTTP was used instead of the blocking solution. No labeled nucleotides were used. The flow cell was incubated under conditions allowing primer extension until the next primer was incorporated. rt dTTP.

[0272] (4) First run check – Under certain conditions, the test solution is introduced into the flow cell to form a stable ternary complex, and the ternary complex is tested to determine… rt The identity of the nucleotide following dTTP. The dinucleotide is labeled TN, where T is the incorporated element. rt dTTP and N is any one of A, C, T or G. In one configuration, the extension step is fluidly separated from the inspection step by replacing the extension mixture with the inspection solution. Alternatively, the extension mixture and the inspection solution may be present together in the flow cell (e.g., by co-delivery or additive delivery, such that steps (3) and (4) are effectively combined).

[0273] (5) Repeat steps (2) to (4) until the signal-to-noise decay indicator completes the first run—the result is a dinucleotide series [TA][TC][TG][TA][TA][TG][TT][TT], where the number and type of nucleotides between each dinucleotide are clearly ambiguous.

[0274] (6) Stripping the first run extension product – treating the flow cell with a chemical denaturant and / or heat to remove the extension primers from the template nucleic acid.

[0275] (7) Reapply primers—incubate the flow cell with primers having the same sequence and reversible terminator as previously used, thereby reforming the primer-template nucleic acid hybrid.

[0276] (8) Optionally, re-detect the first nucleotide as described in step (1).

[0277] (9) Perform the closing operation as in step (2).

[0278] (10) Second run extension—removal from the flow cell and use of a polymerase and three extendable nucleotides (dCTP, dGTP, and dTTP) and a reversible termination nucleotide. rt The extension mixture of dATP was used instead of the blocking solution. No labeled nucleotides were used. The flow cell was incubated under conditions allowing primer extension until the next primer was incorporated. rt dATP.

[0279] (11) Second run check – Under certain conditions, the test solution is introduced into the flow cell to form a stable ternary complex, and the ternary complex is tested to determine… rt The identity of the nucleotide following dATP. Dinucleotides are labeled AN, where A is the incorporated nucleotide. rt dTTP and N is any one of A, C, T, or G. Again, the extension step can be fluidly separated from or overlapped with the inspection step.

[0280] (12) Repeat steps (9) to (11) until the signal-to-noise decay indicator indicates that the second run is complete—the result is a dinucleotide series [AG][AT][AC][AA][AC][AC], in which the number and type of nucleotides between each dinucleotide are clearly ambiguous.

[0281] (13) Repeat steps (6) to (12), using a polymerase, three extendable nucleotides (dGTP, dTTP, and dATP), and a reversibly terminated nucleotide. rtThe third run extension mixture of dCTP was used instead of the second run extension mixture. The result was a dinucleotide series [CC][CA][CT][CT][CC][CT][CT], where the number and type of nucleotides between each dinucleotide were clearly ambiguous.

[0282] (14) Repeat steps (6) to (12), using a nucleotide containing polymerase, three extendable nucleotides (dTTP, dATP, and dCTP), and a reversible termination nucleotide. rt The fourth run extension mixture of dGTP was used instead of the second run extension mixture. The result was a dinucleotide series [GC][GA][GT], where the number and type of nucleotides between each dinucleotide were clearly ambiguous.

[0283] Table 1 shows the extension mixtures used in each of the four runs and the dinucleotide series determined based on the runs.

[0284] Table 1

[0285] Run Extension Mix Dinucleotide Series 1 A / C / G rt T T[TA][TC][TG][TA][TA][TG][TT][TT] 2 C / G / T rt A]] T[AG][AT][AC][AA][AC][AC] 3 G / T / A rt C]] T[CC][CA][CT][CT][CC][CT][CT] 4 T / A / C rt G]] T[GC][GA][GT]

[0286] One or more dinucleotide series from the dinucleotide series shown in Table 1 can be used as characteristic markers for templates. When the background sequence complexity in the desired sample is low, relatively short dinucleotide series may be sufficient as unique characteristic markers for identifying templates in the sample. With increasing sequence complexity, longer series can be more discriminative.

[0287] Furthermore, four dinucleotide series can be combined to determine the template sequence at single nucleotide resolution. Conceptually, the sequence can be determined by aligning dinucleotides. Alignment can begin with the first nucleotide T, as determined according to the first examination step (or as otherwise known). The second nucleotide in the sequence is understood to follow T and can therefore be determined by extending into... rt The run terminated at T (i.e., the first run) is used to determine this. In this case, the second nucleotide is A. The third nucleotide in the sequence is understood to follow A and can therefore be determined based on the extension at T. rt The run that terminates at point A (i.e., the second run) is determined. The process can continue in this manner and is visually presented in the comparison in Table 2.

[0288] Table 2

[0289]

[0290] Example 2

[0291] Mixed extension using only 3 nucleotides

[0292] The primer-template nucleic acid hybrid is attached to the inner surface of an optically transparent flow cell as described in Example 1. The primer-template hybrid has the same sequence as indicated in Example 1, and the primer optionally has a 3' reversible terminator portion at its 3' end.

[0293] The resulting nucleic acid template undergoes the following Sequencing By Binding TM program:

[0294] (1) Optional testing of the first nucleotide—Under certain conditions, a test solution consisting of polymerase and four labeled nucleotides (dATP, dCTP, dGTP, and dTTP) is introduced into a flow cell to form a stable ternary complex consisting of a primer-template hybrid, polymerase, and the next correct nucleotide. The ternary complex is stabilized by (a) a reversible terminator portion at the 3' end of the primer, (b) a polymerase inhibitor (such as a non-catalytic metal), (c) the absence of a catalytic metal, and / or (d) a polymerase mutant whose ability to extend the primer is inhibited. The ternary complex is tested under certain conditions to distinguish the signals generated by each nucleotide in the nucleotides when participating in the ternary complex. The next correct nucleotide generates a signal indicating dTTP binding, and thus the first nucleotide of the extension product is labeled T.

[0295] (2) Activation – The test solution is removed from the flow cell. The primer-template hybrid is activated by (a) deblocking the reversible terminator on the primer, (b) removing the polymerase inhibitor, (c) adding a catalytic metal ion, and / or (d) adding a catalytically active polymerase. Primers are then prepared for extension.

[0296] (3) First run of single nucleotide extension – will include unlabeled dTTP (or rt A mononucleotide extension solution of dTTP and polymerase was added to the flow cell. Primers were allowed to pass through the flow cell after adding the previously omitted dTTP (or...). rt The flow cell is incubated under extended conditions (dTTP). Then, deactivation is performed as in step (2).

[0297] (4) First run of mixed extension—Remove and replace the mononucleotide extension mixture with an extension mixture containing polymerase and three extendable nucleotides (dATP, dCTP, and dGTP). No labeled nucleotides. Incubate the flow cell under conditions that allow primer extension until the next correct nucleotide is the lost dTTP.

[0298] (5) First run check – Under certain conditions, the test solution is introduced into the flow cell to form a stable ternary complex, and the ternary complex is detected to determine the identity of the nucleotide following dTTP. The dinucleotide is labeled TN, where N is the nucleotide following T.

[0299] (6) Repeat steps (2) to (5) until the signal-to-noise decay indicator completes the first run—the result is a dinucleotide series [TA][TC][TG][TA][TA][TG][TT][TT], where the number and type of nucleotides between each dinucleotide are clearly ambiguous.

[0300] (7) Stripping the first run extension product – treating the flow cell with a chemical denaturant and / or heat to remove the extension primers from the template nucleic acid.

[0301] (8) Reapply primers—incubate the flow cell with primers having the same sequence as previously used to reform the primer-template nucleic acid hybrid.

[0302] (9) Optionally, re-detect the first nucleotide as described in step (1).

[0303] (10) Perform activation as in step (2).

[0304] (11) Second run of mixed extension—Add the extension mixture containing polymerase and three extendable nucleotides (CTP, GTP, and TTP) to the flow cell. No labeled nucleotides. Incubate the flow cell under conditions that allow primer extension until the next correct nucleotide is the lost ATP.

[0305] (12) Second run of single nucleotide extension—removal and with unlabeled ATP (or rt Replace the extension mixture with a mononucleotide extension solution containing ATP and polymerase. Allow primers to be added by adding previously omitted ATP (or...) rt The flow cell was incubated under extended conditions (ATP).

[0306] (13) Second run check – Under certain conditions, the test solution is introduced into the flow cell to form a stable ternary complex, and the ternary complex is detected to determine the identity of the nucleotide following ATP. The dinucleotide is labeled AN, where N is any one of A, C, T, or G.

[0307] (14) Repeat steps (10) to (13) until the signal-to-noise decay indicator indicates that the second run is complete—the result is a dinucleotide series [AG][AT][AC][AA][AC][AC], in which the number and type of nucleotides between each dinucleotide are clearly ambiguous.

[0308] (15) Repeat steps (7) through (14), replacing the second run extension mixture with a third run extension mixture containing polymerase and three extendable nucleotides (GTP, TTP, and ATP). The result is a dinucleotide series [CC][CA][CT][CT][CC][CT][CT], where the number and type of nucleotides between each dinucleotide are clearly ambiguous.

[0309] (16) Repeat steps (7) through (14), replacing the second run extension mixture with a fourth run extension mixture containing polymerase and three extendable nucleotides (dTTP, dATP, and dCTP). The result is a dinucleotide series [GC][GA][GT], where the number and type of nucleotides between each dinucleotide are clearly ambiguous.

[0310] Table 3 shows the extension mixtures used in each of the four runs and the dinucleotide series determined based on the runs.

[0311] Table 3

[0312] Run Extension Mix Dinucleotide Series 1 A / C / G T[TA][TC][TG][TA][TA][TG][TT][TT] 2 C / G / T T[AG][AT][AC][AA][AC][AC] 3 G / T / A T[CC][CA][CT][CT][CC][CT][CT] 4 T / A / C T[GC][GA][GT]

[0313] One or more dinucleotide series from the dinucleotide series shown in Table 3 can be used as characteristic markers for the template. Furthermore, four dinucleotide series can be combined to determine the template sequence at single nucleotide resolution using an alignment method similar to that described above in Example 1.

[0314] Example 3

[0315] Sequencing was performed using a single nucleotide resolution method coupled with dinucleotide signatures.

[0316] This example describes a method for identifying template sequences by determining the sequence of the first region of the template at single-nucleotide resolution and identifying a series of characteristic marker dinucleotides for the adjacent tail regions of the template. The characteristic marker tail regions can help align the template with a reference genome and can provide information about the long-range structure of the genome.

[0317] Multiple template nucleic acids are attached to the inner surface of an optically transparent flow cell. Primers are hybridized to the template nucleic acids. The primers contain a 3' reversible terminator motif (e.g., the terminator motif has been introduced via chemical synthesis or polymerase-catalyzed elongation using a reversibly terminated nucleotide).

[0318] The resulting nucleic acid template undergoes the following Sequencing By Binding TM program:

[0319] (A) Single nucleotide resolution SBB of the first region of the template TM The procedure—checking, deblocking, and single nucleotide extension—is performed as follows:

[0320] (1) Check – Under certain conditions, a check solution consisting of polymerase and four labeled nucleotides (dATP, dCTP, dGTP, and dTTP) is introduced into a flow cell to form a stable ternary complex consisting of each primer-template hybrid with polymerase and the next correct nucleotide. The ternary complex is then detected under certain conditions to distinguish the signal generated by each nucleotide in the ternary complex when it participates in the process.

[0321] (2) Deblocking—The test solution is removed from the flow cell and replaced with a deblocking reagent that removes the reversible terminator portion from the primer, thus obtaining an extendable primer.

[0322] (3) Single nucleotide extension—removing the blocking solution from the flow cell and reversibly terminating the four nucleotides ( rt dTTP, rt dATP, rt dCTP and rt dGTP is introduced into the flow cell. No nucleotides are labeled for reversible termination. The flow cell is incubated under conditions that allow the next correct nucleotide to be added to the primers.

[0323] (4) Repeat steps (1) to (3) 99 times to obtain 100 base reads covering the first end region of the template and generate primer extension products hybridized to the first region.

[0324] (B) Determine the characteristic marker dinucleotide series for the second region of the template—perform steps (1) through (5) of Example 1 using primer extension products hybridized to the first region of the template. Repeat the steps as desired or until the signal-to-noise ratio decays to the point indicating the run is complete.

[0325] The sequencing procedure described above produces a collection of reads containing two regions: single-nucleotide resolution regions and dinucleotide sequence marker regions. Reads can be aligned to a reference genome based on the juxtaposition of these two regions. Although the resolution is lower, the dinucleotide marker regions of each read can help confirm alignment with the single-nucleotide resolution regions. The dinucleotide marker regions of each read can also aid in phasing, for example, when single nucleotide polymorphisms (SNPs) are distinguished by the markers and when the read is long enough to cover SNPs that form haplotypes.

[0326] Example 4

[0327] Paired read sequencing was performed using a single nucleotide resolution method coupled with dinucleotide signatures.

[0328] This example describes a method for identifying a template sequence by determining the sequences of two regions of the template at single nucleotide resolution and identifying a series of characteristic marker dinucleotides for the insertion region of the template. The resulting reads can be aligned using known paired read methods (also known in the art as “paired-end methods”) that have the additional benefit of using the characteristic marker regions to confirm the alignment.

[0329] Multiple template nucleic acids are attached to the inner surface of an optically transparent flow cell. Primers are hybridized to the template nucleic acids. The primers contain a 3' reversible terminator motif (e.g., the terminator motif has been introduced via chemical synthesis or polymerase-catalyzed elongation using a reversibly terminated nucleotide).

[0330] The resulting nucleic acid template undergoes the following Sequencing By Binding TM program:

[0331] (A) Single nucleotide resolution SBB of the first region of the template TM The procedure—performed as described in step (A) of Example 3—involves checking, deblocking, and single nucleotide extension. The procedure identifies a 100-base read covering the first end region of the template and generates primer extension products hybridized to the first region.

[0332] (B) Determining the characteristic dinucleotide sequence of the template insertion region—Perform step (B) of Example 1 using the primer extension product hybridized to the first region of the template. Repeat the steps as desired or until the signal-to-noise ratio decays to a predefined point. The result of steps (A) and (B) is a set of reads containing two regions: a single nucleotide resolution region (i.e., the first-end region) and a dinucleotide sequence characteristic region (i.e., the insertion region). The product is the primer extension product hybridized to the first and second regions.

[0333] (C) Single nucleotide resolution SBB of the second-terminal region of the template TM Procedure—Step (A) is performed using primer extension products hybridized to the first and second regions of the template.

[0334] Steps (A) through (C) result in a collection of reads comprising three regions: a first single nucleotide resolution region (first end region), a dinucleotide sequence marker region (insertion region), and a second single nucleotide resolution region (second end region). Reads can be aligned to a reference genome based on the proximity of the first and second end regions. Although the resolution is lower, the insertion region of each read can help confirm alignment of the single nucleotide resolution end regions. The dinucleotide sequence marker region of each read can also aid in phasing, for example, when the length of the region spanning the read is sufficiently long to cover the SNPs forming the haplotype.

[0335] Example 5

[0336] Mixed extension was performed using two extendable nucleotides and two reversibly terminated nucleotides.

[0337] This example describes a method for determining the dinucleotide sequence of a template nucleic acid. Individual or combined dinucleotide sequences can be used as characteristic markers of the template. Two or more dinucleotide sequences can be compared to determine the template sequence at single-nucleotide resolution. Furthermore, a third dinucleotide sequence can provide error checking capability.

[0338] The template nucleic acid is attached to the inner surface of an optically transparent flow cell. Primers are hybridized to the template nucleic acid. The primers contain a 3' reversible terminator motif (e.g., the terminator motif has been introduced via chemical synthesis or polymerase-catalyzed extension using a reversibly terminated nucleotide). The primers are intended to be extended after deblocking the reversible terminator motif so that the following nucleotide sequence is added to the primers:

[0339] 5'-AAATGCATTGGCAGTGTA-3'(SEQ ID NO:4)

[0340] The resulting nucleic acid template undergoes the following Sequencing By Binding TM program:

[0341] (1) First run check—Under certain conditions, a check solution consisting of polymerase and four labeled nucleotides (dATP, dCTP, dGTP, and dTTP) is introduced into the flow cell to form a stable ternary complex consisting of a primer-template hybrid, polymerase, and the next correct nucleotide. The ternary complex is then tested under certain conditions to distinguish the signals generated by each nucleotide in the complex during its participation. The next correct nucleotide generates a signal indicating dATP binding, and therefore the first nucleotide of the extension product is labeled A.

[0342] (2) First run deblocking - Remove the test solution from the flow cell and replace it with a deblocking reagent that removes the reversible terminator portion from the primers to obtain extendable primers.

[0343] (3) First run extension – Remove from the flow cell and run with polymerase, two extendable nucleotides (dATP and dTTP), and two reversibly terminated nucleotides. rt dCTP and rt The extension mixture of dGTP was used instead of the blocking solution. No labeled nucleotides were used. The flow cell was incubated under conditions allowing primer extension until the next primer was incorporated. rt dCTP or rt dGTP. In one configuration, the extension step is fluidly separated from the inspection step by replacing the extension mixture with an inspection solution. Alternatively, the extension mixture and the inspection solution may be present together in the flow cell (e.g., by co-delivery or additive delivery).

[0344] (4) Repeat steps (1) through (3) until the signal-to-noise decay indicator indicates the first run is complete—the result is a dinucleotide sequence [SC][SA][SG][SC][SA][ST][ST], where S = C or G and the number and type of nucleotides between each dinucleotide are clearly ambiguous. The dinucleotide sequence can serve as a characterizing marker for the template. Optionally, additional characterizing information can be obtained by performing the following steps on the template nucleic acid.

[0345] (5) Stripping the first run extension product – treating the flow cell with a chemical denaturant and / or heat to remove the extension primers from the template nucleic acid.

[0346] (6) Reapply primers—incubate the flow cell with primers having the same sequence and reversible terminator as previously used, thereby reforming the primer-template nucleic acid hybrid.

[0347] (7) Second run check – detect the next correct nucleotide as described for step (1).

[0348] (8) Perform the de-closing as in step (2).

[0349] (9) Second run extension – Remove from the flow cell and use a solution containing polymerase and two extendable nucleotides (dCTP and dGTP) and two reversibly terminated nucleotides ( rt dATP and rt The extension mixture of dTTP was used instead of the blocking solution. No labeled nucleotides were used. The flow cell was incubated under conditions allowing primer extension until the next primer was incorporated. rt dATP or rt dTTP.

[0350] (10) Repeat steps (7) to (9) until the signal-to-noise decay indicator completes the second run—the result is a dinucleotide series [WA][WA][WT][WG][WT][WT][WG][WG][WG][WA], where W = A or T and the number and type of nucleotides between each dinucleotide are clearly ambiguous.

[0351] (11) Repeat steps (5) to (10), using a polymerase, two extendable nucleotides (dTTP and dCTP), and two reversibly terminated nucleotides. rt dATP and rt The third run extension mixture of dGTP was used instead of the second run extension mixture. The result was a dinucleotide series [RA][RA][RT][RC][RT][RG][RC][RG][RT][RT], where R = A or G and the number and type of nucleotides between each dinucleotide were clearly ambiguous.

[0352] Table 4 shows the extension mixtures used in each of the three runs and the dinucleotide series determined based on the runs.

[0353] Table 4

[0354]

[0355] One or more dinucleotide series from the dinucleotide series shown in Table 4 can be used as signature markers for the template. When the background sequence complexity in the desired sample is low, a relatively short dinucleotide series may be sufficient as a unique signature marker for the template. In the context of increasing sequence complexity, longer series can have greater discriminative power. The signature can consist of two or three dinucleotide series.

[0356] Furthermore, the first and second dinucleotide series are orthogonal to each other and can be combined to determine the template sequence at single nucleotide resolution. Conceptually, the sequence can be determined by aligning dinucleotides. Alignment can begin with the first nucleotide A, as determined according to the first examination step (or as otherwise known). The second nucleotide in the sequence is understood to follow A and can therefore be determined by extending into... rt The run that terminates at point A (i.e., the second run) is used to determine the second nucleotide. In this case, the second nucleotide is A. The third nucleotide in the sequence is understood to be after A and can therefore be determined based on the extension at point A. rt The run that terminates at point A (i.e., the second run) is determined. The process can continue in this manner and is visually presented in the comparison in Table 5.

[0357] Table 5

[0358]

[0359] The third dinucleotide series can serve as an error check. As shown in Table 6, the dinucleotides from the third run correctly aligned with the sequences identified according to Table 5. Specifically, all A and G bases in the sequence aligned with RN dinucleotides (N being any one of the four nucleotides) and all RN dinucleotides aligned with the sequence. Therefore, the information from the third run indicates that the sequences identified according to the first and second runs are correct.

[0360] Table 6

[0361]

[0362] As explained above, sequences determined from two sequencing runs using orthogonal components with two reversible termination nucleotides and two extension nucleotides can be error-checked using the results of a run performed with a third, different extension mixture. The third run will identify errors associated with three types of nucleotides. One type of nucleotide may be discarded, including sequence reads in which errors have already been identified (e.g., a mismatch between the first and third reads). If desired, the sequence can be further modified using a third extension mixture (i.e., A / ). rt T / rt A fourth run is performed using C / G orthogonal extension mixtures. The combination of all four runs allows for the detection of all random errors. Furthermore, across all four runs, each nucleotide is reversibly terminated in two different extension mixtures, providing sufficient information to correct errors involving all four nucleotide types.

[0363] The error detection strategy described above (three reads with an orthogonal component of two native nucleotides and two blocking nucleotides) allows for the detection of certain (but not all) classes of errors. Additionally, nucleotide mixtures present in the checking step (as described in commonly owned U.S. Patent Application Serial No. 15 / 712,632 and U.S. Patent Application No. 9,951,385, each of which is incorporated herein by reference) can detect different classes of errors, thereby complementing the original strategy and further improving sequencing accuracy by filtering out erroneous reads.

[0364] Example 6

[0365] Mixed extension on octagonal instrument

[0366] This example demonstrates an SBB using a mixture of three extendable nucleotides and one reversibly terminated nucleotide during the extension phase. TMThe feasibility of the procedure in distinguishing products compared to methods that use four reversibly terminated nucleotides during extension or four extendable nucleotides during extension.

[0367] The bonding reaction at the surface of the fiber tip was measured using biolayer interferometry in the form of a porous plate. (Menlo Park, California) Octagonal instrument. A primer-template nucleic acid hybrid with a biotinylated 5' end of the template strand is immobilized onto the tip of an optical fiber functionalized with streptavidin (SA) using a standard procedure. The primer contains a fluorescent label at the 5' end. The Tough-22 primer-template nucleic acid hybrid in this procedure has… Figure 1 The sequence shown is a template sequence to be read downstream of the primer. Figure 2 The template sequence to be read by the primer-template nucleic acid hybrid for Tough-29 is shown in the figure.

[0368] The loop procedure involves the following steps:

[0369] a) Equilibrate for 5 seconds in a solution containing all the components required for incorporating the reversible terminator (step b) except for the reversible terminator nucleotide, magnesium chloride, and DNA polymerase.

[0370] b) Using primers for 30 seconds with a reversible terminator for hybridization under the conditions described in Hutter et al., Nucleosides, Nucleotides and Nucleic Acids 29:879-895 (2010) (DOI:10.1080 / 15257770.2010.536191), which is incorporated herein by reference.

[0371] c) Equilibrate in the solution for 5 seconds as in step a).

[0372] d) Cut the end portion of the reversible terminator under the conditions described by Hutter et al. (ibid.) for 5 seconds.

[0373] Repeat steps a) to d) 10 times. After this, the extended primers are eluted from the fiber tip in dimethylformamide and analyzed by an ABI 3500 capillary electrophoresis instrument.

[0374] Figure 1 The results of performing the above procedure for 10 cycles on the primer-template hybrid of Tough-22 are shown in the figure. Figure 1 Figure A shows the capillary electrophoresis (CE) traces of the primers. The goal is to extend the primers by 10 nucleotides using ten extension steps, each with four reversibly terminated nucleotides. (The text abruptly ends here, likely due to an incomplete sentence or missing information.) Figure 1The CE trace in B shows an extension approximately equivalent to 9.6 nucleotides. It is expected that dATP, dCTP, dGTP, and... rt The ten extension steps of dTTP are achieved using a 42-nucleotide extension primer. (As shown by...) Figure 1 The CE trace in C indicates an extension of approximately 39.7 nucleotides. It is expected that ten extension steps using four extendable nucleotides each, via a 58-nucleotide extension primer, will be employed. (As shown by...) Figure 1 The CE trace in D shows an extension approximately equivalent to 53.1 nucleotides.

[0375] Figure 2 The results of performing the above procedure for 10 cycles on the primer-template hybrid of Tough-29 are shown in the figure. Figure 2 Figure A shows the capillary electrophoresis (CE) traces of the primers. The goal is to extend the primers by 10 nucleotides using ten extension steps, each with four reversibly terminated nucleotides. (The text abruptly ends here, likely due to an incomplete sentence or missing information.) Figure 2 The CE trace in B shows an extension approximately equivalent to 10.0 nucleotides. It is expected that dATP, dCTP, dGTP, and... rt The ten extension steps of dTTP are achieved using a 31-nucleotide extension primer. (As shown by...) Figure 2 The CE trace in C indicates an extension of approximately 31.2 nucleotides. It is expected that ten extension steps using four extendable nucleotides each, with a 58-nucleotide extension primer, will be employed. (As shown by...) Figure 2 The CE trace in D shows an extension approximately equivalent to 55.5 nucleotides.

[0376] The measured primer extension length was determined to be within a reasonable error of the expected length. Therefore, the results demonstrate that a mixture of one reversibly terminated nucleotide and three extendable nucleotides can be used in SBB. TM The program uses this method to increase primer extension length compared to using only reversibly terminated nucleotides.

[0377] Example 7

[0378] Re-phase the nucleic acid sequence

[0379] Inefficient primer extension and cleavage of reversible terminator portions have a cumulative effect on sequencing protocols. Within multiple cycles, primer molecules fail to extend at least once during run accumulation and cause negative phasing within clusters or other integrations of the sequenced nucleic acids. Negative phasing dilutes the correct signal and amplifies the incorrect signal. This problem can be mitigated using the following sequencing protocols by recovering the phased signal and using the signal to coalesce in phase:

[0380] 1. Perform single nucleotide resolution sequencing for 100 cycles or until the signal shifts.

[0381] 2. Perform the following steps to phase-shift the signal:

[0382] a. Extension of a single cycle using a mixture of extendable first, second, and third nucleotide analogs and a reversibly terminated fourth nucleotide analog.

[0383] b. Extension of a single cycle using a mixture of an elongated analogue containing the first, second, and fourth natural nucleotides and a reversibly terminated analogue containing the third nucleotide.

[0384] c. Extension of a single cycle using a mixture of elongation analogs of the first, third, and fourth natural nucleotides and a reversibly terminated analog of the second nucleotide.

[0385] d. Extension of a single cycle using a mixture of elongation analogs of the second, third, and fourth natural nucleotides and reversibly terminated analogs of the first nucleotide.

[0386] 3. After phase shift, perform another 100 cycles of single nucleotide resolution sequencing or until the signal is phase shifted.

[0387] 4. Repeat steps 2 and 3 until the signal and signal-to-noise ratio drop to a level where bases cannot be identified with reasonable certainty.

[0388] Step 2 of this strategy allows the negatively phased sequence to "catch up" with and synchronize with the main "full-length" sequence. The effect of re-phasing step 2 is illustrated in Table 7 below using the following model sequences:

[0389] TGCCATCTCGAAAAAACATTTGGACTGCTCCGCTTCCTCCTGAGACTGAGCT(SEQ ID NO:5)

[0390] The first, second, third, and fourth nucleotides correspond to dA, dC, dG, and dT, respectively.

[0391] Table 7

[0392]

[0393] Example 8

[0394] Paired read sequencing was performed using a single nucleotide resolution method coupled with dark extension.

[0395] This example describes a method for identifying a template sequence by determining the sequences of first and second regions of the template at single nucleotide resolution and by determining the length of the adapter region of the template by counting extension steps. In this example, the adapter region is inserted into the first and second regions such that sequence reads from the first and second regions provide paired reads. In this exemplary method, a loop including a detection step is used to extend the primer through the first region, then a loop without a detection step is used to further extend the extended primer through the adapter region, and then a loop including a detection step is used to further extend the further extended primer through the second region.

[0396] Multiple template nucleic acids were attached to the inner surface of an optically transparent flow cell. Primers were then hybridized to the template nucleic acids. The primers contained a 3' reversible terminator.

[0397] The resulting nucleic acid template undergoes the following Sequencing By Binding TM program:

[0398] (A) Single nucleotide resolution SBB of the first region of the template TM The procedure—performed as described in step (A) of Example 3—involves checking, deblocking, and single nucleotide extension. The procedure identifies a 100-base read covering the first end region of the template and generates primer extension products hybridized to the first region.

[0399] (B) Determine the characteristic marker nucleotide count of the template insertion region—subject the primer extension product hybridized to the first region of the template after step (A) to a modified SBB. TM The procedure is looped. An improvement is to omit the checking step used in step (A), so that each loop includes deblocking and single nucleotide extension using the blocked nucleotides. The loops are repeated the desired number of times. The result of steps (A) and (B) is a set of reads containing two regions: a single nucleotide resolution region (i.e., the first region) and a length that includes the number of nucleotide positions of the linker region immediately adjacent to the first region. The product is the primer extension product hybridized to both the first region and the linker region.

[0400] (C) Single nucleotide resolution SBB of the second region of the template TM Procedure—The procedure in step (A) is performed using primer extension products hybridized to the first and second regions of the template after step (B) is completed.

[0401] Steps (A) through (C) result in a collection of reads comprising three regions: a first single nucleotide resolution region, a signature nucleotide count of the linker region, and a second single nucleotide resolution region. Reads can be aligned to a reference genome based on the proximity of the first and second end regions. The length of the linker region can inform the alignment of the first and second single nucleotide resolution regions by limiting the distance between reads when aligning with the reference genome. The signature region of each read can also aid in phasing, for example, when the length of the region spanning the read is sufficiently long to cover the SNPs forming the haplotype. Exemplary algorithms for assembling paired reads to determine genome sequences include, but are not limited to, those described in Hormozdiari et al., Combinatorial Algorithms for Structural Variation Detection in High-Throughput Sequenced Genomes, Genome Research 19:1270–1278 (2009), or Sindi et al., A geometric approach for classification and comparison of structural variants, Bioinformatics 25:i222–i230 (2009) (each of which is incorporated herein by reference) or other algorithms known in the art.

[0402] The advantage of the paired read technique illustrated in this example is that it avoids detection steps that adversely affect sequencing accuracy and read length. Although not intended to be mechanistically limited, it achieves single-nucleotide resolution SBB. TM Excitation of fluorophores during the fluorescence detection step used in the procedure is believed to cause photodamage to the sequenced nucleic acids and the reagents used for sequencing. Because sequencing is a cumulative process, this photodamage not only affects the accuracy of individual base reads within the sequence reads, but also shortens the read length due to the cumulative effect of signal attenuation and / or errors accumulating during the run.

[0403] Although short reads from multiple genomic segments can be assembled to reconstruct segments derived from larger chromosomes, certain types of structural information cannot be definitively inferred from the assembled sequence. In diploid organisms such as humans, it can be difficult to determine whether polymorphisms observed in reads from two separate segments derive from the same or different parents. However, if two reads are linked by a pair of reads, as illustrated in this example, it can be determined that the reads originate from the same strand and therefore from the same parent. Thus, the pair-read approach described in this paper offers the advantage of identifying structural features of the genome (e.g., haplotypes) that are not readily available from single-read approaches.

[0404] Throughout this application, various publications, patents, and / or patent applications have been cited. The disclosures of these documents are incorporated herein by reference in their entirety.

[0405] Several embodiments have been described. However, it should be understood that various modifications can be made. Therefore, other embodiments are within the scope of the following claims. sequence list <110> Omniom Corporation <120> Compositions and techniques for nucleic acid primer extension <130> 053195-505001WO <150> 62 / 626,836 <151> 2018-02-06 <160> 8 <170> PatentIn version 3.5 <210> 1 <211> 58 <212> DNA <213> Artificial Sequence <220> <223> Synthetic polynucleotides <400> 1 tgagtcaaaa aaaaaaaaaa aaaaggctac tccttttctc ctgcttccaa ttttctga 58 <210> 2 <211> 58 <212> DNA <213> Artificial Sequence <220> <223> Synthetic polynucleotides <400> 2 cgaatgtgct gctgctgctg ctgctgctgc tgctgcggtc tcctgggaaa ggctctca 58 <210> 3 <211> 25 <212> DNA <213> Artificial Sequence <220> <223> Synthetic polynucleotides <400> 3 tagccatctg actaacctac tgttt 25 <210> 4 <211> 18 <212> DNA <213> Artificial Sequence <220> <223> Synthetic polynucleotides <400> 4 aaatgcattg gcagtgta 18 <210> 5 <211> 51 <212> DNA <213> Artificial Sequence <220> <223> Synthetic polynucleotides <400> 5 tgccatctcg aaaaacattt ggactgctcc gcttcctcct gagactgagc t 51 <210> 6 <211> 10 <212> DNA <213> Artificial Sequence <220> <223> Synthetic polynucleotides <400> 6 tgccatctcg 10 <210> 7 <211> 16 <212> DNA <213> Artificial Sequence <220> <223> Synthetic polynucleotides <400> 7 tgccatctcg aaaaac 16 <210> 8 <211> 17 <212> DNA <213> Artificial Sequence <220> <223> Synthetic polynucleotides <400> 8 tgccatctcg aaaaaca 17

Claims

1. A method for characterizing nucleic acids, the method comprising: (a) Under certain conditions, the primer-template nucleic acid hybrid is contacted with a mixture of polymerase and nucleotides to produce an extended primer hybrid. The mixture comprises nucleotide homologs of first, second, third and fourth different base types, wherein the nucleotide homolog of the first base type includes a reversible terminator, and wherein at least one of the nucleotide homologs of the second, third or fourth base types is elongated; (b) Contact the extended primer hybrid with a nucleotide homolog of at least one of the different base types and a polymerase to form a stable ternary complex comprising the extended primer hybrid, the polymerase and a nucleotide homolog of the next base in the template. (c) Detect the stable ternary complex to distinguish the next base from other base types in the template; (d) Determine the presence of a base multivariate in the template nucleic acid, the base multivariate comprising the first base type followed by the next base; (e) Repeat steps (a) to (c) using the extended primer hybrid as the primer-template nucleic acid hybrid. (f) Determine the presence of a series of nucleic acids having at least two base multiple states in the template nucleic acid; The nucleotide homologs of the second, third, and fourth base types are elongated in step (a); (g) Repeat steps (a) to (f), wherein the nucleotide homolog of the second base type includes a reversible terminator, and wherein the nucleotide homologs of the first, third, and fourth base types in the mixture are elongated, and wherein the base multivariate determined in step (c) includes the second base type, followed by the next base. This confirms the presence of two base multiple state series in the template nucleic acid; (h) Repeat steps (a) to (f), wherein the nucleotide homolog of the third base type includes a reversible terminator, and wherein the nucleotide homologs of the first, second, and fourth base types in the mixture are elongated, and wherein the base multivariate determined in step (c) includes the third base type, followed by the next base; and (i) Repeat steps (a) to (f), wherein the nucleotide homolog of the fourth base type includes a reversible terminator, and wherein the nucleotide homologs of the first, second, and third base types in the mixture are elongated, and wherein the base multivariate determined in step (c) includes the fourth base type, followed by the next base. This confirms the presence of a four-base multivariate sequence in the template nucleic acid; Steps (a) to (f) are repeated using the template nucleic acid after the extended primer is removed from the extended primer hybrid and the second primer is hybridized to the template.

2. The method of claim 1, further comprising a computer-aided process for comparing the four base multiple state series to determine the sequence of the nucleic acid.

3. The method of claim 1, wherein the nucleotide homologs of the first and second base types include reversible terminators, and wherein the nucleotide homologs of the third and fourth base types are elongated in step (a).

4. The method of claim 3, further comprising: (g) Repeat steps (a) to (f), wherein the nucleotide homologs of the third and fourth base types include reversible terminators, and wherein the nucleotide homologs of the first and second base types in the mixture are elongated.

5. The method of claim 4, further comprising: (h) Repeat steps (a) to (f), wherein the nucleotide homologs of the first and third base types include reversible terminators, and wherein the nucleotide homologs of the second and fourth base types in the mixture are elongable.

6. The method according to any one of claims 1 to 5, wherein steps (a) to (f) are repeated using a second template nucleic acid molecule whose sequence is identical to that of the template nucleic acid used in step (a).

7. The method of claim 1, wherein the primer-template nucleic acid hybrid of step (a) is the product of polymerase extension during the sequencing process, wherein the sequencing process detects a signal indicating the sequence of the template at the 3' position of the series having at least two base multiple states.

8. The method of claim 7, further comprising a computer-aided process of aligning the sequence of the template with a region of a reference genome, wherein the alignment is performed by aligning the region of the reference genome with the sequence of the template at the 3' position of the series having at least two base multiples.

9. The method of claim 8, further comprising (g) polymerase extension of the extended primers during a second sequencing process, wherein the second sequencing process detects a signal indicating a sequence of the template at the 5' position of the series having at least two base multiple states.

10. The method of claim 9, further comprising a computer-aided process of aligning the sequence of the template with a region of the reference genome, wherein the alignment is performed by aligning the region of the reference genome with the sequence of the template at position 3' of the sequence having at least two base multiples, the sequence having at least two base multiples, and the sequence of the template at position 5' of the sequence having at least two base multiples.

11. The method of claim 1, wherein the nucleotide homologs of the first, second, third and fourth different base types include the exogenous markers detected in step (c).

12. The method according to claim 1, wherein the nucleotide homologs of the first, second, third and fourth different base types do not include the exogenous markers detected in step (c).

13. The method of claim 1, wherein the polymerase comprises the exogenous marker detected in step (c).

14. The method of claim 1, wherein step (b) is performed after the mixture has been removed from contact with the extended primer hybrid.

Citation Information

Patent Citations

  • Method and system employing distinguishable polymerases for detecting ternary complexes and identifying cognate nucleotides

    US11248254B2

  • Method of nucleic acid sequencing

    US20020055100A1

  • Methods for detecting genome-wide sequence variations associated with a phenotype

    US20040002090A1

  • Isothermal amplification of nucleic acids on a solid support

    US20040096853A1

  • Diffraction grating-based encoded micro-particles for multiplexed experiments

    US20040125424A1