Sequencing from multiple primers to increase data rate and density

By using multiple primer hybridization and nucleotide analog extension of unique markers at different locations in the nucleic acid chain, the problem of insufficient sequencing data rate and density in the prior art is solved, and more efficient sequencing data acquisition is achieved.

CN113564238BActive Publication Date: 2025-08-05ILLUMINA CAMBRIDGE LTD
View PDF 57 Cites 0 Cited by

Patent Information

Application Number
CN202110829488.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2014-11-05
Filing Date
2015-11-04
Publication Date
2025-08-05
Estimated Expiration
2035-11-18

AI Technical Summary

Technical Problem

Existing next-generation sequencing technologies have limitations in improving the rate and density of sequencing data, especially when dealing with large genomes, which are difficult to meet research needs.

Method used

Multiple primers are used to hybridize at different positions in the nucleic acid strand and primer extension is performed using nucleotide analogs with unique markers, the identity of the nucleotide base is determined by signal data and assigned to the extended read.

Benefits of technology

It significantly improves the rate and density of sequencing data, enables the identification of multiple nucleotide bases on the same chain at the same time, and enhances the efficiency and accuracy of sequencing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure HDA0003174977580000011
    Figure HDA0003174977580000011
  • Figure HDA0003174977580000021
    Figure HDA0003174977580000021
  • Figure HDA0003174977580000031
    Figure HDA0003174977580000031
Patent Text Reader

Abstract

The present invention relates to a sequencing method that allows for increased sequencing rates and increased sequencing data density. The system can be based on a next generation sequencing method such as sequencing by synthesis (SBS), but uses multiple primers that bind at different positions of the same nucleotide chain.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the Chinese invention patent application filed on November 4, 2015, with application number 201580069430.1 and invention name “Sequencing from multiple primers to increase data rate and density”.

[0002] The present invention relates to sequencing methods that allow increased sequencing speed and increased sequencing data density.The system may be based on next generation sequencing methods such as sequencing by synthesis (SBS), but using multiple primers that bind at different positions of the same nucleotide chain.

[0003] Deciphering DNA sequences is crucial to nearly every branch of biological research. With the advent of Sanger sequencing, scientists gained the ability to elucidate genetic information from any given biological system. The technology has become widely used in laboratories around the world, but has been hampered by inherent limitations in throughput, scalability, speed, and resolution, which often prevent scientists from obtaining the essential information they need. To overcome these obstacles, next-generation sequencing (NGS) was developed, a fundamentally different sequencing approach that has led to many groundbreaking discoveries and sparked a revolution in genomics science.

[0004] Next-generation sequencing data output initially grew at a rate exceeding Moore's Law, more than doubling annually since its invention. In 2007, a single sequencing run could generate approximately one gigabase (Gb) of data. By 2011, this rate had reached nearly one terabase (Tb) of data in a single sequencing run, an increase of almost 1,000-fold in just four years. With the ability to rapidly generate large amounts of sequencing data, next-generation sequencing allows researchers to move quickly from an idea to a complete dataset in a matter of hours or days. Researchers can now sequence 16 human genomes in a single run, generating data in approximately three days, while reagent costs per genome continue to decline.

[0005] By comparison, it took about 10 years to sequence the first human genome using CE technology, with an additional three years to analyze it. The completed project was published in 2003, just a few years before next-generation sequencing was invented, and came with a price tag of nearly $3 billion.

[0006] While the latest high-throughput sequencing instruments are capable of large data outputs, next-generation sequencing technology is highly scalable. The same basic chemistry can be used for lower-yield targeted studies or smaller genomes. This scalability gives researchers the flexibility to design studies that best suit their specific research needs. To sequence small bacterial / viral genomes or target regions (such as exons), researchers can choose to use lower-output instruments and process a smaller number of samples per run, or they can choose to process a large number of samples by multiplexing on a high-throughput instrument. Multiplexing allows for simultaneous sequencing of a large number of samples during a single experiment.

[0007] While next generation sequencing has significantly increased throughput, it is still advantageous to increase sequencing data rate and density when, for example, processing large genomes. The present invention aims to further improve data rate and data density / throughput, particularly when applied to next generation sequencing / synthesis test (SBS) methods.

[0008] Preferably, the term "base calling" refers to the process of assigning bases (nucleobases) to information obtained during sequencing, for example by assigning nucleotides to chromatographic peaks.

[0009] As used herein, and unless otherwise indicated, each of the following terms shall have the following definition.

[0010] A--adenine;

[0011] C-cytosine;

[0012] ·DNA--deoxyribonucleic acid;

[0013] G--guanine;

[0014] RNA – ribonucleic acid;

[0015] T – thymine; and

[0016] U--Uracil.

[0017] "Nucleic acid" shall mean any nucleic acid molecule, including but not limited to DNA, RNA, and hybrids thereof. The nucleic acid bases that form the nucleic acid molecule may be the bases A, C, G, T, and U, and their derivatives. Derivatives of these bases are well known in the art and are exemplified in PCR systems, reagents, and consumables (Perkin Elmer Catalogue 1996-1997, Roche Molecular Systems, Inc., Branchburg, NJ, USA).

[0018] The "type" of a nucleotide refers to A, G, C, T or U.

[0019] "Mass tag" shall mean a molecular entity of predetermined size that is capable of being attached to another entity via a cleavable bond.

[0020] "Solid substrate" shall mean a medium present in a solid phase to which antibodies or reagents can be affixed.

[0021] Where a range of values is provided, it is understood that unless the context clearly dictates otherwise, each intervening value, to the tenth of the lower limit, between the upper and lower limits of that range, and any other specified or intervening value in that range, is encompassed within the invention. The upper and lower limits of these smaller ranges may independently be included in the smaller ranges and are also encompassed within the invention, subject to any specifically excluded limits in the stated range. Where the stated range includes one or both of the limits, ranges excluding one or both of those included limits are also encompassed within the invention.

[0022] One aspect of the present invention provides a method for determining a nucleic acid sequence, comprising

[0023] providing at least one nucleic acid bound to a support;

[0024] hybridizing at least two primers to the same strand of nucleic acid;

[0025] performing primer extension on the same strand of the nucleic acid at each of the at least two primers;

[0026] obtaining signal data corresponding to at least one nucleotide base incorporated during each primer extension;

[0027] The identities of the nucleic acid bases are determined from the signal data and the bases are assigned to extension reads.

[0028] Preferably, the method comprises contacting the nucleic acid with an enzyme under conditions permissive for hybridization in the presence of: (i) at least two primers capable of hybridizing to the same strand of the nucleic acid at different positions, and (ii) four labelled moieties comprising at least one nucleotide analogue selected from dGTP, dCTP, dTTP, dUTP, and dATP, each of the four labelled moieties having a unique label that is different from the unique labels of the other three labelled moieties to give a plurality of extended reads, each extended read extending from one of the at least two primers.

[0029] The method may comprise the step of removing unbound label moieties.

[0030] Preferably, the plurality of extended reads comprises substantially simultaneous extended reads.

[0031] In certain embodiments, the four labeling moieties each comprise a single reversible terminating nucleotide analog or reversible terminator analog selected from dGTP, dCTP, dTTP, dUTP, and dATP, or an oligonucleotide probe comprising at least one nucleotide analog selected from dGTP, dCTP, dTTP, dUTP, and dATP.

[0032] Advantageously, said signal data corresponding to each of said at least two primer extensions may be obtained substantially simultaneously.

[0033] This provides the advantage that the signal data contains data corresponding to multiple signals from the same nucleic acid strand.

[0034] Optionally, the signal data comprises one or more images.

[0035] Optionally, the signal data is detected as color signals corresponding to a plurality of nucleotide analogs incorporated upon extension of said primer, and wherein each color signal corresponds to a different nucleotide analog or combination of nucleotide analogs.

[0036] One aspect of the present invention provides a method for determining a nucleic acid sequence, comprising performing the following steps:

[0037] (a) providing at least one nucleic acid bound to a support;

[0038] (b) contacting the nucleic acid with an enzyme under conditions permissive for hybridization in the presence of: (i) at least two primers capable of hybridizing to the same strand of the nucleic acid at different positions, and (ii) four label moieties comprising at least one nucleotide analogue selected from dGTP, dCTP, dTTP, dUTP, and dATP, each of the four label moieties having a unique label that is different from the unique labels of the other three label moieties.

[0039] (c) removing unbound probe;

[0040] (d) determining the identity of the nucleic acid analogue incorporated in step (b) by determining the identity of the corresponding unique label, and

[0041] (e) Assigning nucleotide bases to extended reads.

[0042] Preferably, the plurality of extended reads comprises simultaneous extended reads.

[0043] In certain embodiments, the method comprises sequencing by synthesis and the enzyme comprises a polymerase.

[0044] Each of the four probes may comprise a single reversible terminator nucleotide derivative or reversible terminator analog selected from dGTP, dCTP, dTTP, dUTP, and dATP.

[0045] In certain embodiments, the method comprises sequencing by ligation and the enzyme comprises a ligase.

[0046] The four probes may comprise oligonucleotides.

[0047] The oligonucleotides may be of the type used for sequencing by ligation techniques.

[0048] The methods may include sequencing by synthesis, sequencing by ligation, pyrosequencing, or nanopore sequencing.

[0049] In certain embodiments, each unique marker is detectable as a different color.

[0050] Advantageously, because two or more primers hybridize to the same chain of nucleic acid at different locations, this allows the sequencing of the chain to begin at two different sites, to provide multiple extended reads, potentially multiplying the density of the data and the rate at which it is obtained. When the system uses nucleic acids in conjunction with solid supports (e.g., in next-generation sequencing), this allows multiple base calls to be made at each specific position. Each base call is assigned to the correct extended read, which allows the determination of the nucleic acid sequence.

[0051] In certain embodiments, this method offers the advantage that multiple nucleotide bases incorporated into a single template strand of nucleic acid at different positions on the same strand can be simultaneously identified in real time.

[0052] Preferably, the nucleic acid is deoxyribonucleic acid (DNA).

[0053] Preferably, the nucleic acid is bound to a solid substrate.

[0054] More preferably, the solid substrate is of the type used in sequencing by synthesis or sequencing by ligation methods.

[0055] Preferably, the solid substrate is a chip.

[0056] Preferably, the solid substrate is a bead.

[0057] Optionally, the unique label is bound to the base via a cleavable linker.

[0058] Optionally, the unique label is bound to the base via a chemically cleavable or photocleavable linker.

[0059] Optionally, the unique label is a dye, a fluorophore, a chromophore, a combined fluorescence energy transfer tag, a mass tag, or an electrophore.

[0060] Preferably, after (a) providing at least one support-bound nucleic acid; the nucleic acid is amplified.

[0061] More preferably, the nucleic acid is amplified by bridge amplification.

[0062] Optionally, at least two primers capable of hybridizing to the same strand of nucleic acid at different positions have overlapping sequences.

[0063] Optionally, the second primer is identical to the first primer with an extra base at the end.

[0064] Additional primers, each adding an extra base, may also be used.

[0065] Optionally, blocked and unblocked primers are used to chemically distinguish the extended reads. This ensures that the correct base calls are assigned to the correct extended reads.

[0066] This can be achieved by using different levels of blocked and unblocked primers for each of the at least two primers used.

[0067] Optionally, bioinformatics information is used to assign the bases to extended reads.

[0068] In certain embodiments, the step of determining the identity of the nucleotide analog comprises detecting signal data corresponding to the plurality of extended reads.

[0069] Signal data corresponding to each of the plurality of extended reads may be detected simultaneously.

[0070] The step of determining the identity of the nucleotide analogs may comprise analyzing a signal intensity profile corresponding to the unique marker detected at each extended read.

[0071] This may include simultaneously measuring signal intensity in the signal data corresponding to each extended read.

[0072] Optionally, the method comprises determining a distribution of intensity measurements in the signal data by generating a histogram of the intensity data.

[0073] The signal data may be detected as color signals corresponding to the plurality of nucleotide analogs incorporated in step (b), wherein each color signal corresponds to a different nucleotide analog or combination of nucleotide analogs.

[0074] The method may comprise a signal processing step comprising one or more of signal deconvolution, signal improvement and signal selection.

[0075] The method can include selecting the highest intensity color signal in each extended read to provide a base call.

[0076] Multiple base calls can be made during each hybridization cycle.

[0077] Optionally, the step of assigning bases to the extended reads comprises determining positions of nucleotide analogs.

[0078] Optionally, the step of assigning bases to the extended reads comprises providing a preliminary base call and a final base call for each extended read.

[0079] Optionally, final base calls are provided by comparing the preliminary base call data to a reference genome.

[0080] Optionally, final base calls are provided by comparing the preliminary base call data to a reference genome.

[0081] In certain embodiments, at least two primers may be overlapping primers.

[0082] In certain embodiments, at least two primers may differ by a single base addition.

[0083] Advantageously, in certain embodiments, the method provides for detecting a signal comprising signal data corresponding to multiple substantially simultaneous primer extensions at different sites on the same nucleic acid strand or molecule.

[0084] Advantageously, no more than four unique labels corresponding to mononucleotide monomer analogs of dGTP, dCTP, dTTP, dUTP and dATP or to oligonucleotide probes comprising analogs of dGTP, dCTP, dTTP, dUTP and dATP are required in this method.

[0085] The same four probe sets are preferably used in the method for each extended read.

[0086] According to the present invention, there is provided a method for determining a nucleic acid sequence, which comprises performing the following steps for each residue of the nucleic acid to be sequenced:

[0087] (a) providing at least one nucleic acid bound to a support;

[0088] (b) contacting the nucleic acid with a nucleic acid polymerase under conditions permissive for polymerase-catalyzed nucleic acid synthesis to give a plurality of extended reads, each extended read extending from a primer: (i) at least two primers capable of hybridizing to the same strand of the nucleic acid at different positions, and (ii) four reversibly terminating nucleic acid analogs or reversible terminators selected from nucleotide analogs of dGTP, dCTP, dTTP, dUTP, and dATP, each of the four analogs having a unique label that is different from the unique labels of the other three analogs.

[0089] (c) removing unbound nucleic acid analogs;

[0090] (d) determining the identity of the nucleic acid analogue incorporated in step (b) by determining the identity of the corresponding unique label, and

[0091] (e) Assigning the bases to extended reads.

[0092] Yet another aspect of the present invention provides a system for determining the sequence of a nucleic acid, the system comprising a sequencing instrument having a solid support for immobilizing at least one nucleic acid, and means for determining the sequence thereof by contacting the nucleic acid with an enzyme in the presence of: (i) at least two primers capable of hybridizing to the same strand of the nucleic acid at different positions and (ii) four probes comprising at least one nucleotide analogue selected from the group consisting of dGTP, dCTP, dTTP, dUTP and dATP,

[0093] Each of the four probes has a unique label that is different from the unique labels of the other three probes.

[0094] under conditions permissive for hybridization to give a plurality of extended reads, each extended read being extended from one of the at least two primers;

[0095] Remove unbound probe;

[0096] determining the identity of the incorporated nucleic acid analog by determining the identity of the corresponding unique label, and

[0097] Assign nucleotide bases to extended reads.

[0098] Another aspect of the present invention provides a sequencing method comprising:

[0099] providing at least one nucleic acid hybridized to a support;

[0100] hybridizing at least two primers to the same strand of the nucleic acid to provide two extended reads;

[0101] Extension of each primer is performed with a reversible terminator base;

[0102] An image of the bound reversible terminator base is acquired for each base added.

[0103] A plurality of bases present at positions on the support are determined from the image and assigned to extended reads.

[0104] Yet another aspect of the present invention provides a system for determining a nucleic acid sequence, comprising a sequencing device having a solid support for immobilizing at least one nucleic acid and means for:

[0105] providing at least one nucleic acid bound to a support;

[0106] hybridizing at least two primers to the same strand of nucleic acid;

[0107] performing at least two primer extensions on the same strand of the nucleic acid with each of the at least two primers;

[0108] obtaining signal data corresponding to at least one nucleotide base upon incorporation in each of at least two primer extensions;

[0109] The identities of the nucleic acid bases are determined from the signal data and the bases are assigned to extension reads.

[0110] Preferably, the system comprises a signal processor for processing the signal data.

[0111] Yet another aspect of the present invention provides a kit for determining nucleic acid sequences, comprising sequencing reagents and instructions for performing the method. BRIEF DESCRIPTION OF THE DRAWINGS

[0112] Figure 1 The standard construct is shown below and comprises P5, SBS3, then a T residue followed by the genomic insert, SBS8' followed by an A residue and P7.

[0113] Figure 2 The grayscale of the three channels (green, orange and red) for observation is shown. st Cycle through individual scans.

[0114] Figure 3 The intensity histogram shows that the combined signals result in different intensity information.

[0115] Figure 4 Again the intensity histogram is shown.

[0116] To illustrate the present invention, a flow cell with a v2 PhiX cluster was used.

[0117] Figure 1 The standard construct is shown below and comprises P5, SBS3, then a T residue followed by the genomic insert, SBS8' followed by an A residue and P7.

[0118] Figure 1 SBS3 / SBS8 std PE construct

[0119] Bold = P5, Underline = SBS3, Italic = SBS8, Italic and Bold = P7

[0120] AATGATACGGCGACCACCGAGATCT ACACTCTTTCCCTACCACGACGCTTCTTCCGATC T---insert---

[0121] TTACTATGCCGCTGGTGGCTCTAGATGTGAGAAAGGGATGGTGCTGCGAGAAGGCTAGA---insert---

[0122] AGATCGGAAGAGCGGTTCAGCAGGAATGCCGAGACCGATCTCGTATGCCGTCTTCTGCTTGTCTAGCCTTCTCGCCAAGTCGTCCTTACGGCTCTGGCTAGAGCATACGGCAGAAGACGAAC

[0123] SBS3, SBS3+T and SBS8' primers were hybridized individually and in combination as shown in the table below.

[0124] Table 1

[0125] Lane Primer hybridization <![CDATA[Desired 1 st Loop]]> 1 SBS3 T 2 SBS3+T All 4 (as genome) 3 SBS8' A 4 SBS3 / SBS8' T+A 5 SBS3+T / SBS8' All 4 (as genome) + A

[0126] The first sequencing cycle was performed and the first image of the flow cell was acquired. Figure 2 The grayscale of the three channels (green, orange and red) for observation is shown. st Individual scans were performed in cycles. Lanes 1 and 3 have the strongest signals in two different channels, while lane 4 has the strongest signal in all three channels due to the incorporation of a mixture of T and A bases at SBS3 and SBS8' in this lane. T incorporation at SBS3 shows strong signals in the green and orange channels, but very weak signals in the red channel. A incorporation at SBS8' shows strong signals in the orange and red channels, but a mixture of weak and strong signals in the green channel. SBS3 / SBS8' shows strong signals in the green, orange, and red channels due to the incorporation of T and A bases at both primers.

[0127] from Figure 3 As can be seen, the intensity histograms show that the combined signals result in different intensity information, which can then be assigned to base information to allow efficient base calling. The top histogram is from SBS3 alone, illustrating T base calls, the bottom histogram is from SBS8' alone, illustrating A base calls, and the center histogram shows the combined SBS3 / SBS8' histogram, where both A and T base calls are made, one from each extended read.

[0128] Similarly, Figure 4Again, intensity histograms are shown, this time the top histogram is from SBS3+T alone, where the base call can be any of the 4 bases associated with genomic DNA, the bottom histogram is from SBS8' alone, illustrating the A base call, and the center histogram shows the combined SBS3+T / SBS8' histogram, where the base call can be any of the 4 bases associated with genomic DNA from one extended read plus the A base call from the other extended read giving a double intensity peak from the A / A cluster.

[0129] Once the intensity reads and images have been obtained, base calls can be made for each position with multiple calls made at each position for each cycle of sequencing. Each call at each position then needs to be assigned to the correct extended read. This can be done chemically, for example by making one of the reads brighter than the other by using a mixture of blocked and unblocked primers. Another option is to use bioinformatics information to determine which is statistically more "correct" base call, for example if we use a human insert or an E. coli insert, etc.

[0130] In certain embodiments, the multiple primers can be overlapping primers. For example, SBS3 and SBS3+T can be used for sequence 2 conserved bases to give better accuracy by interrogating each base multiple times over multiple cycles, in which case each base is interrogated twice using two primers that differ by a single base addition.

[0131] In some cases, blocked and unblocked primers are used to chemically distinguish extended reads. For example, one of the primers will be completely unblocked at the beginning, while the other primer will be a mixture of blocked / unblocked. So, for example, primer 1 will be 100% unblocked and give 100% signal on sequencing. Then "primer 2" will be, for example, a mixture of 25% unblocked and 75% blocked, which means that on sequencing you will get 25% signal from that primer site. So, you will end up with 2 reads giving data at 2 different levels (100% and 25%), which will make it easier to distinguish which base belongs to which read.

[0132] When using bioinformatics data to assign bases to extended reads, it is possible to bioinformatically search for the most likely read from the mixture of bases read, e.g. if your read is:

[0133] A / G,A / G,T / T,C / G,C / C,A / T

[0134] Then, one can search the sequenced genome for possible matches and possibly distinguish the two reads as:

[0135] R1, human, AGTCCT

[0136] R2, Escherichia coli, GATGCA

[0137] This is in contrast to any other possible reads from the mixture of bases mentioned above (eg, AATCCA was not found in either genome). Longer reads will be more likely to be unique to each genome.

[0138] Sequencing methods

[0139] Method described herein can be used in combination with various nucleic acid sequencing technologies. Particularly suitable technology is wherein nucleic acid is attached to the fixed position in the array so that their relative position does not change, and wherein the array is repeatedly imaged those. In different color channels, the embodiment of image is obtained, for example, being particularly suitable for conforming to the different marks for distinguishing a kind of nucleotide base type from another kind of nucleotide base type. In some embodiments, the process of measuring the target nucleic acid nucleotide sequence can be an automated process. Preferred embodiments include synthetic sequencing (" SBS ") technology.

[0140] SBS techniques typically involve enzymatic extension of a nascent nucleic acid chain by repeated addition of nucleotides to a template strand. In conventional SBS methods, a single nucleomonomer can be provided to a target nucleic acid in each delivery in the presence of a polymerase. However, in the methods described herein, more than one type of nucleomonomer can be provided to a target nucleic acid in the presence of a polymerase during delivery.

[0141] SBS can utilize nucleotide monomers with terminator moiety or those lacking any terminator moiety.The method using the nucleotide monomer lacking terminator, for example, includes, using pyrophosphate sequencing and the order-checking of the nucleotide of gamma-phosphate / phosphate labeling, as further elaborated below.In the method using the nucleotide monomer lacking terminator, the nucleotide number added in each cycle is normally variable, and depends on template sequence and nucleotide delivery mode.For the SBS technology utilizing the nucleotide monomer with terminator moiety, terminator can be effectively irreversible under order-checking conditions, as the situation of traditional Sanger order-checking of dideoxynucleotide utilized, or terminator can be reversible, as the situation of the order-checking method developed by Solexa (now Illumina, Inc.).

[0142] SBS technology can utilize nucleotide monomers with or without a labeling portion. Thus, incorporation events can be detected based on characteristics of the label, such as fluorescent labels; characteristics of the nucleotide monomer such as molecular weight or charge; byproducts of nucleotide incorporation, such as phosphate / salt release, etc. In embodiments, where two or more different nucleotides are present in the sequencing reagent, the different nucleotides can be distinguished from each other, or alternatively, the two or more different labels can be indistinguishable under the detection technology used. For example, different nucleotides present in the sequencing reagent can have different labels and can be distinguished using appropriate optics exemplified by sequencing methods developed by Solexa (now Illumina, Inc.).

[0143] Embodiments include pyrosequencing technology. Pyrosequencing detects the release of inorganic pyrophosphate (PPi) due to the incorporation of specific nucleotides into nascent chains (Ronaghi, M., Karamohamed, S., Pettersson, B., Uhlen, M. and Nyren, P. (1996) "Real-time DNA sequencing using detection of pyrophosphate release." Analytical Biochemistry 242(1), 84-9; Ronaghi, M. (2001) "Pyrosequencing sheds light on DNA sequencing." Genome Res. 11(1), 3-11; Ronaghi, M., Uhlen, M. and Nyren, P. (1998) "A sequencing method based on real-time pyrophosphate." Science 281(5375),363; U.S. Patent No. 6,210,891; U.S. Patent No. 6,258,568 and U.S. Patent No. 6,274,320, the foregoing disclosures are incorporated herein by reference in their entireties). In pyrosequencing, released PPi can be immediately converted to adenosine triphosphate (ATP) by ATP sulfurylase, and the resulting ATP level can be detected by photons generated by luciferase. The nucleic acids to be sequenced can be attached to features in an array, and imaging can be performed to capture the chemiluminescent signal generated by the incorporation of nucleotides at the features of the array. An image can be obtained after treating the array with a specific nucleotide type (e.g., A, T, C, or G). The image obtained after adding each nucleotide type will differ depending on which features in the array are detected. These differences in the image reflect the different sequence content of the features on the array. However, the relative position of each feature in the image will remain unchanged. The images can be stored, processed, and analyzed using the methods described herein. For example, images obtained after treating an array with various nucleotide types can be processed in the same manner as images obtained with different detection channels for sequencing methods based on reversible terminators, as exemplified herein.

[0144] In another exemplary type of SBS, cycle sequencing is accomplished by the stepwise addition of reversible terminator nucleotides, which, for example, include cleavable or fluorescent cleavable tags as described below: WO 04 / 018497 and U.S. Patent No. 7,057,026, which are incorporated herein by reference in their entirety. This method is being commercialized by Solexa (now Illumina, Inc.) and is also described in WO 91 / 06678 and WO 07 / 123,744, each of which is incorporated herein by reference in its entirety. The availability of fluorescently labeled terminators (which can reverse the two terminators) and cleaved fluorescent tags promotes effective cyclic reversible termination (CRT) sequencing. Polymerases can also be co-designed to effectively incorporate and extend from these modified nucleotides.

[0145] Preferably, in reversible terminator-based sequencing embodiments, the label is substantially incapable of inhibiting extension under the SBS reaction conditions. However, the detection label may be removable, for example by cleavage or degradation. After the label is incorporated into the nucleic acid features of the array, an image can be captured. In a specific embodiment, each cycle involves the simultaneous delivery of four different nucleotide types to the array, with each nucleotide type having a spectrally distinct label. Four images can then be obtained, each using a detection channel selected for one of the four different labels. Alternatively, different nucleotide types can be added sequentially, and an image of the array can be obtained between each addition step. In such embodiments, each image will display the nucleic acid features that have incorporated a particular nucleotide type. Due to the different sequence content of each feature, different images will have different features present or absent. However, the relative positions of the features in the images will remain unchanged. Images obtained from such reversible terminator-SBS methods can be stored, processed, and analyzed as described herein. After the image capture step, the label can be removed, and the reversible terminator moiety can be removed for subsequent nucleotide addition and detection cycles. Once the label is detected within a specific cycle and in subsequent cycles, it can provide the advantages of reduced background signal and crosstalk between cycles. Examples of useful tags and removal methods are described below.

[0146] In a specific embodiment, some or all of the nucleotide monomers may include reversible terminators. In such embodiments, reversible terminators / cleavable fluorescent agents may include fluorescent agents (Metzker, Genome Res. 15: 1767-1776 (2005) that are connected to a ribose moiety via 3' ester bonds. Other methods separate terminator chemistry from fluorescently labeled cutting (Ruparel et al., Proc Natl Acad Sci USA 102: 5932-7 (2005) that are incorporated herein by reference in their entirety). Ruparel et al. have described reversible terminators for extending the use of small 3' allyl groups for development, but they can be easily unsealed by the short-lived treatment of palladium catalysts. Fluorophores are attached to bases via photocleavable joints, and the joints can be easily cut by being exposed to long wavelength UV light for 30 seconds. Therefore, disulfide bond reduction or photocleavage can be used as cleavable connexons. Another method of reversible termination is to use the natural termination that occurs after placing a bulky dye on dNTP. The presence of a bulky dye with a charge on the dNTP can act as an effective terminator through steric and / or electrostatic hindrance. Unless the dye is removed, the presence of an incorporation event prevents further incorporation. The cleavage of the dye can remove fluorescence and effectively reverse termination. Examples of modified nucleotides are also described in U.S. Patent No. 7,427,673 and U.S. Patent No. 7,057,026, which are incorporated herein by reference in their entirety.

[0147] Additional exemplary SBS systems and methods that can be used with the methods and systems described herein are described in U.S. Patent Application Publication No. 2007 / 0166705, U.S. Patent Application Publication No. 2006 / 0188901, U.S. Patent No. 7,057,026, U.S. Patent Application Publication No. 2006 / 0240439, U.S. Patent Application Publication No. 2006 / 0281109, PCT Publication No. WO 05 / 065814, U.S. Patent Application Publication No. 2005 / 0100900, PCT Publication No. WO 06 / 064199, PCT Publication No. WO 07 / 010,251, U.S. Patent Application Publication No. 2012 / 0270305, and U.S. Patent Application Publication No. 2013 / 0260372, the contents of each of which are incorporated herein by reference in their entirety.

[0148] Some embodiments can utilize four different markers to detect four different nucleotides.For example, the method and system described in U.S. Patent Publication No. 2013 / 0079232 can be used to perform SBS, and this patent is incorporated herein by reference in its entirety. As a first example, a pair of nucleotide types can be detected with identical wavelength, but based on the intensity difference of a pair of members and another member, or based on the change of a pair of members (for example, via chemical modification, photochemical modification or physical modification), the change causes the obvious signal that the signal detected by other members compared to the pair appears or disappears. As a second example, three of four different nucleotide types can be detected under specific conditions, and the fourth nucleotide type lacks a marker detectable under these conditions, or is detected by minimum detection (for example, the minimum detection caused by background fluorescence etc.) under these conditions. The first three nucleotide types can be measured based on the signal of their respective existence to be incorporated into nucleic acid, and the fourth nucleotide type can be determined to be incorporated into nucleic acid based on the absence or minimum detection of any signal. As a third example, a kind of nucleotide type can include the marker detected in two different channels, and detects other nucleotide types in no more than one channel. The three exemplary configurations described above are not considered mutually exclusive and can be used in various combinations. An exemplary embodiment that combines all three examples is a fluorescence-based SBS method that uses a first nucleotide type detected in a first channel (e.g., dATP with a label detected in the first channel when excited by a first excitation wavelength), a second nucleotide type detected in a second channel (e.g., dCTP with a label detected in the second channel when excited by a second excitation wavelength), a third nucleotide type detected in the first and second channels (e.g., dTTP with at least one label detected in both channels when excited by the first and / or second excitation wavelengths), and a fourth nucleotide type lacking a label that is not detected or minimally detected in either channel (e.g., dGTP without a label).

[0149] Furthermore, as described in the materials incorporated into U.S. Patent Application Publication No. 2013 / 0079232, sequencing data can be obtained using a single channel. In this so-called single-dye sequencing approach, a first nucleotide type is labeled, but the label is removed after the first image is generated, and the second nucleotide type is labeled only after the first image is generated. The third nucleotide type retains its label in both the first and second images, and the fourth nucleotide type remains unlabeled in both images.

[0150] Some embodiments can utilize ligation technology to carry out sequencing. Such technology utilizes DNA ligase to incorporate oligonucleotides and identifies the incorporation of such oligonucleotides. Oligonucleotides generally have different labels related to the identity of specific nucleotides in the sequence to which the oligonucleotides hybridize. As with other SBS methods, images can be obtained after treating the nucleic acid feature array with labeled sequencing reagents. Each image will show the nucleic acid features that have been incorporated with a specific type of label. Due to the different sequence content of each feature, different images will have or not different features, but the relative positions of the features in the image will remain unchanged. The images obtained from the sequencing method based on connection can be stored, processed and analyzed as described herein. Exemplary SBS systems and methods that can be used with the methods and systems described herein are described in U.S. Patent No. 6,969,488, U.S. Patent No. 6,172,218 and U.S. Patent No. 6,306,597, which are incorporated herein by reference in their entirety.

[0151] Sequencing by ligation is a well-known method for sequencing that requires repeated or prolonged illumination of a dibase probe with light. Exemplary systems using sequencing by synthesis include the SOLiD TM System (Life Technologies, Carlsbad, CA). In short, the method of sequencing by ligation includes hybridizing a sequencing primer to an adapter sequence fixed on a template bead. A set of four fluorescently labeled two-base probes competes for attachment to the sequencing primer. The specificity of the two-base probe is achieved by interrogating every second base in each ligation reaction. After a series of ligation cycles, the extension product is removed and the template is reset with a sequencing primer complementary to the n-1 position for a second round of ligation cycles. Multiple cycles of ligation, detection, and cleavage are performed with the number of cycles that determine the final read length.

[0152] In addition, the methods and technical solutions described herein can be particularly suitable for sequencing from nucleic acid arrays, where multiple sequences can be read simultaneously from multiple positions on the array because each nucleotide at each position can be identified based on its identifiable label. Exemplary methods are described in US 2009 / 0088327; US 2010 / 0028885; and US 2009 / 0325172.

[0153] Some embodiments can be achieved by utilizing nanopore sequencing (Deamer, DW & Akeson, M. "Nanopores and nucleic acids: prospects for ultrarapid sequencing." Trends Biotechnol. 18, 147-151 (2000); Deamer, D. and D. Branton, "Characterization of nucleic acids by nanopore analysis". Acc. Chem. Res. 35: 817-825 (2002); Li, J., M. Gershow, D. Stein, E. Brandin, and J. A. Golovchenko, "DNA molecules and configurations in a solid-state nanopore microscope" Nat. Mater. 2: 611-615 (2003), the disclosures of which are incorporated herein by reference in their entireties). In such embodiments, the target nucleic acid passes through the nanopore. The nanopore can be a synthetic pore or a biological membrane protein, such as α-hemolysin. As the target nucleic acid passes through the nanopore, each base pair can be identified by measuring the fluctuations in pore conductance (US Patent 7,001,792; Soni, GV & Meller, "A. Progress toward ultrafast DNA sequencing using solid-state nanopores." Clin. Chem. 53, 1996-2001 (2007); Healy, K. "Nanopore-based single-molecule DNA analysis." Nanomed. 2, 459-481 (2007); Cockroft, SL, Chu, J., Amorin, M. & Ghadiri, MR "A single-molecule nanopore device detects DNA polymerase activity with single-nucleotide resolution." J. Am. Chem. Soc. 130, 818-820 (2008), the disclosures of which are incorporated herein by reference in their entireties). Data obtained from nanopore sequencing can be stored, processed, and analyzed as described herein. In particular, the data can be processed as images according to the exemplary processing of optical images and other images listed herein.In certain embodiments, a single pore can be used to hold a DNA strand hybridized to two or more primers. Light-based sequencing-by-synthesis can then be performed.

[0154] Some embodiments can utilize methods involving real-time monitoring of DNA polymerase activity. Nucleic acid incorporation can be detected by fluorescence resonance energy transfer (FRET) interactions between nucleotides containing fluorophore polymerases and gamma-phosphate labels, as described, for example, in U.S. Patent No. 7,329,492 and U.S. Patent No. 7,211,414 (the disclosures are incorporated herein by reference in their entirety), or nucleotide incorporation can be detected using zero-mode waveguides and fluorescent nucleotide analogs and engineered polymerases. The zero-mode waveguides are described, for example, in U.S. Patent No. 7,315,019 (the disclosure is incorporated herein by reference in its entirety), and the fluorescent nucleotide analogs and engineered polymerases are described, for example, in U.S. Patent No. 7,405,281 and U.S. Patent Application Publication No. 2008 / 0108082 (the disclosures are incorporated herein by reference in their entirety). The illumination can be confined to a zeptoliter-scale volume surrounding the surface-tethered polymerase, allowing low-background observation of the incorporation of fluorescently labeled nucleotides (Levene, MJ et al. "Zero-mode waveguides for single-molecule analysis at high concentrations." Science 299, 682-686 (2003); Lundquist, PM et al. "Parallel confocal detection of single molecules in realtime." Opt. Lett. 33, 1026-1028 (2008); Korlach, J. et al. "Selective aluminum passivation for targeted immobilization of single DNA polymerase molecules in zero-mode waveguide nanostructures." Proc. Natl. Acad. Sci. USA 105, 1176-1181 (2008), the disclosures of which are incorporated herein by reference in their entireties). Images obtained from these methods can be stored, processed, and analyzed as described herein. In certain embodiments, two polymerases may be present at the bottom of each well, each generating sequencing data from a different primer attached to the same template molecule.

[0155] Some SBS embodiments include detection of protons released when nucleotides are incorporated into extension products. For example, sequencing based on detection of released protons can use electrical detectors available from Ion Torrent (Guilford, CT, a Life Technologies subsidiary) and related technologies or sequencing methods and systems described in the following patents: US 2009 / 0026082 A1; US 2009 / 0127589 A1; US 2010 / 0137143 A1; or US 2010 / 0282617 A1, each incorporated herein by reference. More specifically, the methods listed herein can be used to produce clonal populations of amplicons for detecting protons. In embodiments using pyrophosphate sequencing, the use of two or more primers for each strand of nucleic acid may mean that more signal per unit time needs to be deconvoluted. In pyrophosphate sequencing, a drop in pH twice that of normal that is washed out may mean that one of the primers has a TT incorporation, or that each primer has a single T incorporation.

[0156] Said method can advantageously be carried out with multiple forms so that a plurality of different target nucleic acids are operated simultaneously.In a specific embodiment, different target nucleic acids can be processed in common reaction vessels or on the surface of specific substrate.This allows to conveniently deliver sequencing reagents in a variety of ways, remove unreacted reagents and detect incorporation events.In the embodiment using surface-bound target nucleic acid, target nucleic acid can be array form.In array form, target nucleic acid can be attached to surface in a spatially distinguishable manner generally.Target nucleic acid can be attached to beads or other particles or be bound to polymerase or be attached to other molecular combinations on surface by direct covalent attachment.The array can be included in the single copy of the target nucleic acid in each site (also referred to as feature), or multiple copies with identical sequence can be present in each site or feature.Can be by the amplification method production of bridge amplification or emulsion PCR as described in further detail below.

[0157] The methods described herein can use arrays having various densities of features, including, for example, at least about 10 features / cm 2 , 100 features / cm 2 , 500 features / cm 2 , 1,000 features / cm 2 , 5,000 features / cm 2 , 10,000 features / cm 2 , 50,000 features / cm 2 , 100,000 features / cm 2 , 1,000,000 features / cm 2 5,000,000 features / cm 2, or higher.

[0158] An advantage of the methods described herein is that they provide rapid and efficient detection of multiple target nucleic acids in parallel. Thus, the present disclosure provides integrated systems capable of preparing and detecting nucleic acids using techniques known in the art, such as those listed above. Thus, the integrated systems of the present disclosure may include fluid components capable of delivering amplification reagents and / or sequencing reagents for one or more immobilized DNA fragments, the systems comprising components such as pumps, valves, reservoirs, and fluid lines. A flow cell may be configured and / or used in an integrated system for detecting target nucleic acids. Exemplary flow cells are described, for example, in US 2010 / 0111768 A1 and US Serial No. 13 / 273,666, each of which is incorporated herein by reference in its entirety. As exemplified by the flow cell, one or more fluid components of the integrated system may be used in both amplification methods and detection methods. Using the nucleic acid sequencing embodiment as an example, one or more fluid components of the integrated system may be used in both the amplification methods described herein and for delivering sequencing reagents in sequencing methods such as those exemplified above. Alternatively, the integrated system may include separate fluid systems for performing the amplification method and the detection method. Examples of integrated sequencing systems capable of creating amplified nucleic acids and also determining nucleic acid sequences include, but are not limited to, the MiSeq. TM The platform (Illumina, Inc., San Diego, CA) and the apparatus described in U.S. Serial No. 13 / 273,666, which is incorporated herein by reference in its entirety.

[0159] Nucleic acid amplification

[0160] In some embodiments, immobilized DNA fragments are amplified using cluster amplification methods as disclosed in U.S. Patent Nos. 7,985,565 and 7,115,400, the contents of each of which are incorporated herein by reference in their entirety. The materials incorporated into U.S. Patent Nos. 7,985,565 and 7,115,400 describe solid-phase nucleic acid amplification methods that allow amplified products to be fixed on a solid support to form clusters or "communities" comprising immobilized nucleic acid molecules. Each cluster or community on such an array is formed by a plurality of identical immobilized polynucleotide chains and a plurality of identical immobilized complementary polynucleotide chains. The arrays so formed herein are generally referred to as "cluster arrays." The product of the solid-phase amplification reaction described in U.S. Patent Nos. 7,985,565 and 7,115,400 is a so-called "bridged" structure, formed by annealing a fixed polynucleotide chain and a pair of unfixed complementary chains, the two chains preferably being immobilized on a solid support at the 5' end via covalent attachment. Cluster amplification methods are examples of methods in which immobilized nucleic acid templates are used to produce immobilized amplicons. Other suitable methods can also be used to produce immobilized amplicons from immobilized DNA fragments prepared according to the methods provided herein. For example, one or more clusters or colonies can be formed by solid-phase PCR in which one or two primers of each amplification primer pair are immobilized.

[0161] In other embodiments, immobilized DNA fragmentation is increased in solution.For example, in some embodiments, immobilized DNA fragmentation is cut or otherwise discharged from solid support, and then amplification primer is hybridized with the released molecule in solution.In other embodiments, amplification primer is hybridized with immobilized DNA fragmentation in one or more initial amplification steps, then carries out subsequent amplification step in solution.Therefore, in some embodiments, immobilized nucleic acid template can be used for producing solid phase amplicon.

[0162] It should be understood that any amplification method described herein or generally known in the art can be used for universal or target-specific primers to amplify immobilized DNA fragments. Suitable methods for amplification include but are not limited to polymerase chain reaction (PCR), strand displacement amplification (SDA), transcription-mediated amplification (TMA) and nucleic acid sequence-based amplification (NASBA), as described in U.S. Patent number 8,003,354, which is incorporated herein by reference in its entirety. The above-mentioned amplification method can be used to amplify one or more nucleic acids of interest. For example, PCR including multiplex PCR, SDA, TMA, NASBA etc. can be used to amplify immobilized DNA fragments. In some embodiments, primers specific for nucleic acids of interest are included in the amplification reaction.

[0163] Other suitable methods for amplifying nucleic acids can include oligonucleotide extension and ligation, rolling circle amplification (RCA) (Lizardi et al., Nat. Genet. 19: 225-232 (1998), which is incorporated herein by reference) and oligonucleotide ligation assay (OLA) (see generally U.S. Patent Nos. 7,582,420, 5,185,243, 5,679,524 and 5,573,907; EP 0 320 308 B1; EP 0 336 731 B1; EP 0 439 182 B1; WO 90 / 01069; WO 89 / 12696; and WO 89 / 09835, all of which are incorporated herein by reference) technology. It should be understood that these amplification methods can be designed to amplify immobilized DNA fragments. For example, in some embodiments, the amplification method can include ligation probe amplification or oligonucleotide ligation assay (OLA) reactions containing primers specifically for the nucleic acid of interest. In some embodiments, the amplification method can include a primer extension ligation reaction containing primers specifically for the nucleic acid of interest. As a non-limiting example of primer extension and ligation primers that can be specifically designed to amplify the nucleic acid of interest, the amplification can include primers used in the GoldenGate assay (Illumina, Inc., San Diego, California) as shown in U.S. Patent Nos. 7,582,420 and 7,611,869, each of which is incorporated herein by reference in its entirety.

[0164] Exemplary isothermal amplification methods that can be used in the methods of the present disclosure include, but are not limited to, multiple displacement amplification (MDA) as described, for example, in Dean et al., Proc. Natl. Acad. Sci. USA 99:5261-66 (2002), or isothermal strand displacement nucleic acid amplification as described, for example, in U.S. Pat. No. 6,214,587, each of which is incorporated herein by reference in its entirety. Other non-PCR-based methods that can be used in the present disclosure include, for example, strand displacement amplification (SDA) as described in Walker et al., Molecular Methods for Virus Detection, Academic Press, Inc., 1995; U.S. Patent Nos. 5,455,166 and 5,130,238, and Walker et al., Nucl. Acids Res. 20: 1691-96 (1992), or hyperbranched strand displacement amplification as described in Lage et al., Genome Research 13: 294-307 (2003), each of which is incorporated herein by reference in its entirety. Allelic amplification methods can be used for strand displacement Phi 29 polymerase or Bst DNA polymerase large fragment, 5'->3' in vitro for random primer amplification of genomic DNA. The use of these polymerases takes advantage of their high processivity and strand displacement activity. High processivity allows the polymerase to produce fragments of 10-20 kb in length. As described above, a polymerase with low processivity and strand displacement activity (such as Klenow polymerase) can be used under isothermal conditions to generate smaller fragments. Additional description of amplification reactions, conditions, and components is described in detail in the disclosure of U.S. Patent No. 7,670,810, which is incorporated herein by reference in its entirety.

[0165] Another nucleic acid amplification method useful in the present disclosure is a tag PCR (Tagged PCR) using a dual-domain primer population with a constant 5' region followed by a random 3' region, which is described, for example, in Grothues et al. Nucleic Acids Res. 21(5): 1321-2 (1993), incorporated herein by reference in its entirety. A first round of amplification is performed to allow multiple starts on heat-denatured DNA based on individual hybridization of the randomly synthesized 3' region. Due to the nature of the 3' region, the start site is considered to be random throughout the genome. Therefore, unbound primers can be removed and further replication can occur using primers complementary to the constant 5' region.

[0166] The present invention provides:

[0167] 1. A method for determining a nucleic acid sequence, comprising:

[0168] providing at least one nucleic acid bound to a support;

[0169] hybridizing at least two primers to the same strand of nucleic acid;

[0170] performing primer extension on the same strand of the nucleic acid at each of the at least two primers;

[0171] obtaining signal data corresponding to at least one nucleotide base incorporated during each primer extension;

[0172] The identities of the nucleic acid bases are determined from the signal data and the bases are assigned to extension reads.

[0173] 2. The method of claim 1 , wherein the method comprises contacting the nucleic acid with an enzyme under conditions permissive for hybridization in the presence of: (i) at least two primers capable of hybridizing to the same strand of the nucleic acid at different positions, and (ii) four labelled moieties comprising at least one nucleotide analogue selected from dGTP, dCTP, dTTP, dUTP, and dATP, each of the four labelled moieties having a unique label that is different from the unique labels of the other three labelled moieties.

[0174] 3. The method of claim 2, wherein the method comprises a step of removing unbound labeled portions.

[0175] 4. The method of item 2 or 3, wherein the plurality of extended reads comprise substantially simultaneous extended reads.

[0176] 5. The method of any one of items 2 to 4, wherein the four labeling moieties each comprise a single reversible terminator nucleotide analog or reversible terminator analog selected from dGTP, dCTP, dTTP, dUTP, and dATP or an oligonucleotide probe comprising at least one nucleotide analog selected from dGTP, dCTP, dTTP, dUTP, and dATP.

[0177] 6. The method of any one of items 1 to 5, wherein the enzyme comprises a polymerase or a ligase.

[0178] 7. The method of any of the preceding items, wherein the support comprises a chip or a bead.

[0179] 8. The method of any one of items 2 to 7, wherein the unique label comprises a dye, a fluorophore, a chromophore, a combined fluorescence energy transfer tag, a mass tag, or an electrophore.

[0180] 9. The method of any of the preceding items, wherein at least two primers have overlapping sequences.

[0181] 10. The method of any of the preceding items, wherein at least two primers differ by a single base addition.

[0182] 11. The method of any of the preceding items, wherein the plurality of primers comprises blocked and unblocked primers to chemically distinguish each extended read.

[0183] 12. The method of any of the preceding items, wherein bioinformatics information is used to assign bases to the extended reads.

[0184] 13. The method of any preceding item, wherein said signal data corresponding to each of said at least two primer extensions are obtained substantially simultaneously.

[0185] 14. The method of any one of items 2 to 13, wherein the step of determining the identity of the nucleotide base comprises analyzing a signal intensity profile corresponding to a unique label detected on each of the plurality of extended reads.

[0186] 15. The method of any preceding clause, wherein the signal data comprises one or more images.

[0187] 16. The method of claim 15, wherein the signal data is detected as color signals corresponding to multiple nucleotide analogs incorporated on the primer extension, and wherein each color signal corresponds to a different nucleotide analog or a combination of nucleotide analogs.

[0188] 17. The method of any of the preceding items, wherein the method comprises a signal processing step comprising one or more of signal deconvolution, signal refinement, and signal selection.

[0189] 18. The method of any of the preceding items, wherein multiple base calls are made during each hybridization cycle.

[0190] 19. The method of any of the preceding items, wherein the step of assigning the base to an extended read comprises determining the position of the nucleotide base.

[0191] 20. The method of any of the preceding items, wherein the step of assigning the bases to extended reads comprises providing a preliminary base call and a final base call for each extended read.

[0192] 21. The method of claim 20, wherein final base calls are provided by comparing preliminary base call data to a reference genome.

[0193] 22. The method of any one of items 2 to 21, wherein the same set of four marker portions is used for each extended read.

[0194] 23. A system for determining a nucleic acid sequence comprising a sequencing instrument, wherein the sequencing instrument

[0195] A device having a solid support for immobilizing at least one nucleic acid and for:

[0196] providing at least one nucleic acid bound to a support;

[0197] hybridizing at least two primers to the same strand of nucleic acid;

[0198] performing at least two primer extensions on the same strand of the nucleic acid at each of the at least two primers;

[0199] obtaining signal data corresponding to at least one nucleotide base incorporated during each of at least two primer extensions;

[0200] The identities of the nucleic acid bases are determined from the signal data and the bases are assigned to extension reads.

[0201] 24. The system of claim 23, further comprising a signal processor for processing the signal data.

[0202] 25. A kit for determining a nucleic acid sequence, comprising sequencing reagents and instructions for performing the method of any one of items 1 to 22. Sequence Listing <110> Illumina Cambridge Ltd <120> Sequencing from multiple primers to increase data rate and density <130> P166912.WO.01 <160> 2 <170> BiSSAP 1.3 <210> 1 <211> 120 <212> DNA <213> Homo sapiens <400> 1 aatgatacgg cgaccaccga gatctacact ctttccctac cacgacgctc ttccgatcta 60 gatcggaaga gcggttcagc aggaatgccg agaccgatct cgtatgccgt cttctgcttg 120 <210> 2 <211> 120 <212> DNA <213> Homo sapiens <400> 2 ttactatgcc gctggtggct ctagatgtga gaaagggatg gtgctgcgag aaggctagat 60 ctagccttct cgccaagtcg tccttacggc tctggctaga gcatacggca gaagacgaac 120

Claims

1. An electronic system comprising a flow cell comprising a support having a plurality of bound polynucleotides, wherein the electronic system comprises a processor configured to perform a method comprising: activating a fluid component to hybridize a plurality of first and second primers to the plurality of bound polynucleotides, wherein at least the first and second primers hybridize to the same strand of each polynucleotide, and wherein the plurality of first and second primers comprises blocked and unblocked primers to chemically distinguish each extended read of the first and second primers; activating the fluid component to extend the first and second primers hybridized to the same strand of each polynucleotide with signal-generating labeled nucleotide analogs, wherein each type of labeled nucleotide analog is labeled with a unique label; obtaining signal data corresponding to at least one nucleotide analog incorporated at each of the extended first and second primers; determining the identity of the nucleotide analog from the signal data; and The identities of the nucleotide analogs are assigned to the extended reads of the first and second primers based on the brightness of the extended reads. 2 . The system of claim 1 , wherein the processor is configured to perform a method of obtaining signal data by obtaining one or more images.

3. The system of claim 1, wherein the processor is configured to perform a method of obtaining signal data by obtaining color signal data corresponding to a plurality of nucleotide analogs incorporated at the primer.

4. The system of claim 3, wherein each color signal corresponds to a different nucleotide analog or combination of nucleotide analogs.

5. The system of claim 1, wherein the processor is configured to perform a method of determining the identity of the nucleotide analog by one or more of signal deconvolution, signal improvement, and signal selection.

6. The system of claim 1, wherein the processor is configured to perform a method for determining the identity of the nucleotide analogs by analyzing a signal intensity profile corresponding to a unique label detected at each primer. 7 . The system of claim 1 , further comprising assigning the nucleotide analogs to each extended read using bioinformatics information.

8. The system of claim 1, wherein the processor is configured to perform a method of assigning the identity of each labeled nucleotide analog by performing a preliminary base call and then a final base call for the nucleotide analog.

9. The system of claim 8, wherein the processor is configured to perform the final base calling by comparing the preliminary base calling data to a reference genome.

10. An electronic system for determining the sequence of a nucleic acid comprising at least two hybridization primers, wherein each of said hybridization primers comprises a labeled nucleotide analog, wherein each of said labeled nucleotide analogs is labeled with a unique label, and wherein said system comprises: a flow cell comprising nucleic acid bound to a support and hybridized to two primers, wherein a first primer comprises a first labeled nucleotide analog and a second primer comprises a second labeled nucleotide analog; an image capture system that captures an image of the first and second labeled nucleotide analogs on the hybridized primer; and A signal processor obtains signal data corresponding to the first and second labeled nucleotide analogs from the image capture system and identifies the labeled nucleotide analogs by calculating an intensity histogram of the signal data from the first and second labeled nucleotide analogs.

11. The system of claim 10, wherein the signal processor is configured to perform a method of determining the identity of the nucleotide analogues by one or more of signal deconvolution, signal improvement, and signal selection.

12. The system of claim 10, wherein the signal processor is configured to perform a method of obtaining signal data by obtaining one or more images.

13. The system of claim 12, wherein the processor is configured to perform a method of obtaining signal data by obtaining color signal data corresponding to a plurality of nucleotide analogs incorporated at the primer.

14. The system of claim 13, wherein each color signal corresponds to a different nucleotide analog or combination of nucleotide analogs.

15. The system of claim 10, wherein the signal processor is configured to perform a method for determining the identity of the nucleotide analogs by analyzing a signal intensity profile corresponding to a unique label detected at each primer.

16. The system of claim 10, wherein the signal processor is configured to assign identities of the nucleotide analogs by performing a preliminary base call and then a final base call for each labeled nucleotide analog.

17. The system of claim 16, wherein the signal processor is configured to make the final base call by comparing the preliminary base call to a reference genome.

18. An electronic system comprising a flow cell comprising a support having a plurality of bound polynucleotides, wherein the electronic system comprises: a flow cell comprising a plurality of bound polynucleotides; one or more fluid components; Image capture system; and A signal processor configured to perform a method comprising: activating a fluid component to hybridize a plurality of first and second primers to the plurality of bound polynucleotides, wherein at least the first and second primers hybridize to the same strand of each polynucleotide, wherein the first and second primers comprise different primers selected from different primer mixtures comprising different levels of blocked and unblocked primers to chemically distinguish each extended read of the first and second primers; activating the fluid component to extend the first and second primers hybridized to the same strand of each polynucleotide with a signal-generating labeled nucleotide base; obtaining signal data from the image capture system corresponding to at least one nucleotide base incorporated at each of the extended first and second primers; determining the identity of the nucleotide analog from the signal data; and The identities of the nucleotide analogs are assigned to the extended reads of the first and second primers based on the brightness of the extended reads.

19. The system of claim 18, wherein the signal processor is configured to perform a method of obtaining signal data by obtaining one or more images from the image capture system.

20. The system of claim 18, wherein the signal processor is configured to perform a method of obtaining signal data by obtaining color signal data corresponding to a plurality of nucleotide analogs incorporated at the primer.

21. The system of claim 20, wherein each color signal corresponds to a different nucleotide analog or combination of nucleotide analogs.

22. The system of claim 18, wherein the signal processor is configured to perform a method of determining the identity of the nucleotide analog by one or more of signal deconvolution, signal improvement, and signal selection.

23. The system of claim 18, the signal processor configured to perform a method for determining the identity of the nucleotide analog by analyzing a signal intensity profile corresponding to a unique label detected at each primer.

24. The system of claim 18, further comprising assigning the nucleotide analogs to each extended read using bioinformatics information.

25. The system of claim 18, wherein the signal processor is configured to assign the identity of the nucleotide analogs by performing a preliminary base call and then a final base call for each labeled nucleotide base.

26. The system of claim 25, wherein the signal processor is configured to make the final base calls by comparing the preliminary base call data to a reference genome.

Citation Information

Patent Citations

  • Method for detecting a target nucleic acid sequence

    EP0320308B1

  • Method of amplifying and detecting nucleic acid sequences

    EP0336731B1

  • Improved method of amplifying target nucleic acids applicable to both polymerase and ligase chain reactions

    EP0439182B1

  • Improvement in milk-coolers

    US169196A

  • Method of nucleic acid amplification

    US20050100900A1