Systems and methods of sequencing polynucleotides with alternative scatterplots
Patent Information
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- ILLUMINA INC
- Filing Date
- 2024-06-25
- Publication Date
- 2026-05-06
AI Technical Summary
Current DNA sequencing technologies face challenges in accurately distinguishing between the G nucleotide and empty wells in two-channel detection systems, leading to incorrect base calls due to inefficient seeding and amplification processes, particularly in patterned flowcells where G nucleotides are not labeled, resulting in erroneous 'poly G' calls.
The method involves labeling at least four nucleotides with a mixture of fluorescent dyes to create alternative scatterplot shapes, such as diamond configurations, which allow for the differentiation of G nucleotides from empty wells by varying the intensity and ratio of emissions at different wavelengths, enabling accurate sequence determination.
This approach enhances the accuracy of DNA sequencing by reducing incorrect base calls and improving the ability to distinguish between G nucleotides and empty wells, thereby increasing the reliability of sequence data.
Smart Images

Figure US2024035409_02012025_PF_FP_ABST
Abstract
Description
ILLINC.766WO / IP-2551-PCT PATENT SYSTEMS AND METHODS OF SEQUENCING POLYNUCLEOTIDES WITH ALTERNATIVE SCATTERPLOTS INCORPORATION BY REFERENCE TO ANY PRIORITY APPLICATIONS
[0001] This application claims priority to U.S. Provisional Application No.63 / 511,382 filed June 30, 2023, the content of which is incorporated by reference in its entirety. BACKGROUND Field
[0002] The present disclosure relates to DNA sequencing systems and methods. In particular, this disclosure relates to improved detection methods using alternative scatterplot shapes in two-channel detection systems. Background
[0003] Current sequencing technologies involve determining DNA or RNA sequences by deciphering four natural bases in the genome, A, T (U), G, and C. One approach to polynucleotide sequencing involves using two optical channels to detect fluorescent emissions from dyes attached to different types of nucleotides. Illumina sequencing systems use onboard real-time analysis (RTA) to turn raw image data into basecalls (ACGT). This process can be massively parallelized and occurs in real time on the instrument.
[0004] In a standard two-channel detection method, scatterplots may be used for analysis and determination of a polynucleotide sequence where the x-axis represents the intensity of fluorescence at one wavelength (channel 1), and the y-axis represents intensity of fluorescence at another wavelength (channel 2). Each data point on the scatterplot represents a single nucleotide that has been detected by the sequencing instrument, and the position of the dot on the scatterplot indicates the signal intensity for each channel. By analyzing the pattern of dots on the scatterplot, researchers can determine the sequence of nucleotides in the sample. The use of two channels for polynucleotide sequencing allows for the simultaneous detection of multiple nucleotides, which can increase the speed and accuracy of the sequencing process.
[0005] For two-channel base-calling systems, the four DNA bases are encoded using two bits of information from each color channel, where the G base is encoded using two channels being in the off state. In a typical sequencing system, the seeding and amplification of reads is not expected to be completely efficient, which results in some empty wells. In 2-channel SBS on patterned flowcells, wells not occupied by a DNA cluster may be incorrectly called G because the G nucleotide is not labeled. This also occurs to a lesser extent with random flowcells due to spurious spots that are assigned as clusters. An “empty-detection” algorithm is usually required to distinguish “G” nucleotides and empty wells. SUMMARY
[0006] An aspect of the disclosure is directed to a method of sequencing polynucleotides bound to a flowcell, including: detecting fluorescent emissions from a first labeled nucleotide at a first wavelength and a first intensity; detecting fluorescent emissions from a second labeled nucleotide at a second wavelength and a second intensity, wherein the first wavelength is different from the second wavelength; detecting fluorescent emissions from a third labeled nucleotide at the first and second wavelengths at third and fourth intensities; detecting fluorescent emissions from a fourth labeled nucleotide at the first and second wavelengths at fifth and sixth intensities; and determining the sequence of the polynucleotides based on the detected fluorescent emissions and intensities. I
[0007] In some embodiments, the flowcell may comprise wells configured to bind polynucleotides. In some embodiments, systems and methods may be configured for one or both of one excitation, two channel chemistry or two excitation, two channel chemistry. In some embodiments, the first labeled nucleotide and the second labeled nucleotide may be labeled with fluorescent dyes having the same excitation wavelength. In some embodiments, the first labeled nucleotide and the second labeled nucleotide may be labeled with fluorescent dyes having different stokes shifts.
[0008] In some embodiments, at least two of the first through sixth intensities are greater than zero. In some embodiments, at least three of the first through sixth intensities are greater than zero. In some embodiments, at least four of the first through sixth intensities are greater than zero. In some embodiments, at least five of the first through sixth intensities are greater than zero. In some embodiments, all of the first through sixth intensities are greater than zero. Insome embodiments, the first intensity is less than the third intensity; further wherein the second intensity is less than the sixth intensity; and wherein four labeled nucleotides form a diamond configuration. In some embodiments, the ratio of the third intensity to the fourth intensity is approximately equal to the ratio of the sixth intensity to the fifth intensity.
[0009] In some embodiments, methods may comprise the step of detecting fluorescent emissions from a fifth labeled nucleotide at the first and second wavelengths at seventh and eighth intensities. In some embodiments, a method may further include the step of determining the presence of an empty well based on a substantial absence of the detected fluorescent emissions and intensities.
[0010] In some embodiments, a kit for determining the sequence of a polynucleotide may include: a first mixture of a first nucleotide-first fluorescent dye conjugate detectable in a first wavelength channel and a first nucleotide-second fluorescent dye conjugate detectable in a second wavelength channel; a second mixture of a second nucleotide-first fluorescent dye conjugate detectable in the first wavelength channel and a second nucleotide-second fluorescent dye conjugate detectable in the second wavelength channel; a third mixture of a third nucleotide-first fluorescent dye conjugate detectable in the first wavelength channel and a third nucleotide-second fluorescent dye conjugate detectable in the second wavelength channel; and a fourth mixture of a fourth nucleotide-first fluorescent dye conjugate detectable in the first wavelength channel and a fourth nucleotide-second fluorescent dye conjugate detectable in the second wavelength channel. In some embodiments, the ratio of intensity of emissions in the first channel versus the second channel for the third mixture is approximately equal to the ratio of intensity of emissions in the second channel versus the first channel for the fourth mixture. In some embodiments, a kit may comprise mixtures of labeled nucleotides that result in a scatterplot with nucleotide clouds in one of the following configurations, diamond, L-Shaped, and diagonal configurations.
[0011] In some aspects the disclosure is related to a system for determining the sequence of a polynucleotide, including: a machine-readable memory; and a processor configured to execute machine-readable instructions, which, when executed by the processor, cause the system to perform steps including: detecting fluorescent emissions from a first labeled nucleotide at a first wavelength and a first intensity; detecting fluorescent emissions from a second labeled nucleotide at a second wavelength and a second intensity, wherein the first wavelength is different from the second wavelength; detecting fluorescent emissions from a third labeled nucleotide at thefirst and second wavelengths at third and fourth intensities; detecting fluorescent emissions from a fourth labeled nucleotide at the first and second wavelengths at fifth and sixth intensities; and determining the sequence of the polynucleotides based on the detected fluorescent emissions and intensities.
[0012] An aspect of the disclosure relates to mitigating potential issues with assigning bases (or empty wells) that have a low probability of occurring, due to, for example, low base diversity. Clusters that are not incorporating or are empty wells on a patterned flowcell generate may erroneously generate G basecalls because the G nucleotide is not labeled. Furthermore, when clusters sequence the whole insert and are no longer incorporating, poly G will be called for the remainder of the read. By partially labeling the G, a fifth cloud may be present in a two- dimensional plot of intensities in 2-channel SBS. An aspect of the disclosure is directed to addressing the fifth cloud. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Features of examples of the present disclosure will become apparent by reference to the following detailed description and drawings, in which like reference numerals correspond to similar, though perhaps not identical, components. For the sake of brevity, reference numerals or features having a previously described function may or may not be described in connection with other drawings in which they appear. While the disclosure has been illustrated and described in detail in the drawings and foregoing description, such illustration and description are to be considered illustrative or exemplary and not restrictive. The disclosure is not limited to the disclosed embodiments. Variations to the disclosed embodiments can be understood and effected by those skilled in the art in practicing the claimed disclosure, from a study of the drawings, the disclosure and the appended claims.
[0014] FIG. 1A schematically illustrates an example sequencing system which can perform embodiments of the disclosed sequencing technology.
[0015] FIG. 1B schematically illustrates an example imaging system to be used in embodiments of the disclosed sequencing technology.
[0016] FIG.2 is a flowchart illustrating a process of sequencing polynucleotides bound to a flow cell with two detection channels.
[0017] FIG.3A shows a two-dimensional scatterplot of an experiment to demonstrate labeled bases with one excitation, two channel (1Ex-2Ch) chemistry.
[0018] FIG. 3B shows an illustration of one embodiment of a 1Ex-2Ch chemistry method employing a labeled G nucleotide.
[0019] FIG. 4 shows an embodiment of four dyes with emission spectra compatible with 1Ex-2Ch chemistry.
[0020] FIG.5A shows an embodiment of three dyes with emission spectra compatible with 1Ex-2Ch chemistry.
[0021] FIG. 5B shows experimental results of a typical 1Ex-2Ch system, with the positions of the nucleotides matching those in Fig.5A.
[0022] FIG. 6 depicts three example embodiments of Diamond Scatterplots compatible with two excitation, two channel (2Ex-2Ch) chemistry.
[0023] FIG.7A illustrates an X-Y scatterplot of the fluorescent emissions of a system where all four nucleotides are labeled.
[0024] FIG. 7B illustrates an X-Y scatterplot of the fluorescent emissions of a system where all four nucleotides are labeled.
[0025] FIG.7C illustrates an X-Y scatterplot of the fluorescent emissions of a system where all four nucleotides are labeled.
[0026] FIG.8 shows an X-Y scatterplot of the fluorescent emissions of a system where three nucleotides are labeled, and a “G” nucleotide is unlabeled.
[0027] FIG. 9A illustrates Fig. 9A displays an experimental example of an L-shape scatterplot including an X-Y scatterplot of the fluorescent emissions of a two-channel system, where three nucleotides are labeled, and a “G” nucleotide is unlabeled.
[0028] Fig. 9B shows an experimental example of an L-shaped scatterplot with two nucleotides encoded at different intensities of the signal state of channel two.
[0029] FIG. 10A displays a panel of the results of five experimental examples in five different X-Y scatterplots of differing degrees of labeled G on a NextSeq2K system.
[0030] FIG. 10B shows a graph of the resulting Q30 scores from each of the experiments in Fig.10A.
[0031] FIG. 11 shows an X-Y scatterplot of the fluorescent emissions of a system where four nucleotides are labeled, and empty wells are distinguishable from the four labeled nucleotides.
[0032] FIG. 12 schematically illustrates a system including a memory comprising a sequencing module. DETAILED DESCRIPTION
[0033] All patents, applications, published applications and other publications referred to herein are incorporated herein by reference to the referenced material and in their entireties. If a term or phrase is used herein in a way that is contrary to or otherwise inconsistent with a definition set forth in the patents, applications, published applications and other publications that are herein incorporated by reference, the use herein prevails over the definition that is incorporated herein by reference.
[0034] One aspect of the disclosure provides for systems and methods for determining the sequence of a polynucleotide by labeling at least four nucleotides in a sequencing by synthesis (SBS) system. In some embodiments, the at least four nucleotides may be labeled with a mixture of fluorescent dyes that results in alternative scatterplot shapes in a X-Y scatterplot of emissions in a first channel and emissions in a second channel. The scatterplots are “alternative” in that they differ from classically shaped scatterplots which have a set of cloud formations which arrange in the shape of a square or rectangle, as discussed in more detail below. In some embodiments, the polynucleotides have four unmodified bases, and each of the A, T, C and G nucleotides are labeled. In this embodiment, a two-dimensional scatterplot showing the measured wavelengths and intensities from each fluorescent label will have at least four distinct cloud formations, one for each fluorescent label.
[0035] As an example of one labeling arrangement which results in an alternative scatterplot, a first base A may be labeled with a first fluorophore bound to only a specific percentage of the A bases being incorporated into an insert during SBS sequencing reactions. For example, only 60% of the A bases may be labeled with the first fluorophore. A second base G may be labeled with a second fluorophore bound to only a percentage of the G bases being incorporated into the insert during SBS sequencing reactions. For example, only 60% of the G bases may be labeled with the second fluorophore. A third base T may be labeled with both the first and second fluorophore on each of the T bases to be incorporated into the insert during SBS sequencingreactions, but where the majority of labeled T bases are labeled with the first fluorophore. For example, about 66% of the T bases may be labeled with the first fluorophore and about 33% of the T bases may be labeled with the second fluorophore. And finally, the fourth base C may be labeled with both the first and second fluorophore, but with a majority of labeled C bases being labeled with the second fluorophore and a minority of the bases being labeled with the first fluorophore. For example, about 33% of the C bases may be labeled with the first fluorophore and about 66% of the C bases may be labeled with the second fluorophore. In this embodiment, some wells on a flowcell will be empty and not contain any clusters of inserts as cluster formation does not always occur in every flowcell well. In these empty wells, a dark state cloud may be formed on the two- dimensional scatterplot from the lack of fluorescent signals corresponding to the empty wells. Thus, in this embodiment there may be five clouds formed on a two-dimensional scatterplot. In the above labeling arrangement, a two-dimensional scatterplot which is formed from the various labeled nucleotides will have a diamond shape, as will be disclosed more fully with reference to Figs. 3A and 3B below.
[0036] One other aspect of the disclosure is a method of sequencing polynucleotides bound in clusters to a flowcell by using different amounts of two different labels some of the nucleotides. This method can include detecting fluorescent emissions from a first labeled nucleotide at a first wavelength and a first intensity. For example, wherein a first reversibly terminated nucleotide A is labeled with a blue dye which emits light at a first wavelength. The method then includes detecting fluorescent emissions from a second reversibly terminated nucleotide G which is labeled with a green dye which emits light at a second wavelength, and wherein the first wavelength is different from the second wavelength. The method then detects fluorescent emissions from a third reversibly terminated nucleotide C which is labeled with both the blue and green dyes at predetermined concentrations and which emits light at the first and second wavelengths. Finally, the method includes detecting fluorescent emissions from a fourth reversibly terminated nucleotide T which is labeled with different concentrations of both the blue and green dyes and which emit light at the first and second wavelengths. It should be realized that the intensity of fluorophores used in embodiments of the invention can be any detectable percentage of a full-intensity fluorophore by mixing labeled and unlabeled nucleotides together prior to contact with the flowcell. For example, a mixture of 50% labeled C nucleotides and 50% unlabeled nucleotides will result in the sequencing system detecting an intensity of C nucleotidesthat is 50% of the expected full fluorescence of the fluorophore as compared to a mixture of all labeled C nucleotides. The ratio of emissions at the first and second wavelengths for nucleotide T, for example, may be adjusted to mirror the ratio of emissions at the first and second wavelengths for nucleotide C. Any ratio of emissions for a base may be employed where the emissions result in a distinguishable intensity from the other bases.
[0037] Moreover, in some aspects, the disclosure herein relates to a system for performing the above methods of sequencing polynucleotides bound to a flowcell. The system would include a memory linked to one or more processors which are configured to execute machine-readable instructions, which, when executed by the one or more processors, cause the system to perform the above method steps.
[0038] Another embodiment is related to the above method, but instead of using nucleotides which have predetermined concentrations of two different fluorophores, to instead use an increased intensity dye to increase the intensity of a nucleotide that may be encoded within a channel. In some embodiments, the method may produce nucleotide clouds in an L-shaped configuration. For example, the method may include labeling a first labeled nucleotide resulting in emissions at a first wavelength and a first intensity. For example, wherein a first reversibly terminated nucleotide T is labeled with a green dye which emits light at a first wavelength, the method then includes detecting fluorescent emissions from a second reversibly terminated nucleotide A which is labeled with a green dye which emits light at the first wavelength, but at a second intensity that is less than the first intensity. The method then detects fluorescent emissions from a third reversibly terminated nucleotide C which is labeled with a blue dye that emits light at a second wavelength that is different from the first wavelength. The method then includes detecting the absence of fluorescent emissions from an unlabeled fourth reversibly terminated nucleotide G. Moreover, during the library preparation process, particularly modified nucleotides, such as met- C, may have a specific percentage of the modified nucleotides labeled such that a detectable reduction in the light being emitted at the first wavelength can be detected. Definitions
[0039] The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.
[0040] It is noted that, as used in this specification and the appended claims, the singular forms "a", "an" and "the" include plural referents unless expressly and unequivocally limited to one referent. It will be apparent to those skilled in the art that various modifications and variations can be made to various embodiments described herein without departing from the spirit or scope of the present teachings. Thus, it is intended that the various embodiments described herein cover other modifications and variations within the scope of the appended claims and their equivalents.
[0041] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as is commonly understood by one of ordinary skill in the art. The use of the term “including” as well as other forms, such as “include”, “includes,” and “included,” is not limiting. The use of the term “having” as well as other forms, such as “have”, “has,” and “had,” is not limiting. As used in this specification, whether in a transitional phrase or in the body of the claim, the terms “comprise(s)” and “comprising” are to be interpreted as having an open-ended meaning. That is, the above terms are to be interpreted synonymously with the phrases “having at least” or “including at least.” For example, when used in the context of a process, the term “comprising” means that the process includes at least the recited steps but may include additional steps. When used in the context of a compound, composition, or device, the term “comprising” means that the compound, composition, or device includes at least the recited features or components, but may also include additional features or components.
[0042] All references cited herein are incorporated herein by reference in their entirety. To the extent publications and patents or patent applications incorporated by reference contradict the disclosure contained in the specification, the specification is intended to supersede and / or take precedence over any such contradictory material.
[0043] As used herein, common organic abbreviations are defined as follows: °C Temperature in degrees Centigrade dATP Deoxyadenosine triphosphate dCTP Deoxycytidine triphosphate dGTP Deoxyguanosine triphosphate dTTP Deoxythymidine triphosphate ddNTP Dideoxynucleotide triphosphate ffA Fully functionalized A nucleotide ffC Fully functionalized C nucleotide ffG Fully functionalized G nucleotide ffN Fully functionalized nucleotideffT Fully functionalized T nucleotide h Hour(s) RT Room temperature SBS Sequencing by Synthesis A Adenosine (may refer to nucleotide base calls) C Cytosine G Guanine T Thymine
[0044] As used herein, a “peptide” refers to two or more amino acids joined together by an amide bond (that is, a “peptide bond”). Peptides comprise up to or include 50 amino acids. Peptides may be linear or cyclic. Peptides may be Į, ȕ, Ȗ, į, or higher, or mixed. Peptides may comprise any mixture of amino acids as defined herein, such as comprising any combination of D, L, Į, ȕ, Ȗ, į, or higher amino acids.
[0045] As used herein, a “protein” refers to an amino acid sequence having 51 or more amino acids.
[0046] As used herein, “nucleobase” is a heterocyclic base such as adenine, guanine, cytosine, thymine, uracil, inosine, xanthine, hypoxanthine, or a heterocyclic derivative, analog, or tautomer thereof. A nucleobase can be naturally occurring or synthetic. Non-limiting examples of nucleobases are adenine, guanine, thymine, cytosine, uracil, xanthine, hypoxanthine, 8-azapurine, purines substituted at the 8 position with methyl or bromine, 9-oxo-N6-methyladenine, 2- aminoadenine, 7-deazaxanthine, 7-deazaguanine, 7-deaza-adenine, N4-ethanocytosine, 2,6- diaminopurine, N6-ethano-2,6-diaminopurine, 5-methylcytosine, 5-(C3-C6)- alkynylcytosine, 5- fluorouracil, 5-bromouracil, thiouracil, pseudoisocytosine, 2-hydroxy-5-methyl-4- triazolopyridine, isocytosine, isoguanine, inosine, 7,8-dimethylalloxazine, 6-dihydrothymine, 5,6- dihydrouracil, 4-methyl-indole, ethenoadenine and the non-naturally occurring nucleobases described in U.S. Pat. Nos. 5,432,272 and 6,150,510 and PCT applications WO 92 / 002258, WO 93 / 10820, WO 94 / 22892, and WO 94 / 24144, and Fasman ("Practical Handbook of Biochemistry and Molecular Biology", pp.385-394, 1989, CRC Press, Boca Raton, LO), all herein incorporated by reference in their entireties.
[0047] As used herein, the term “nucleotide” is intended to mean a molecule that includes a sugar and at least one phosphate group, and in some examples also includes a nucleobase. A nucleotide that lacks a nucleobase may be referred to as “abasic.” In some embodiments, a “nucleotide” includes a nitrogen containing heterocyclic base, a sugar, and one ormore phosphate groups. Nucleotides are monomeric units of a nucleic acid sequence. Examples of nucleotides include, for example, ribonucleotides or deoxyribonucleotides. In ribonucleotides (RNA), the sugar is a ribose, and in deoxyribonucleotides (DNA), the sugar is a deoxyribose, i.e., a sugar lacking a hydroxyl group that is present at the 2' position in ribose. The nitrogen containing heterocyclic base can be a purine base or a pyrimidine base. Purine bases include adenine (A) and guanine (G), and modified derivatives or analogs thereof. Pyrimidine bases include cytosine (C), thymine (T), and uracil (U), and modified derivatives or analogs thereof. The C-1 atom of deoxyribose is bonded to N-1 of a pyrimidine or N-9 of a purine. The phosphate groups may be in the mono-, di-, or tri-phosphate form. These nucleotides are natural nucleotides, but it is to be further understood that non-natural nucleotides, modified nucleotides or analogs of the aforementioned nucleotides can also be used.
[0048] Examples of nucleotides may include deoxyribonucleotides, modified deoxyribonucleotides, ribonucleotides, modified ribonucleotides, peptide nucleotides, modified peptide nucleotides, modified phosphate sugar backbone nucleotides, and mixtures thereof. Examples of nucleotides include adenosine monophosphate (AMP), adenosine diphosphate (ADP), adenosine triphosphate (ATP), thymidine monophosphate (TMP), thymidine diphosphate (TDP), thymidine triphosphate (TTP), cytidine monophosphate (CMP), cytidine diphosphate (CDP), cytidine triphosphate (CTP), guanosine monophosphate (GMP), guanosine diphosphate (GDP), guanosine triphosphate (GTP), uridine monophosphate (UMP), uridine diphosphate (UDP), uridine triphosphate (UTP), deoxyadenosine monophosphate (dAMP), deoxyadenosine diphosphate (dADP), deoxyadenosine triphosphate (dATP), deoxythymidine monophosphate (dTMP), deoxythymidine diphosphate (dTDP), deoxythymidine triphosphate (dTTP), deoxycytidine diphosphate (dCDP), deoxycytidine triphosphate (dCTP), deoxyguanosine monophosphate (dGMP), deoxyguanosine diphosphate (dGDP), deoxyguanosine triphosphate (dGTP), deoxyuridine monophosphate (dUMP), deoxyuridine diphosphate (dUDP), and deoxyuridine triphosphate (dUTP).
[0049] Examples of nucleotides may also be intended to encompass any nucleotide analogue which is a type of nucleotide that includes a modified nucleobase, sugar, backbone, and / or phosphate moiety compared to naturally occurring nucleotides. Nucleotide analogues also may be referred to as “modified nucleic acids.” Example modified nucleobases include inosine, xathanine, hypoxathanine, isocytosine, isoguanine, 2-aminopurine, 5-methylcytosine, 5-hydroxymethyl cytosine, 2-aminoadenine, 6-methyl adenine, 6-methyl guanine, 2-propyl guanine, 2-propyl adenine, 2-thiouracil, 2-thiothymine, 2-thiocytosine, 15-halouracil, 15-halocytosine, 5- propynyl uracil, 5-propynyl cytosine, 6-azo uracil, 6-azo cytosine, 6-azo thymine, 5-uracil, 4- thiouracil, 8-halo adenine or guanine, 8-amino adenine or guanine, 8-thiol adenine or guanine, 8- thioalkyl adenine or guanine, 8-hydroxyl adenine or guanine, 5-halo substituted uracil or cytosine, 7-methylguanine, 7-methyladenine, 8-azaguanine, 8-azaadenine, 7-deazaguanine, 7- deazaadenine, 3-deazaguanine, 3-deazaadenine or the like. As is known in the art, certain nucleotide analogues cannot become incorporated into a polynucleotide, for example, nucleotide analogues such as adenosine 5'-phosphosulfate. Nucleotides may include any suitable number of phosphates, e.g., three, four, five, six, or more than six phosphates. Nucleotide analogues also include locked nucleic acids (LNA), peptide nucleic acids (PNA), and 5-hydroxylbutynl-2'- deoxyuridine (“super T”).
[0050] In some embodiments, the term “modification” as used herein is intended to refer not only to a chemical modification of a nucleic acids, but also to a variation in nucleic acid conformation or composition, interaction of an agent with a nucleic acid (e.g., bound to the nucleic acid), and other perturbations associated with the nucleic acid. As such, a location or position of a modification is a locus (e.g., a single nucleotide or multiple contiguous or noncontiguous nucleotides) at which such modification occurs within the nucleic acid. For a double-stranded template, such a modification may occur in the strand complementary to a nascent strand synthesized by a polymerase processing the template or may occur in the displaced strand. For example, modified nucleotides may include 5-methylcytosine, N6-methyladenosine, N3- methyladenosine, N7-methylguanosine, 5-hydroxymethylcytosine, pseudouridine, thiouridine, isoguanosine, isocytosine, dihydrouridine, queuosine, wyosine, inosine, triazole, diaminopurine, ȕ-D-glucopyranosyloxymethyluracil (a.k.a., ȕ-D-glucosyl-HOMedU, ȕ-glucosyl- hydroxymethyluracil, “dJ,” or “base J”), 8-oxoguanosine, and 2ƍ-O-methyl derivatives of adenosine, cytidine, guanosine, and uridine. Modified DNA and RNA bases are further described, for example, in Narayan P, et al. (1987) Mol Cell Biol 7(4):1572-5; Horowitz S, et al. (1984) Proc Natl Acad Sci U.S.A. 81(18):5667-71; “RNA's Outfits: The nucleic acid has dozens of chemical costumes,” (2009) C&EN; 87(36):65-68; Kriaucionis, et al. (2009) Science 324 (5929): 929-30; and Tahiliani, et al. (2009) Science 324 (5929): 930-35; Matray, et al. (1999) Nature 399(6737):704-8; Ooi, et al. (2008) Cell 133: 1145-8; Petersson, et al. (2005) J Am Chem Soc.127(5):1424-30; Johnson, et al. (2004) 32(6):1937-41; Kimoto, et al. (2007) Nucleic Acids Res. 35(16):5360-9; Ahle, et al. (2005) Nucleic Acids Res 33(10):3176; Krueger, et al., Curr Opinions in Chem Biology 2007, 11(6):588); Krueger, et al. (2009) Chemistry & Biology 16(3):242; McCullough, et al. (1999) Annual Rev of Biochem 68:255; Liu, et al. (2003) Science 302(5646):868-71; Limbach, et al. (1994) Nucl. Acids Res.22(12):2183-2196; Wyatt, et al. (1953) Biochem. J.55:774-782; Josse, et al. (1962) J. Biol. Chem.237:1968-1976; Lariviere, et al. (2004) J. Biol. Chem. 279:34715-34720; and in International Application Publication No. WO / 2009 / 037473, the disclosures of which are incorporated herein by reference in their entireties.
[0051] Modifications may further include the presence of non-natural base pairs in the nucleic acid, including but not limited to hydroxypyridone and pyridopurine homo- and hetero- base pairs, pyridine-2,6-dicarboxylate and pyridine metallo-base pairs, pyridine-2,6- dicarboxamide and a pyridine metallo-base pairs, metal-mediated pyrimidine base pairs T-Hg(II)- T and C-Ag(I)-C, and metallo-homo-basepairs of 2,6-bis(ethylthiomethyl)pyridine nucleobases Spy, and alkyne-, enamine-, alcohol-, imidazole-, guanidine-, and pyridyl-substitutions to the purine or pyridimine base (Wettig, et al. (2003) J Inorg Biochem 94:94-99; Clever, et al. (2005) Angew Chem Int Ed 117:7370-7374; Schlegel, et al. (2009) Org Biomol Chem 7(3):476-82; Zimmerman, et al. (2004) Bioorg Chem 32(1):13-25; Yanagida, et al. (2007) Nucleic Acids Symp Ser (Oxf) 51:179-80; Zimmerman (2002) J Am Chem Soc 124(46):13684-5; Buncel, et al. (1985) Inorg Biochem 25:61-73; Ono, et al. (2004) Angew Chem 43:4300-4302; Lee, et al. (1993) Biochem Cell Biol 71:162-168; Loakes, et al. (2009), Chem Commun 4619-4631; and Seo, et al. (2009) J Am Chem Soc 131:3246-3252, the disclosures of which are incorporated herein by reference in their entireties). Other types of modifications include, e.g, a nick, a missing base (e.g., apurinic or apyridinic sites), a ribonucleoside (or modified ribonucleoside) within a deoxyribonucleoside-based nucleic acid, a deoxyribonucleoside (or modified deoxyribonucleoside) within a ribonucleoside-based nucleic acid, a pyrimidine dimer (e.g., thymine dimer or cyclobutane pyrimidine dimer), a cis-platin crosslinking, oxidation damage, hydrolysis damage, other methylated bases, bulky DNA or RNA base adducts, photochemistry reaction products, interstrand crosslinking products, mismatched bases, and other types of “damage” to the nucleic acid. Modified nucleotides can be caused by exposure of the DNA to radiation (e.g., UV), carcinogenic chemicals, crosslinking agents (e.g., formaldehyde), certainenzymes (e.g., nickases, glycosylases, exonucleases, methylases, other nucleases, glucosyltransferases, etc.), viruses, toxins and other chemicals, thermal disruptions, and the like.
[0052] As used herein, the term “polynucleotide” refers to a molecule that includes a sequence of nucleotides that are bonded to one another. A polynucleotide is one nonlimiting example of a polymer. Examples of polynucleotides include deoxyribonucleic acid (DNA), ribonucleic acid (RNA), and analogues thereof such as locked nucleic acids (LNA) and peptide nucleic acids (PNA). A polynucleotide may be a single stranded sequence of nucleotides, such as RNA or single stranded DNA, a double stranded sequence of nucleotides, such as double stranded DNA, or may include a mixture of a single stranded and double stranded sequences of nucleotides. Double stranded DNA (dsDNA) includes genomic DNA, and PCR and amplification products. Single stranded DNA (ssDNA) can be converted to dsDNA and vice-versa. Polynucleotides may include non-naturally occurring DNA, such as enantiomeric DNA, LNA, or PNA. The precise sequence of nucleotides in a polynucleotide may be known or unknown. The following are examples of polynucleotides: a gene or gene fragment (for example, a probe, primer, expressed sequence tag (EST) or serial analysis of gene expression (SAGE) tag), genomic DNA, genomic DNA fragment, exon, intron, messenger RNA (mRNA), transfer RNA, ribosomal RNA, ribozyme, cDNA, recombinant polynucleotide, synthetic polynucleotide, branched polynucleotide, plasmid, vector, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probe, primer or amplified copy of any of the foregoing.
[0053] The terms “oligonucleotide” and “polynucleotide” may be used interchangeably herein. The different terms are not intended to denote any particular difference in size, sequence, or other property unless specifically indicated otherwise. For clarity of description, the terms may be used to distinguish one species of polynucleotide from another when describing a particular method or composition that includes several polynucleotide species.
[0054] The term “nucleic acid” and “polynucleotide” may be used interchangeably to refer to a deoxyribonucleotide or ribonucleotide polymer in either single- or double-stranded form, and unless otherwise limited, encompasses known analogs of natural nucleotides that hybridize to nucleic acids in manner similar to naturally occurring nucleotides, such as peptide nucleic acids (PNAs) and phosphorothioate DNA. Unless otherwise indicated, a particular nucleic acid sequence includes the complementary sequence thereof. Nucleotides include, but are not limited to, ATP, dATP, CTP, dCTP, GTP, dGTP, UTP, TTP, dUTP, 5-methyl-CTP, 5-methyl-dCTP, ITP, dITP, 2-amino-adenosine-TP, 2-amino-deoxyadenosine-TP, 2-thiothymidine triphosphate, pyrrolo- pyrimidine triphosphate, and 2-thiocytidine, as well as the alphathiotriphosphates for all of the above, and 2ƍ-O-methyl-ribonucleotide triphosphates for all the above bases. Modified bases include, but are not limited to, 5-Br-UTP, 5-Br-dUTP, 5-F-UTP, 5-F-dUTP, 5-propynyl dCTP, and 5-propynyl-dUTP.
[0055] As used herein, a “nucleoside” is structurally similar to a nucleotide, but is missing the phosphate moieties. An example of a nucleoside analogue would be one in which the label is linked to the base and there is no phosphate group attached to the sugar molecule. The term “nucleoside” is used herein in its ordinary sense as understood by those skilled in the art. Examples include, but are not limited to, a ribonucleoside comprising a ribose moiety and a deoxyribonucleoside comprising a deoxyribose moiety. A modified pentose moiety is a pentose moiety in which an oxygen atom has been replaced with a carbon and / or a carbon has been replaced with a sulfur or an oxygen atom. A “nucleoside” is a monomer that can have a substituted base and / or sugar moiety. Additionally, a nucleoside can be incorporated into larger DNA and / or RNA polymers and oligomers.
[0056] The term “purine base” is used herein in its ordinary sense as understood by those skilled in the art, and includes its tautomers. Similarly, the term “pyrimidine base” is used herein in its ordinary sense as understood by those skilled in the art, and includes its tautomers. A non-limiting list of optionally substituted purine-bases includes purine, adenine, guanine, hypoxanthine, xanthine, alloxanthine, 7-alkylguanine (e.g. 7-methylguanine), theobromine, caffeine, uric acid and isoguanine. Examples of pyrimidine bases include, but are not limited to, cytosine, thymine, uracil, 5,6-dihydrouracil and 5-alkylcytosine (e.g., 5-methylcytosine).
[0057] The term “nucleobase” as used herein, is a purine base or a pyrimidine base. Non-limiting examples of purine nucleobases include adenine (A), guanine (G), and derivatives or analogs thereof. Non-limiting examples of pyrimidine nucleobases include cytosine (C), thymine (T), uracil (U), and derivatives or analogs thereof.
[0058] As used herein, when an oligonucleotide or polynucleotide is described as “comprising” a nucleoside or nucleotide described herein, it means that the nucleoside or nucleotide described herein forms a covalent bond with the oligonucleotide or polynucleotide. Similarly, when a nucleoside or nucleotide is described as part of an oligonucleotide or polynucleotide, such as “incorporated into” an oligonucleotide or polynucleotide, it means that thenucleoside or nucleotide described herein forms a covalent bond with the oligonucleotide or polynucleotide. In some such embodiments, the covalent bond is formed between a 3^ hydroxy group of the oligonucleotide or polynucleotide with the 5^ phosphate group of a nucleotide described herein as a phosphodiester bond between the 3^ carbon atom of the oligonucleotide or polynucleotide and the 5^ carbon atom of the nucleotide.
[0059] As used herein, the term “array” refers to a population of different probe molecules that are attached to one or more substrates such that the different probe molecules can be differentiated from each other according to relative location. An array can include different probe molecules that are each located at a different addressable location on a substrate. Alternatively, or additionally, an array can include separate substrates each bearing a different probe molecule, wherein the different probe molecules can be identified according to the locations of the substrates on a surface to which the substrates are attached or according to the locations of the substrates in a liquid. Exemplary arrays in which separate substrates are located on a surface include, without limitation, those including beads in wells as described, for example, in U.S. Patent No. 6,1055,331 B1, US 2002 / 0102578 and PCT Publication No. WO 00 / 63437. Exemplary formats that can be used in the invention to distinguish beads in a liquid array, for example, using a microfluidic device, such as a fluorescent activated cell sorter (FACS), are described, for example, in US Pat. No. 6,524,793. Further examples of arrays that can be used in the invention include, without limitation, those described in U.S. Pat Nos. 5,329,807; 5,336,1027; 5,561,071; 5,583,911; 5,658,734; 5,837,858; 5,874,919; 5,919,523; 6,836,969; 6,987,768; 6,987,776; 6,988,920; 6,997,006; 6,991,893; 6,1046,313; 6,316,949; 6,382,591; 6,514,751 and 6,610,382; and WO 93 / 17126; WO 95 / 11995; WO 95 / 35505; EP 742987; and EP 799897.
[0060] A nucleotide analog may be attached to or associated with one or more photo- detectable labels to provide a detectable signal. In some embodiments, a photo-detectable label may be a fluorescent compound, such as a small molecule fluorescent label. Fluorescent molecules (fluorophores) suitable as a fluorescent label include, but are not limited to: 1,5 IAEDANS; 1,8- ANS; 4-methylumbelliferone; 5-carboxy-2,7-dichlorofluorescein; 5-carboxyfluorescein (5-FAM); fluorescein amidite (FAM); 5-carboxynapthofluorescein; tetrachloro-6-carboxyfluorescein (TET); hexachloro-6-carboxyfluorescein (HEX); 2,7-dimethoxy-4,5-dichloro-6-carboxyfluorescein (JOE); VIC®; NED™; tetramethylrhodamine (TMR); 5-carboxytetramethylrhodamine (5- TAMRA); 5-HAT (Hydroxy Tryptamine); 5-hydroxy tryptamine (HAT); 5-ROX (carboxy-X-rhodamine); 6-carboxyrhodamine 6G; 6-JOE; Light Cycler® red 610; Light Cycler® red 640; Light Cycler® red 670; Light Cycler® red 705; 7-amino-4-methylcoumarin; 7-aminoactinomycin D (7-AAD); 7-hydroxy-4-methylcoumarin; 9-amino-6-chloro-2-methoxyacridine; 6-methoxy-N- (4-aminoalkyl)quinolinium bromide hydrochloride (ABQ); Acid Fuchsin; ACMA (9-amino-6- chloro-2-methoxyacridine); Acridine Orange; Acridine Red; Acridine Yellow; Acriflavin; Acriflavin Feulgen SITSA; AFPs-AutoFluorescent Protein-(Quantum Biotechnologies); Texas Red; Texas Red-X conjugate; Thiadicarbocyanine (DiSC3); Thiazine Red R; Thiazole Orange; Thioflavin 5; Thioflavin S; Thioflavin TCN; Thiolyte; Thiozole Orange; Tinopol CBS (Calcofluor White); TMR; TO-PRO-1; TO-PRO-3; TO-PRO-5; TOTO-1; TOTO-3; TriColor (PE-Cy5); TRITC (TetramethylRodamine-lsoThioCyanate); True Blue; TruRed; Ultralite; Uranine B; Uvitex SFC; WW 781; X-Rhodamine; X-Rhodamine-5-(and-6)-Isothiocyanate (5(6)-XRITC); Xylene Orange; Y66F; Y66H; Y66W; YO-PRO-1; YO-PRO-3; YOYO-1; interchelating dyes such as YOYO-3, Sybr Green, Thiazole orange; members of the Alexa Fluor® dye series (from Molecular Probes / Invitrogen) which cover a broad spectrum and match the principal output wavelengths of common excitation sources such as Alexa Fluor 350, Alexa Fluor 405, 430, 488, 500, 514, 532, 546, 555, 568, 594, 610, 633, 635, 647, 660, 680, 700, and 750; members of the Cy Dye fluorophore series (GE Healthcare), also covering a wide spectrum such as Cy3, Cy3B, Cy3.5, Cy5, Cy5.5, Cy7; members of the Oyster® dye fluorophores (Denovo Biolabels) such as Oyster- 500, -550, -556, 645, 650, 656; members of the DY-Labels series (Dyomics), for example, with maxima of absorption that range from 418 nm (DY-415) to 844 nm (DY-831) such as DY-415, - 495, -505, -547, -548, -549, -550, -554, -555, -556, -560, -590, -610, -615, -630, -631, -632, -633, -634, -635, -636, -647, -648, -649, -650, -651, -652, -675, -676, -677, -680, -681, -682, -700, -701, -730, -731, -732, -734, -750, -751, -752, -776, -780, -781, -782, -831, -480XL, -481XL, -485XL, -510XL, -520XL, -521XL; members of the ATTO series of fluorescent labels (ATTO-TEC GmbH) such as ATTO 390, 425, 465, 488, 495, 520, 532, 550, 565, 590, 594, 610, 611X, 620, 633, 635, 637, 647, 647N, 655, 680, 700, 725, 740; members of the CAL Fluor® series or Quasar® series of dyes (Biosearch Technologies) such as CAL Fluor® Gold 540, CAL Fluor® Orange 560, Quasar® 570, CAL Fluor® Red 590, CAL Fluor® Red 610, CAL Fluor® Red 635, Quasar® 570, and Quasar® 670. In some embodiments, a first photo-detectable label interacts with a second photo-detectable moiety to modify the detectable signal, e.g., via fluorescence resonance energy transfer (“FRET”; also known as Förster resonance energy transfer).
[0061] The fluorescent labels utilized by the systems and methods disclosed herein can have different peak absorption wavelengths, for example, ranging from 400 nm to 800 nm. In some embodiments, the peak absorption wavelengths of the fluorescent labels can be, or be about, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720, 730, 740, 750, 760, 770, 780, 790, 800 nm, or a number or a range between any two of these values. In some embodiments the peak absorption wavelengths of the fluorescent labels can be at least, or at most, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720, 730, 740, 750, 760, 770, 780, 790, or 800 nm.
[0062] The fluorescent labels can have different peak emission wavelength, for example, ranging from 400 nm to 800 nm. In some embodiments, the peak emission wavelengths of the fluorescent labels can be, or be about, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720, 730, 740, 750, 760, 770, 780, 790, 800 nm, or a number or a range between any two of these values. In some embodiments the peak emission wavelengths of the fluorescent labels can be at least, or at most, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720, 730, 740, 750, 760, 770, 780, 790, or 800 nm.
[0063] The fluorescent labels can have different Stokes shift, for example, ranging from 10 nm to 200 nm. In some embodiments, the stoke shift can be, or be about, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200 nm, or a number or a range between any two of these values. In some embodiments, the stoke shift can be at least, or at most, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 nm.
[0064] In some embodiments, the distance between the peak emission wavelengths of any two fluorescent labels can vary, for example, ranging from 10 nm to 200 nm. In some embodiments, the distance between the peak emission wavelengths of any two fluorescent labels can be, or be about, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200 nm, or a number or a range between any two of these values. In some embodiments, the distance between the peak emission wavelengths of any two fluorescent labels can be at least, orat most, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 nm.
[0065] A “light source” may be any device capable of emitting energy along the electromagnetic spectrum. A light source may be a source of visible light (VIS), ultraviolet light (UV) and / or infrared light (IR). “Visible light” (VIS) generally refers to the band of electro- magnetic radiation with a wavelength from about 400 nm to about 750 nm. “Ultraviolet (UV) light” generally refers to electromagnetic radiation with a wavelength shorter than that of visible light, or from about 10 nm to about 400 nm range. “Infrared light” or infrared radiation (IR) generally refers to electromagnetic radiation with a wavelength greater than the VIS range, or from about 750 nm to about 50,000 nm. A light source may also provide full spectrum light. Light sources may output light from a selected wavelength or a range of wavelengths. In some embodiments of the invention, the light source may be configured to provide light above or below a predetermined wavelength, or may provide light within a predetermined range. A light source may be used in combination with a filter, to selectively transmit or block light of a selected wavelength from the light source. A light source may be connected to a intensity source by one or more electrical connectors; an array of light sources may be connected to a intensity source in series or in parallel. A intensity source may be a battery, or a vehicle electrical system or a building electrical system. The light source may be connected to a intensity source via control electronics (control circuit); control electronics may comprise one or more switches. The one or more switches may be automated, or controlled by a sensor, timer or other input, or may be controlled by a user, or a combination thereof. For example, a user may operate a switch to turn on a UV light source; the light source may be applied on a constant basis until it is turned off, or it may be pulsed (repeated on / off cycles) until it is turned off. In some embodiments, the light source may be switched from a continuously-on state to a pulsed state, or vice versa. In some embodiments, the light source may be configured to be brightening or darkening over time.
[0066] For operation, the light source may be connected to an intensity source capable of providing sufficient intensity to illuminate the sample. Control electronics may be used to switch the intensity on or off based on input from a user or some other input, and can also be used to modulate the intensity to a suitable level (e.g. to control brightness of the output light). Control electronics may be configured to turn the light source on and off as desired. Control electronics may include a switch for manual, automatic, or semi-automatic operation of the light sources. Theone or more switches may be, for example, a transistor, a relay or an electromechanical switch. In some embodiments, the control circuit may further comprise an AC-DC and / or a DC-DC converter for converting the voltage from the voltage source to an appropriate voltage for the light source. The control circuit may comprise a DC-DC regulator for regulation of the voltage. The control circuit may further comprise a timer and / or other circuitry elements for applying electric voltage to the optical filter for a fixed period of time following the receipt of input. A switch may be activated manually or automatically in response to predetermined conditions, or with a timer. For example, control electronics may process information such as user input, stored instructions, or the like.
[0067] One or more of a plurality of light sources may be provided. In some embodiments, each of the plurality of light sources may be the same. Alternatively, one or more of the light sources may vary. The light characteristics of the light emitted by the light sources may be the same or may vary. A plurality of light sources may or may not be independently controllable. One or more characteristic of the light source may or may not be controlled, including but not limited to whether the light source is on or off, brightness of light source, wavelength of light, intensity of light, angle of illumination, position of light source, or any combination thereof.
[0068] In some embodiments, light output from a light source may be from about 350 to about 750 nm, or any amount or range therebetween, for example from about 350 nm to about 360, 370, 380, 390, 400, 410, 420, 430 or about 450 nm, or any amount or range therebetween. In other embodiments, light from a light source may be from about 550 to about 700 nm, or any amount or range therebetween, for example from about 550 to about 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690 or about 700 nm, or any amount or range therebetween. In some embodiments, the wavelength of the light generated by the light source can vary, for example, ranging from 400 nm to 800 nm. In some embodiments, the wavelength of the light generated by the light source can be, or be about, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720, 730, 740, 750, 760, 770, 780, 790, 800 nm, or a number or a range between any two of these values. In some embodiments, the wavelength of the light generated by the light source can be at least, or at most, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720, 730, 740, 750, 760, 770, 780, 790, or 800 nm. The light source may be capable of emittingelectromagnetic waves in any spectrum. In some embodiments, the light source may have a wavelength falling between 10 nm and 100 ^m. In some embodiments, the wavelength of light may fall between 100 nm to 5000 nm, 300 nm to 1000 nm, or 400 nm to 800 nm. In some embodiments, the wavelength of light may be less than, and / or equal to 10 nm, 100 nm, 200 nm, 300 nm, 400 nm, 500 nm, 600 nm, 700 nm, 800 nm, 900 nm, 1000 nm, 1100 nm, 1200 nm, 1300 nm, 1500 nm, 1750 nm, 2000 nm, 2500 nm, 3000 nm, 4000 nm, or 5000 nm.
[0069] In one example, a light source may be a light-emitting diode (LED) (e.g., gallium arsenide (GaAs) LED, aluminum gallium arsenide (AlGaAs) LED, gallium arsenide phosphide (GaAsP) LED, aluminum gallium indium phosphide (AlGaInP) LED, gallium(III) phosphide (GaP) LED, indium gallium nitride (InGaN) / gallium(III) nitride (GaN) LED, or aluminum gallium phosphide (AlGaP) LED). In another example, a light source can be a laser, for example a vertical cavity surface emitting laser (VCSEL) or other suitable light emitter such as an Indium-Gallium-Aluminum-Phosphide (InGaAIP) laser, a Gallium-Arsenic Phosphide / Gallium Phosphide (GaAsP / GaP) laser, or a Gallium-Aluminum-Arsenide / Gallium-Aluminum-Arsenide (GaAIAs / GaAs) laser. Other examples of light sources may include but are not limited to electron stimulated light sources (e.g., Cathodoluminescence, Electron Stimulated Luminescence (ESL light bulbs), Cathode ray tube (CRT monitor), Nixie tube), incandescent light sources (e.g., Carbon button lamp, Conventional incandescent light bulbs, Halogen lamps, Globar, Nernst lamp), electroluminescent (EL) light sources (e.g., Light-emitting diodes—Organic light-emitting diodes, Polymer light-emitting diodes, Solid-state lighting, LED lamp, Electroluminescent sheets Electroluminescent wires), gas discharge light sources (e.g., Fluorescent lamps, Inductive lighting, Hollow cathode lamp, Neon and argon lamps, Plasma lamps, Xenon flash lamps), or high-intensity discharge light sources (e.g., Carbon arc lamps, Ceramic discharge metal halide lamps, Hydrargyrum medium-arc iodide lamps, Mercury-vapor lamps, Metal halide lamps, Sodium vapor lamps, Xenon arc lamps). Alternatively, a light source may be a bioluminescent, chemiluminescent, phosphorescent, or fluorescent light source.
[0070] Optical filters may be tuned in terms of clarity or haze, translucency, transparency or opacity, light transmittance (LT), switching speed, durability, photostability, contrast ratio, state of light transmittance (e.g. dark state or light state). “Light transmittance” (LT) refers to the quantity of light that is transmitted or passes through an optical filter, or device or apparatus comprising same. LT may be expressed with reference to a change in light transmissionand / or a particular type of light or wavelength of light (e.g. from about 10% visible light transmission (LT) to about 90% LT, or the like). LT may alternately be expressed as absorbance, and may optionally include reference to one or more wavelengths that are absorbed. According to some embodiments, an optical filter may be selected, or configured to have in one state, a LT of less than 80%, or less than 70%, or less than 60%, or less than 50%, or less than 40%, or less than 30%, or less than 20% or less than 10%, or any amount or range therebetween. According to some embodiments, an optical filter may be selected, or configured to have in another state, a LT of greater than 80%, or greater than 70%, or greater than 60%, or greater than 50%, or greater than 40%, or greater than 30%, or greater than 20% or greater than 10%, or any amount or range therebetween.
[0071] A filter can be a bandpass filter and can have peak transmittance of varying wavelength, ranging from 400 nm to 800 nm. In some embodiments, the peak transmittance can be, or be about, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720, 730, 740, 750, 760, 770, 780, 790, 800 nm, or a number or a range between any two of these values. In some embodiments, the peak transmittance can be at least, or at most, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720, 730, 740, 750, 760, 770, 780, 790, or 800 nm. The width of the transmission window of a filter can vary, for example, ranging from 1 nm to 50 nm. In some embodiments, the width of the filter can be, or be about, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50 nm, or a number or a range between any two of these values. In some embodiments, the width of the filter can be at least, or at most, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, or 50 nm. A shortpass filter may be considered a special bandpass filter having the lower limit of the transmission window close to 0 nm. A longpass filter may be considered a special bandpass filter having the upper limit of the transmission window close to infinity. A bandstop filter may be defined as complementary to some bandpass filter.
[0072] As used herein, an “optical channel” is a predefined profile of optical frequencies (or equivalently, wavelengths). For example, a first optical channel may have wavelengths of 500 nm–600 nm. To take an image in the first optical channel, one may use a detector which is only responsive to 500 nm–600 nm light, or use a bandpass filter having a transmission window of 500 nm–600 nm to filter the incoming light onto a detector responsive to300 nm–800 nm light. A second optical channel may have wavelengths of 300 nm–450 nm and 850 nm–900 nm. To take an image in the second optical channel, one may use a detector responsive to 300 nm–450 nm light and another detector responsive to 850 nm–900 nm light and then combine the detected signals of the two detectors. Alternatively, to take an image in the second optical channel, one may use a bandstop filter which rejects 451 nm–849 nm light in front of a detector responsive to 300 nm–900 nm light.
[0073] As used herein, an “1ex-2Ch” refers to one excitation, two channel systems and methods, where one excitation source may be used to excite fluorescent labels with emissions in two channels.
[0074] As used herein, an “2Ex-2Ch” refers to two excitation, two channel systems and methods, where two excitation sources may be used to excite fluorescent labels with emissions that may be detected in two channels. Example Sequencer
[0075] In FIG. 1A, an example sequencing system 100 which can perform the disclosed sequencing technology is illustrated. The sequencing system 100 can be configured to utilize disclosed sequencing methods based on a single optical excitation and a single detection channel. Non-limiting examples of the sequencing reactions utilized can include variations of sequencing-by-synthesis processes, such as those used in Illumina® dye sequencing or HeliScope® single molecule sequencing.
[0076] The sequencing system 100 can include an optics system 102 configured to generate raw sequencing data using sequencing reagents supplied by a fluidics system 104 that is part of the sequencing system 100. The raw sequencing data can include fluorescent images captured by the optics system 102. The sequencing system 100 can further include a computer system 106 that can be configured to control the optics system 102 and the fluidics system 104 via communication channels 108a and 108b. For example, a computer interface 110 of the optics system 102 can be configured to communicate with the computer system 106 through the communication channel 108a.
[0077] During sequencing reactions, the fluidics system 104 can direct the flow of reagents through one or more reagent tubes 112 to and from a flowcell 114 positioned on a mounting stage 116. The reagents can include, for example, fluorescently labeled nucleotides,buffers, enzymes, and cleavage reagents. The flowcell 114 can include at least one fluidic channel. The flowcell 114 can be a patterned array flowcell or a random array flowcell. The flowcell 114 can include multiple clusters of single-stranded polynucleotides to be sequenced in the at least one fluidic channel. The lengths of the polynucleotides can vary ranging, for example, from about 50 bases, 100 bases, 150 bases, 200 bases, 300 bases, 500 bases, to about 1000 bases. The polynucleotides can be attached to one or more fluidic channels of the flowcell 114. In some embodiments, the flowcell 114 can include a plurality of wells, wherein each well can include a cluster comprising multiple identical copies of a target polynucleotide to be sequenced. The mounting stage 116 can be configured to allow proper alignment and movement of the flowcell 114 in relation to the other components of the optics system 102. In one embodiment, the mounting stage 116 can be used to align the flowcell 114 with a lens 118.
[0078] The optics system 102 can include a two light sources 120, such as lasers or a LEDs, with each light source configured to generate light having wavelengths distributed at around a predetermined wavelength. However, embodiments are not limited to any particular wavelength of light. The light source only needs to be configured to generate the correct wavelength of light which excites the fluorescent labels attached to the nucleotides on the flowcell.
[0079] The light generated by the light source 120 can pass through fiber optic cables 122 to excite fluorescent labels in the flowcell 114. The lens 118, mounted on a focuser 124, can move along the z-axis. The focused fluorescent emissions can be detected by detectors 126, for example charge-coupled device (CCD) sensors or a complementary metal oxide semiconductor (CMOS) sensors. In some embodiments, nucleotide incorporations can be detected with zeromode waveguides as described, for example, in Levene et al. Science 299, 682-686 (2003); Lundquist et al. Opt. Lett.33, 1026-1028 (2008); and Korlach et al. Proc. Natl. Acad. Sci. USA 105, 1176-1181 (2008), the disclosures of which are incorporated herein by reference in their entireties.
[0080] A filter assembly 128 of the optics system 102 can be configured to filter the fluorescent emissions from the fluorescent labels in the flowcell 114. The filter assembly 128 can include a plurality of optical filters, where a correct filter can be selected depending on the particular fluorophores used in a sequencing reaction. In one alternate embodiment, the computer system 106 may automatically determine which optical filter should be used for a sequencing reaction, e.g., by scanning labels and / or barcodes attached to a sample vial and determining the particular fluorophores to be used in a sequencing reaction based on the labels and / or barcodes, orby retrieving information stored in the memory relating to previous sequencing reactions, and then control the filter assembly 128 to select and use the desired optical filter. The selected filter can be a longpass filter, a shortpass filter, a bandstop filter, or a bandpass filter, depending on the types of fluorescent molecules being used in the system. For example, the selected filter can be a bandpass filter selected to match the peak of the emission spectrum of a particular fluorescent label.
[0081] In some embodiments, the detectors 126 include one or more sub-detector while the filters of the filter assembly 128 may be mechanically switched or rotated in front of the sub- detector, such that differently filtered images can be taken by the sub-detector sequentially. In some embodiments, the detectors 126 include one sub-detector and the filter assembly 128 may include at least one layer of switchable material which has a light transmittance that is variable upon application of a stimulus, where the stimulus may be light, electricity, temperature, or any combination thereof. As a result, the filter assembly 128 can provide a plurality of optical filters such that differently filtered images can be taken by the sub-detector sequentially. In some embodiments, the detectors 126 may each include one sub-detector and the filter assembly 128 may include one or more switchable filters base on the micro-electromechanical system technology, such that differently filtered images can be taken by the sub-detector sequentially.
[0082] In some embodiments, the detectors 126 can include two or more sub-detectors to be selected depending on the set of fluorophores used, for example a first detector coupled with a first filter and a second detector coupled with a second filter. In some embodiments, the optics system 102 may include two or more dichroic mirrors / beamsplitters configured to split the fluorescent emissions, such that after splitting the fluorescent emissions with the dichroic mirrors, the detectors 126 can take two differently filtered images simultaneously (or close in time) using the two sub-detectors coupled with two different filters. In some embodiments, the detectors 126 can include two or more sub-detectors stacked along the incoming direction of the fluorescent emissions. Different wavelengths of the fluorescent emissions may differentially decay or be differentially absorbed along the incoming direction, such that sub-detectors at different positions along the incoming direction can be selected depending on the set of fluorophores used, or be configured to take differently filtered images simultaneously (or close in time).
[0083] In use, a sample having a polynucleotide to be sequenced may be loaded into the flowcell 114 and placed in the mounting stage 116. The computer system 106 may then activatethe fluidics system 104 to begin a sequencing cycle. During sequencing reactions, the computer system 106 may instruct the fluidics system 104, through the communication interface 108b, to supply reagents, for example labeled nucleotide analogs, to the flowcell 114. Through the communication interface 108a and the computer interface 110, the computer system 106 may control the light source 120 of the optics system 102 to generate light at around a predetermined wavelength and excite nucleotide analogs incorporated into growing primers hybridized to the polynucleotide being sequenced, for example. The computer system 106 may control the detector 126 of the optics system 102 to capture images of the diffraction-limited spots of DNA clusters having the fluorescently labeled nucleotide analogs. The computer system 106 can receive the fluorescent images from the detector 126 and process the fluorescent images received to determine the nucleotide sequence of the polynucleotide being sequenced.
[0084] In FIG.1B, an example of an imaging system 10000 to be used in the disclosed sequencing technology is illustrated. For example, the imaging system 10000 may be used in the example sequencing system 100 illustrated in FIG. 1A. The imaging system 10000 may include a light source 11000 that can provide light at one or more wavelengths to excite fluorophores at targeted points on a sample. The light source 11000 can include one or more lasers, light-emitting diodes, or other optical sources, such that the light source 11000 can provide a variety of wavelengths of light. In some embodiments, the light source 11000 can be configured to selectively provide light with a predetermined range of wavelengths that are tuned to the set of fluorophores being used. In some embodiments, the light source 11000 can be configured to output light at an optical frequency corresponding to a wavelength in a predefined range of wavelengths of light. In some embodiments, a user of the disclosed sequencing systems may choose a specific optical frequency to be output from the light source 11000, depending on the particular fluorophores used in a sequencing reaction. In one alternate embodiment, the computer system 106 may automatically determine which optical frequency should be output from the light source 11000, e.g., by scanning labels and / or barcodes attached to a sample vial and determining the particular fluorophores to be used in a sequencing reaction based on the labels and / or barcodes, or by retrieving information stored in the memory relating to previous sequencing reactions, and then control the light source 11000 to select and output the desired optical frequency.
[0085] The imaging system 10000 may include an optical path 12000 from the light source 11000 to the sample 13000, e.g., a microfluidic device including one or more flow chamberswhere one or more sequencing reactions occur. In some embodiments, the optical path 12000 can include a combination of one or more of mirrors, lenses, prisms, quarter wave plates, half wave plates, polarizers, filters, dichroic mirrors, beam splitters, beam combiners, objective lenses, wide field optics configured to spread light from a light source over a relatively large region of a sample, etc. The optical path 12000 can be configured to direct light from the light source 11000 to the sample 13000. In addition, the optical path 12000 may include optical components which can be configured to direct light emitted from the sample 13000 to an integration detection system 15000. In some embodiments, a portion of the optical elements that are used to direct light from the light source 11000 to the sample 13000 are also used to direct light from the sample 13000 to the integration detection system 15000. Further examples of optical paths and optical systems may be found in U.S. Pat. No. 7,589,315, U.S. Pat. No. 8,951,781, or U.S. Pat. No. 9,193,996, each of which is incorporated by reference herein in its entirety.
[0086] The imaging system 10000 may include a scanning system 14000 to effectively move light relative to the sample 13000 to scan the sample to generate an image. In some embodiments, the scanning system 14000 can be implemented within the optical path 12000. For example, the scanning system 14000 can include one or more scanning mirrors that move relative to one another within the optical path 12000 to effectively move the light from the light source 11000 across the sample. In some embodiments, the scanning system 14000 can be implemented as a mechanical system that physically moves the sample 13000 so that the sample moves relative to the light from the light source 11000. In some embodiment, the scanning system 14000 can be a combination of optical components in the optical path 12000 and a mechanical system for physically moving the sample 13000 so that the light from the light source 11000 and the sample 13000 move relative to one another.
[0087] The imaging system 10000 may include an integration detection system 15000 that includes one or more light detectors as well as associated electronic circuitry, processors, data storage, memory, and the like to acquire and process image data of the sample 13000. In some embodiments, the integration detection system 15000 can include photomultiplier tubes, avalanche photodiodes, image sensors (e.g., CCDs, CMOS sensors, etc.), and the like. In some embodiments, the light detectors of the integration detection system 15000 can include components to amplify light signals and may be sensitive to single photons. In some embodiments, the light detectors of the integration detection system 15000 can have a plurality of channels or pixels. The integrationdetection system 15000 can acquire one or more images based on the light detected from the sample 13000.
[0088] In some embodiments, the optical path 12000 may include an array generator 12100 that can generate a plurality of light exposure regions on the sample 13000. In some embodiments, the array generator 12100 can generate a certain light exposure pattern on the sample 13000. These light exposure regions can be scanned over the sample 13000 using the scanning system 14000 to selectively illuminate areas of the sample 13000 for imaging. The integration detection system 15000 can integrate signals corresponding to particular points on the sample 13000 as the plurality of light exposure regions are scanned over the sample 13000. For example, for an individual point on the sample 13000, the integration detection system 1500 can selectively aggregate detected signals corresponding to the individual point where the individual point is illuminated at different times by different light exposure regions. In some embodiments, the combination of the array generator 12100 and the integration detection system 15000 can detect light simultaneously, or near-simultaneously, from a plurality of points on the sample 13000. In some embodiments, the combination of the array generator 12100 and the integration detection system 15000 can integrate the detected light from a plurality of points on the sample over time.
[0089] In some embodiments, a plurality of sequencing reactions may be run parallelly in a plurality of flow chambers of the sample 13000. For example, a plurality of sequencing reactions may be performed for a plurality of biological specimen. In some embodiments, the plurality of sequencing reactions may use different sets of fluorophores. In some embodiments, the light source 11000, the array generator 12100, and the scanning system 14000 can be configured to selectively illuminate different areas of the sample 13000 with different optical frequencies of light, depending on the different sets of fluorophores used for the sequencing reactions occurring in different areas of the sample 13000. Methods of Sequencing Clusters of Labeled Polynucleotides
[0090] Fig.2 shows a flowchart of a method 100 of sequencing clusters of labeled polynucleotides bound to a flowcell according to one embodiment. The method may begin at a start step 102, by, for example, gathering DNA samples, and putting the sample into a sequencing system. Sequencing systems according to the disclosure may include flow cell sequencers such as those produced by ILLUMINA®, INC. (San Diego, CA). For example, when the method 100begins at step 102, a flowcell may have been prepared by embedding the flowcell with fragmented polynucleotides (e.g., fragmented single- or double-stranded polynucleotide fragments). Fragmented polynucleotides may be generated from a deoxyribonucleic acid (DNA) sample. DNA samples may be from various sources, for example, a biological sample, a cell sample, an environmental sample, or any combination thereof. The lengths of fragmented polynucleotide fragments may range from, for example, 100 bases to 1000 bases.
[0091] Polynucleotide fragments may be bridge-amplified into clusters of polynucleotide fragments attached to the inside surface of one or more channels of a flowcell. An inside surface of the one or more flowcell channels may include two types of primers, for example a first primer type (P1) and a second primer type (P2) and the DNA fragments may be amplified by well-known methods.
[0092] After generating clusters within the flowcell, the method 100 may begin a Sequencing by Synthesis process. During each sequencing cycle, four types of nucleotide analogs may be added and incorporated onto the growing primer-polynucleotides. Each type of nucleotide analogs may be linked to one or more fluorescent label. For example, a first type of nucleotide may be an analog of deoxyguanosine triphosphate (dGTP), and a mixture of dGTP / fluorescent conjugates may be prepared where dGTP is conjugated with two types of fluorescent labels.
[0093] In some embodiments, a third of the dGTP may be conjugated via a linker with a first type of fluorescent label that, after an excitation, can emit at a first emission wavelength, and the other two-thirds of the dGTP may be conjugated via a linker that is dark and does not emit light after an excitation. A second type of nucleotide may be an analog of deoxythymidine triphosphate (dTTP), and a mixture of dTTP / fluorescent conjugates may be prepared where two- thirds of the dTTP is conjugated via a linker with the first type of fluorescent label, and the remaining third conjugated via a linker with a second type of fluorescent label that, after an excitation, can emit at a second emission wavelength. A third type of nucleotide may be an analog of deoxycytidine triphosphate (dCTP), and a mixture of dCTP / fluorescent conjugates may be prepared where one-third of the dCTP is conjugated via a linker with the first type of fluorescent label, and the remaining two-thirds are conjugated via a linker with the second type of fluorescent label. A fourth type of nucleotide may be an analog of deoxyadenosine triphosphate (dATP), and a mixture of dATP / fluorescent conjugates may be prepared where a third of the dATP may beconjugated via a linker with the second type of fluorescent label, and the other two-thirds of the dGTP may be conjugated via a linker that is dark and does not emit light after an excitation.
[0094] The linkers may include one or more cleavage groups. Prior to the subsequent sequencing cycle, the fluorescent labels may be removed from the nucleotide analogs. For example, a linker attaching a fluorescent label to a nucleotide analog may include an azide and / or an alkoxy group, for example on the same carbon, such that the linker may be cleaved after each incorporation cycle by a phosphine reagent, thereby releasing the fluorescent label from subsequent sequencing cycles.
[0095] Once the clusters are created on the flowcell, the method 100 moves to a step 110 of detecting fluorescent emissions from a first labeled nucleotide at a first wavelength and a first intensity. For example, step 110 may include exciting all of the clusters of labeled polynucleotides on a flowcell at a first excitation wavelength. A single light source such as a laser or an LED source may excite a fluorescent label at the predetermined excitation wavelength. In some embodiments, the single laser or the LED source may be non-tunable. As described in Fig. 2 here, “wavelength” refers to a detection wavelength, or range of detection wavelengths. Accordingly, Step 110 may also include detecting any fluorescent emissions from the clusters within a first detection wavelength range to detect the presence of a first labeled nucleotide. In general, the first detection wavelength will be red-shifted to a longer wavelength relative to the first excitation wavelength. For reference, step 110 of detecting fluorescent emissions from a first labeled nucleotide at a first wavelength and a first intensity, encompasses either example of dATP and dGTP above, wherein a mixture of nucleotide analogs / fluorescent conjugates is prepared with a fluorescent and non-fluorescent linkers.
[0096] Image processing yields base calling, described in further detail below, where the complementary nucleotides added to the molecules in a cluster during a cycle are identified. In some embodiments, the fluorescent images may be stored for later processing offline. In some embodiments, the fluorescent images may be processed to determine the sequence of the growing primer-polynucleotides in each cluster in real time. After detecting the fluorescence emissions from the clusters at step 110, the method 100 moves to a step 120, where the method detects fluorescent emissions from a second labeled nucleotide at a second wavelength and a second intensity, wherein the first wavelength is different from the second wavelength.
[0097] When the method detects fluorescent emissions from a first wavelength in step 110 different than a second wavelength in step 120, the method may be described as employing at least two channel base calling. A “channel” may be used to describe an emission / detection process for a labeled or unlabeled base that is specific for a particular excitation wavelength and detection wavelength pair. Thus, a channel may include a particular first range of wavelengths of light used to excite a fluorescent dye and a second range of wavelengths of light used to detect the fluorescent emissions from the excited dyes. The disclosure provides for, inter alia, two-channel base calling with one excitation and two excitation sources. The disclosure provides for, inter alia, two-channel base calling with four labeled nucleotides where each nucleotide is labeled with at least two different fluorescent labels. For reference, step 110 of detecting fluorescent emissions from a second labeled nucleotide at a second wavelength and a second intensity, wherein the first wavelength is different from the second wavelength. encompasses, alternately, detecting fluorescent emissions from dATP as the second nucleotide analog, if dGTP is the first nucleotide analog above, or vice versa.
[0098] After detecting the fluorescence emissions from the clusters at step 120, the method 100 moves to a step 130, where the method detects fluorescent emissions from a third labeled nucleotide at the first and second wavelengths at third and fourth intensities. For example, at step 130 fluorescent emissions from a third labeled nucleotide may be detected at the first wavelength at a third intensity that is greater than the first intensity. Also, at step 130 fluorescent emissions from a third labeled nucleotide may be detected at the second wavelength at a fourth intensity that is about the same as the second intensity. While not required, in general, the third intensity and the fourth intensity are not equal. In some embodiments, the third intensity may be at least one of 200%, 150%, 110%, 100%, 90%, 75%, 50%, 25% and 10% greater than the fourth intensity, or vice versa where the third intensity is less than the fourth intensity. In some embodiments, fluorescent emissions from a third labeled nucleotide at the first and second wavelength may be detected at some non-zero intensity other than the first intensity.
[0099] After detecting the fluorescence emissions from the clusters at step 110, the method 100 moves to a step 140, where the method detects fluorescent emissions from a fourth labeled nucleotide at the first and second wavelengths at fifth and sixth intensities. For example, at step 140 fluorescent emissions from a third labeled nucleotide may be detected at the first wavelength at a fifth intensity that is about the same as the first intensity. Also at step 140fluorescent emissions from a third labeled nucleotide may be detected at the second wavelength at a sixth intensity that is greater than the second intensity.
[0100] The intensity of the detected emissions at the first wavelength may be quantified as an absolute measurement in terms of photon flux or counts per second. As described herein, the first and second intensities may also be quantified as relative measurement for a cluster labeled with a single nucleotide as compared to a cluster labeled with two different nucleotides. While not required, in general the fifth intensity and the sixth intensity are not equal. In some embodiments, the ratio of third to fourth intensities is inverted compared to the ratio of the fifth and sixth intensities. For example, if the ratio of intensity of the third to fourth intensities, for wavelength 1 and 2 respectively, is 1:2, then the ratio of the fifth to sixth intensities may be 2:1.
[0101] In some embodiments, the disclosure provides for systems and methods for DNA sequencing using detection schemes other than fluorescence. For example, some DNA sequencing systems and methods employ voltage detectors instead of light detectors to detect nucleotides. In some embodiments, these techniques may or may not require amplification of the DNA or RNA sample into a cluster, for example, to be analyzed. In some embodiments, these techniques may or may not require the labelling of the DNA or RNA sample in order to be analyzed. The methods disclosed herein may be applied to voltage systems via an analogous process from the fluorescent detection systems. In some embodiments, instead of detecting fluorescent emissions from a labelled nucleotide and determining the nucleotide based on the emission wavelength and intensity, a voltage detection system may detect various parameters of a voltage signal, such as amplitude, frequency, or waveform characteristics, to extract relevant information to determine a nucleotide. In some embodiments, various parameters of a voltage signal, such as amplitude, frequency, or waveform characteristics may be characterized as a channel, with “on” and “off” states. For example, a voltage signal corresponding to a “T” nucleotide may be detected as a voltage signal with a characteristic frequency and at a first amplitude. In some embodiments, a voltage signal corresponding to a “C” nucleotide may be detected at the same characteristic frequency but at a second amplitude.
[0102] The method 100 then moves to a decision step 145 to determine if the method 100 should repeat for an additional cycle of reading sequence data. A determination may be made at decision step 145 whether to detect more nucleotides based on, for example, the quality of thesignal or after a predetermined number of bases. If more nucleotides are to be detected, then the method 100 may loop back to the step 110 to start a next sequencing cycle.
[0103] The method 100 may then move to a step 150 wherein the system may determine the nucleotide sequence of the polynucleotide based on the detected fluorescent emissions and intensities. For example, at step 150 the method may identify clusters that have added nucleotides with emissions at wavelengths and intensities corresponding to the first, second, third and / or fourth labeled nucleotides. In some embodiments, at step 150 the method may identify clusters which had no fluorescent emissions following excitation at the first and / or second excitation wavelength and determine that the cluster corresponds to an empty well or to a short insert that has completed sequencing.
[0104] As mentioned above, a cluster may be identified by the lack of any fluorescence after excitation at either the first or second wavelengths. Because the lack of fluorescent emissions is dark, undetectable, or otherwise “missing” on an image, the system may track the position of each cluster on a flow cell. Here, the term “undetectable” refers to fluorescent emissions from a nucleotide that are intentionally or unintentionally reduced to an intensity that is not effectively distinguishable from noise by a detection scheme. Once a cluster has been identified at a particular position, the system may then note in subsequence sequencing rounds whether there is a fluorescent emission at the position of the cluster. If no fluorescent emissions are noted at a known position of a cluster, the system may determine that no nucleotide was sequenced—from either an empty well or a finished cluster.
[0105] In some embodiments, prior to the next sequencing cycle, the fluorescent labels attached to each nucleotide may be removed from the incorporated nucleotide analogs, and the reversible 3ƍ blocks may be removed so that another nucleotide analog may be added onto each extending primer-polynucleotide. If a determination is made at the decision step 145 that there are no more additional rounds of sequencing necessary, the method 100 then moves to an end step 160 and terminates the method 100. If the system is employing offline fluorescent imaging processing, when there is no additional nucleotide to be detected at decision block 145, the fluorescent images comprising the fluorescent signals detected may be processed after step 160, and the bases of the nucleotides incorporated into the fragmented polynucleotide may be determined remotely. For each nucleotide base determined, a quality score may be determined. After all the fluorescent images are processed, the method 100 may terminate at the step 160.
[0106] An aspect of the disclosure is directed to providing a sequencing system with fluorescent emissions with different relative intensities as compared to prior systems which may have only utilized an “on” versus “off” intensity. An increase in brightness may be quantified as an absolute measurement in terms of photon flux or counts per second. The increase in brightness may also be quantified as relative measurement for a cluster labeled with a single nucleotide as compared to a cluster labeled with two different nucleotides. For SNR optimization reasons, some sequencing systems employ a square constraint on a scatterplot of a two-channel detection, wherein a cluster labeled with a single nucleotide is constrained to have a similar intensity as a cluster labeled with two nucleotides (e.g. base C in a corner cloud). Such constraints in intensity / brightness of a cluster may be controlled by diluting labeled nucleotides with a non- fluorescing tag, such that all clusters are emitting at approximately half of the potential brightness. Relaxing this constraint according to this disclosure may improve increase signal to noise ratio (SNR).
[0107] It should be realized that on any flowcell, the different clusters may have varying brightness. For example, some clusters can be bright, and some clusters can be dim in comparison to each other. In embodiments, the intensity values vary between base calling cycles and thus the classification of bright and dim may also change between cycles. Some examples of intensity value ratios of emissions between bright and dim clusters include 0.55:0.45, 0.60:0.40, 0.65:0.35, 0.70:0.30, 0.75:0.25, 0.80:0.20, 0.85:0.15, 0.90:0.10, and 0.95:0.05. During each sampling event (e.g., each illumination stage or each image acquisition stage), a detector may image clusters with different intensities or clusters generating different types of signals.
[0108] In some embodiments, a method according to this disclosure may allow for each cluster to be labeled with a single dye, and that dye, or faction of that dye, may be chosen to increase the brightness of the cluster. In some embodiments, methods and systems may not require clusters to be labeled with two nucleotides. Some embodiments may result in an increase in the brightness of a cluster relative to an intensity of fluorescent emissions from a labeled nucleotide labeled with equal proportions of two fluorescent dyes, wherein the two fluorescent dyes are detected in two different channels. Some embodiments may result in an increase in the brightness of a cluster relative to an intensity of fluorescent emissions from a nucleotide forming a corner cloud in a two-excitation, two-channel detection method.
[0109] The brightness of any cluster may also be affected by fragment length distribution of the sample. The varying brightness of the cluster population can have the effect of elongating the ‘on’ populations in the base calling scatterplot. By way of example, without normalizing the level of amplification before trying to increase the brightness of all the clusters, some over-amplified AT rich sequences may be much brighter than similar GC rich sequences, and become even more ‘over amplified’ than the GC rich clusters. In some embodiments, it may be advantageous to normalize each cluster’s intensity by its mean intensity in the first 10 cycles to reduce population intensity variation. For example, in the first ten cycles, for every non- guanine(G) base call, two radii can be calculated: the distance of the population intensity from the origin, and the distance of the corresponding Gaussian mean from the origin. Cluster scaling can include normalizing to the mean of the ratio of these two radii averaged over, for example, the first 10 cycles. All cluster intensities can be normalized by this scaling factor before phase correction and base calling are performed. Cluster scaling can advantageously increase throughput and decrease error rates, for example, for samples with large fragment length distributions.
[0110] It is also possible to improve the brightness of all clusters, for example, by carrying out a higher number of amplification cycles, or by changing the chemistry / fraction of the dyes, or by changing the detection scheme (two to three-channel detection as described herein). In some embodiments, the methods and systems may be employed in a two-channel system, however, one of skill in the art will understand that the labeling in one channel with different / same dyes for two nucleotides may be applied to, for example, a four-channel system.
[0111] In some embodiments, the disclosure provides for systems and methods for DNA sequencing using detection schemes other than fluorescence. For example, some DNA sequencing systems and methods employ voltage detectors instead of light detectors to detect specific nucleotides within a polynucleotide. In some embodiments, these techniques may or may not require amplification of the DNA or RNA sample into a cluster, for example, to be analyzed. In some embodiments, these techniques may or may not require the labelling of the DNA or RNA sample in order to be analyzed. The methods disclosed herein may be applied to voltage-based systems via an analogous process to the fluorescent detection systems. In some embodiments, instead of detecting fluorescent emissions from a labelled nucleotide and determining the nucleotide based on the emission wavelength and intensity, a voltage detection system may detect various parameters of a voltage signal, such as amplitude, frequency, or waveform characteristics,to extract relevant information to determine the presence of a particular nucleotide. In some embodiments, various parameters of a voltage signal, such as amplitude, frequency, or waveform characteristics may be characterized as a channel, with “on” and “off” states. For example, a voltage signal corresponding to a “T” nucleotide may be detected as a voltage signal with a characteristic frequency and at a first amplitude. In some embodiments, a voltage signal corresponding to a “C” nucleotide may be detected at the same characteristic frequency but at a second amplitude. Diamond Configuration In One Excitation, Two Channel Systems And Methods
[0112] An aspect of the disclosure is directed to alternative scatterplots that may be used with 1Ex-2Ch systems and methods. A diamond configuration in a 2D scatterplot refers to the arranging of clouds of data points in a diamond-shaped pattern. This pattern may be made by at least two pairs of approximately symmetrically positioned clouds with regard to the y=x plane on the x and y axes, resulting in a diamond shape. This diamond configuration can be utilized in fluorescence-based assays, where two channels are used to detect different fluorescent signals. The intensity of one fluorescent signal is represented by the x-axis of the scatterplot, while the intensity of the second fluorescent signal is represented by the y-axis. Each scatterplot data point may represent a single fluorescent measurement and resulting base call.
[0113] In a diamond configuration, the points located at the top and bottom corners of the diamond have a relatively high intensity in one channel and a relatively low intensity in the second channel. Conversely, the points located at the upper right and lower left corners of the diamond have high intensity in the second channel and low intensity in the first channel. Accordingly, the diamond configuration may be used to identify and distinguish different populations based on their fluorescence characteristics. In a polynucleotide sequencing experiment, for example, the diamond configuration can be used to discriminate between different types of bases based on their fluorescent signals (e.g., A, C, G, T).
[0114] The use of this alternative scatterplot shape has the benefits of removing the need for a 'dark' (unlabeled) base, and in the one excitation, two channel implementation allows one ffN per base / cloud. Using one ffN per base / cloud permits the use of bright dyes with crosstalk and the use of less bright dyes with reduced crosstalk offers several benefits over traditional scatterplot configurations. Having the corner cloud represented by two ffNs can increase sequencecontext issues, increase noise and reduce quality, and thus the context specific variation in the dual clouds is reduced. Also, by eliminating the need for an unlabeled base, this configuration simplifies the experimental process and reduces the potential for incorrect base calls for dark base homopolymers, where a series of G bases are overcalled with high confidence.
[0115] Fig. 3A and Fig. 3B show two-dimensional scatterplots with 1Ex-2Ch chemistry, where using alternate dyes, or ratios of existing dyes, may be used to obtain a diamond- shaped scatterplot. Note that while shown to illustrate an 1Ex-2Ch system, the diamond scatterplot shape may also be used for 2Ex-2Ch systems as described herein with differences in the dyes used. The Y axes 201 corresponds to the intensity of fluorescent emissions at a first wavelength, which in this example is a green channel. The X axes 205 corresponds to the intensity at a second wavelength, which in this example is a blue channel. Fig. 3A shows a scatterplot containing experimental data from a typical 1Ex-2Ch system. Nucleotide “A” 210 is shown as a cloud in the top left corner, with nucleotide “C” 220 as a cloud in the top right, nucleotide “T” 230 as a cloud in the lower right, and nucleotide “G” 240 as a cloud at the origin of the scatterplot in the lower left corner. The respective location of each nucleotide may be changed, with nucleotide “G” switching position with any other nucleotide by switching the fluorophore associated with nucleotide “A” to be conjugated with nucleotide “G,” and vice versa. Fig.3A also includes arrows indicating that the centers of each cloud may move on the chart by using different dyes or mixtures of dye percentages or concentrations to change the amount of emission intensity that is detected in either channel 1 or channel 2 for each cluster of nucleotides.
[0116] Fig. 3B shows an illustration of a two-dimensional scatterplot with 1Ex-2Ch chemistry in a diamond configuration, but relocating the center of each cloud as suggested in Fig. 3A to form the diamond shape. Note that, as an example that the respective location of each nucleotide may be changed, Fig.3B shows nucleotide “T” 212 in the top center location, whereas in Fig. 3A the nucleotide “T” 230 was in the lower right corner. Accordingly, nucleotide “T” 212 in Fig.3B has been switched relative to nucleotide “T” 230 in Fig.3A but also relocated –by using a different dye or mixture of dyes. Nucleotide “T” 212 is shown as a circle shaded with one-third blue and two-thirds green, which illustrates a potential ratio of green and blue dyes, or a single dye with blue and green emissions, that will locate nucleotide “T” 212 in the top center location, i.e. at a maximum normalized intensity of green and a partial intensity of blue.
[0117] Nucleotide “C” 222 in Fig.3B was not switched relative to nucleotide “C” 220 in Fig. 3A, but was relocated to the middle right location in the scatterplot–by using a different dye or mixture of dyes. Nucleotide “C” 222 is shown as a circle shaded with one-third green and two-thirds blue, which illustrates a potential ratio of green and blue dyes, or a single dye with blue and green emissions, that will locate nucleotide “C” 222 in the right center location, i.e. at a maximum normalized intensity of blue and a partial intensity of green.
[0118] Nucleotide “A” 232 in Fig.3B took the place of nucleotide “T” 230 in Fig.3A, and was also relocated –by using a different dye or mixture of dyes. Nucleotide “A” 232 is shown as a circle shaded with one-third blue and two-thirds black, which illustrates a potential ratio of blue dyes and non-fluorescent labels, or a single dye that will locate nucleotide “A” 232 in the bottom center location, i.e. at a nonzero normalized intensity of blue and zero green intensity.
[0119] Nucleotide “G” 242 in Fig.3B was not switched relative to nucleotide “G” 240 in Fig.3A, but was relocated to the middle left location in the scatterplot–by using a different dye or mixture of dyes. Nucleotide “G” 242 is shown as a circle shaded with one-third green and two- thirds black, which illustrates a potential ratio of green dyes and non-fluorescent labels, or a single dye that will locate nucleotide “G” 242 in the left center location, i.e., at a nonzero normalized intensity of green and zero blue intensity.
[0120] Fig. 3A and 3B demonstrate an embodiment encompassed by a method of sequencing polynucleotides bound to a flowcell, comprising: detecting fluorescent emissions from a first labeled nucleotide at a first wavelength and a first intensity; detecting fluorescent emissions from a second labeled nucleotide at a second wavelength and a second intensity, wherein the first wavelength is different from the second wavelength; detecting fluorescent emissions from a third labeled nucleotide at the first and second wavelengths at third and fourth intensities; detecting fluorescent emissions from a fourth labeled nucleotide at the first and second wavelengths at fifth and sixth intensities; and determining the sequence of the polynucleotides based on the detected fluorescent emissions and intensities. Specifically, nucleotide “G” 242 in Fig.3B may correspond to a first labeled nucleotide with emissions at a first wavelength and a first intensity, where the first wavelength is an emission wavelength within a green channel and the first intensity is approximately 33% of a maximum intensity in the green channel (with arbitrary units). A first labeled nucleotide that corresponds nucleotide “G” 242 in Fig. 3B may be labeled with a single dye, a single partially labeled dye that is diluted with a non-fluorescent label, and two dyes.
[0121] Nucleotide “A” 232 in Fig. 3B may correspond to a second labeled nucleotide with emissions at a second wavelength and a second intensity, wherein the first wavelength is different from the second wavelength, where the second wavelength is within a blue channel, and the second intensity is approximately 33% of a maximum intensity in the blue channel (with arbitrary units). The emissions of Nucleotide “A” 232 illustrated in Fig.3B may be replicated with any of a single dye, a single partially labeled dye that is diluted with a non-fluorescent label, and a combination of two dyes.
[0122] Nucleotide “T” 212 in Fig. 3B may correspond to a third labeled nucleotide with emissions at the first and second wavelengths at third and fourth intensities, wherein the third intensity is approximately 66% of a maximum intensity in the green channel, and the fourth intensity is approximately 33% of a maximum intensity in the blue channel. Nucleotide “C” 222 in Fig. 3B may correspond to a fourth labeled nucleotide at the first and second wavelengths at fifth and sixth intensities, wherein the fifth intensity is approximately 33% of a maximum intensity in the green channel, and the sixth intensity is approximately 66% of a maximum intensity in the blue channel. The emissions of Nucleotide “C” or “T” illustrated in Fig.3B may be each replicated with any of a single dye, a single partially labeled dye that is diluted with a non-fluorescent label, and a combination of two dyes.
[0123] Fig.4 is a line graph that shows a model of a simulated emission spectra leading to the diamond configuration shown in Fig. 3B. The Y axis 301 corresponds to the intensity of fluorescent emissions (in arbitrary units) and the X axis 305 corresponds to wavelength (in nanometers), across an emission band corresponding to a blue channel 380, and a green channel 390. The emission spectrum associated with nucleotide “A” 310 falls almost entirely within the blue channel 380. The emission spectrum associated with nucleotide “G” 340 falls almost entirely within the green channel 390. The emission spectra associated with nucleotide “C” 320 and nucleotide “T” 330 show emission in both the blue channel 380 and green channel 390. The integrated emissions of nucleotide “C” 320 in the green channel 390 are shown to approximately match the integrated emissions of nucleotide “G” 340. Similarly, the integrated emissions of nucleotide “T” 330 in the green channel 390 are shown to approximately match the integrated emissions of nucleotide “A” 310.
[0124] These emission spectra are compatible with 1Ex-2Ch chemistry, by, for example, using single dyes with progressively larger stokes shift. For example, the simulatedemission spectrum of nucleotide “A” 310 may be replicated in an experiment by a single fluorescent dye that absorbs a higher energy wavelength excitation and emits the light with a medium quantum yield leading to a partial intensity of fluorescent emissions in the blue channel 380. The simulated emission spectrum of nucleotide “C” 320 may be replicated by a single fluorescent dye with a higher quantum yield than the dye used for nucleotide “A” 310, leading to a larger intensity of fluorescent emissions in the blue channel 380. Systems and methods according to the disclosure provide for a single ffN per base / cloud, which can lead to increasing signal and reducing noise.
[0125] In some embodiments, each of the green and blue channels may be separately normalized, such that the quantum yields of the dyes emitting in the blue channel do not need to precisely match the quantum yields of dyes in the green channel. The emission spectra corresponding to each of the nucleotides may also be reduced in intensity by diluting the percentage of bright dyes linked to a nucleotide by the addition of non-fluorescent labels.
[0126] In some embodiments, the emission spectra of Figs.3A and 3B are compatible with 1Ex-2Ch chemistry by using mixtures of dyes. The emission corresponding to nucleotide “A” 310 may be created by the mixture of a bright blue dye diluted with a non-fluorescent label. The emissions corresponding to nucleotide “G” 340 may be created by the mixture of a bright green dye diluted with a non-fluorescent label. As currently shown, the emission spectra of A and G are sufficiently separated such that a mixture of the two dyes used to identify those nucleotides may be unlikely to reproduce the spectra corresponding to either C or T. Instead, a mixture of the dye used for nucleotide A and another third dye with an emission centered between the wavelengths of the blue and green channels may contribute to an additive emission spectrum for nucleotide C 320.
[0127] Similarly, a mixture of the dye used for nucleotide “G” and the same third dye as above with an emission centered between the wavelengths of the blue and green channels may result in an additive emission spectrum for nucleotide “C” 320. Alternatively, a fourth dye with an emission centered between the wavelengths of the blue and green channels may contribute to an additive emission spectrum for nucleotide “C” 320. In some embodiments, a mixture of a third dye and a fourth dye each with emissions centered between the wavelengths of the blue and green channels may result in an additive emission spectrum for nucleotide “C” 320 and / or nucleotide “T” 330. For example, a mixture of a third dye and a fourth dye may replicate the emissionspectrum of either nucleotide “C” 320 and / or nucleotide “T” 330, if the third dye and a fourth dye have emissions centered near the transition between the wavelengths of the blue and green channels, with the emission of the third dye at a slightly higher energy than the fourth dye. If the emissions spectrum of the two dyes overlap, then the peak of the additive emission spectra may be moved to higher or lower energies based on the ratio of the two dyes.
[0128] An aspect of the disclosure is directed to methods and systems that may avoid or minimize common microscopy problems associated with crosstalk wherein fluorophores emitting light at one set of wavelengths disrupts or affects detection of signals at other wavelengths. As described above, in a sequencing-by-synthesis system, fluorescent dyes are attached to each of the four nucleotides (A, C, G, and T) that are added to a growing DNA strand during synthesis. As each nucleotide is incorporated, one or more excitation wavelengths is used to excite fluorescent dyes attached to the growing DNA strands, and the corresponding dyes each emit a fluorescent signal that is detected by a camera or other imaging system. Accordingly, crosstalk can refer to both excitation crosstalk (cross-excitation), where multiple fluorophores are excited with a single excitation wavelength which is intentional in a one excitation, two channel scheme, and emission bleed-through, where unwanted fluorescent signal from a dye with an overlapping emission (and excitation) spectrum is detected in another emission detection channel. Overlapping spectra may contribute to false negatives or positives, or otherwise obscure data. However, in some sequencing systems, the fluorescent dyes used may have the potential to crosstalk with one another during the sequencing process, leading to inaccurate sequencing results.
[0129] Fig. 5A illustrates an example of cross-talk in a one excitation two channel system, and, in reference to Fig. 4, illustrates the benefits of implementing methods and systems disclosed herein to maximize signal to noise. The left panel of Fig. 5A shows an illustration of a typical detection scheme for 1Ex-2Ch system, and three corresponding emission spectra. The spectrograph displays a model emission spectrum of a short Stokes shift dye 411, which here corresponds to nucleotide “T.” The emissions spectrum of short Stokes shift dye 411 shows emission predominantly in the blue channel 402, but with some emission intensity from a diminishing tail of fluorescent emissions in the green channel 406. The spectrograph also shows the model emission spectrum of a medium Stokes shift dye 421, which here corresponds to nucleotide “C.” The emissions spectrum of medium Stokes shift dye 421 shows emission in both the blue channel 402 and the green channel 406, and at a higher intensity than the emissions ofeither the short or long Stokes shift dye. As described above, the overall intensity of emissions may be reduced by diluting the fraction of nucleotides with bright dyes with non-fluorescent labels. Finally, the spectrograph displays a model emission spectrum of a long Stokes shift dye 431, which here corresponds to nucleotide “A.” The emissions spectrum of long Stokes shift dye 431 shows emission predominantly in the green channel 406, but with some emission intensity from a diminishing tail of fluorescent emissions in the blue channel 402.
[0130] The right panel of Fig. 5A shows the results from processing the emissions in the left panel in a two-dimensional scatterplot. The Y axis 401 corresponds to the intensity of fluorescent emissions within the green channel 402. The X axis 405 correspond to the intensity in a blue channel 406. The right panel of Fig.5A shows the relative positions of the four nucleotides in a typical 1Ex-2Ch system.
[0131] Figure 5B shows experimental results of a typical 1Ex-2Ch system, with the positions of the nucleotides matching those in Fig.5A. Nucleotide “A” 410 is shown as a cloud in the top left corner, with nucleotide “C” 420 as a cloud in the top right, nucleotide “T” 430 as a cloud in the lower right, and nucleotide “G” 440 as a cloud at the origin of the scatterplot in the lower left corner. The respective location of each nucleotide may be exchanged, with nucleotide “G,” for example, switching position with any other nucleotide by switching the fluorophore associated with nucleotide “A” to be conjugated with nucleotide “G,” and vice versa.
[0132] Fig.5B also includes two arrows pointing from each of nucleotide “A” 410, and nucleotide “T” 430 towards the Y-axis and X-axis, respectively. The arrows point out an example of “crosstalk,” where each cloud has non-zero absorbance in the unintended channel because both the emission spectra of short Stokes shift dye 411 and of long Stokes shift dye 431 have non-zero emissions in both the green channel 406 and blue channel 402, respectively. Generally, to avoid crosstalk, pairs of dyes with good separation between their respective excitation and emission spectra would be chosen. However, because the signals from each dye can be very close in wavelength, there is a risk of crosstalk between the dyes. This means that the signal from one dye may be detected by the camera as if it were a signal from a different dye. For example, the signal from the A nucleotide dye may be detected as if it were the signal from the G nucleotide dye, leading to an incorrect base call.
[0133] Typical measures to address issues with base calling attempt to minimize crosstalk, by, for example using dyes with very distinct spectra that are less likely to overlap. Thenumber of dyes to be used in a sequencing system depends on the number of target molecules that need to be detected and the number of channels that are used to detect them. In general, as the number of target molecules and channels increases, the number of dyes required also increases. For example, in a polynucleotide sequencing experiment, where four different bases need to be detected in four different channels, as many as four different dyes may be used to ensure that the emissions from each base can be distinguished from each other.
[0134] In a 1Ex-2Ch system, where just one wavelength of light is used to excite two dyes, one ffN may be used per cloud. The emissions from each dye is then separated into two different channels based on the dyes’ emission wavelengths (instead of excitation and emission wavelengths). The emissions from each dye may also cover two channels and appear as a corner cloud on a scatterplot. Differentiating the emissions of dyes into separate channels is achieved through dyes that absorb light at the same wavelength but emit at different wavelengths, to ensure that the emissions from each dye can be distinguished from each other.
[0135] In a 2Ex-2Ch system, as described below, two different fluorescent dyes may be used to label each nucleotide molecules, with each dye being excited by a different wavelength of light. The emissions from each dye are then detected separately in two different channels. Because this system excites each dye by a different wavelength of light, a different type of dye may be used to generate emissions for each channel. Usually, a two excitation two channel system will use two ffNs per cloud to modulate the intensities in each channel. Diamond Configuration For Two Excitation, Two Channel Systems And Methods
[0136] An aspect of the disclosure is directed to the diamond scatterplot configuration as described above, but implemented on a 2Ex-2Ch system. A diamond scatterplot from a 2Ex- 2Ch system as described herein is an alternative scheme that avoids a ‘dark’ base and therefore may provide additional advantages. In this embodiment, two clouds are created off the axes having a mixture of nucleotides to modulate intensities in the two channels and achieve a diamond shape. Different dyes may be used for each nucleotide in 1Ex-2Ch systems, but in 2Ex-2Ch systems the chemistry of the same two dyes might be used for four nucleotides. In two excitation two channel chemistry, the centers of each cloud may be relocated by using different dyes or mixtures of dyes to change the amount of emission intensity that is detected in either channel one or channel two for each cluster of nucleotides.
[0137] Fig. 6 shows a successful proof of concept for the diamond configuration for a 2Ex-2Ch sequencing system, which was performed on an ILLUMINA® MiSeq sequencing system. The number of channels used by the MiSeq system depends on the specific reagent kit and sequencing application being used. For example, the MiSeq Reagent Kit v2 typically uses two channels, while the MiSeq Reagent Kit v3 typically uses four channels. The MiSeq Reagent Kit v3 also has a higher output than the v2 kit, with the ability to generate up to 25 million paired-end reads per run, compared to up to 15 million paired-end reads with the v2 kit. Here, the MiSeq system is configured in a 2Ex-2Ch setup, and, as described above, uses two ffNs per cloud to modulate the intensities in the two-dimensional base calling scatterplot.
[0138] The top left panel of Fig. 6 lists the ffNs and concentrations of ffNs used for each nucleotide. The three remaining panels show two-dimensional scatterplots from three different sequencing runs on the MiSeq system in a two excitation, two channel configuration. The X and Y axes of the scatterplots correspond to intensities of fluorescent signals detected in a detection channel. The bottom left panel shows nucleotide “T” 510 as a cloud in the top center location and corresponds to partial detections of fluorescence in both channels. Nucleotide “T” 510 is labelled with fluorescent dyes ffT-NR455BoC and ffT-NR550S0. The bottom left panel shows nucleotide “C” 520 as a cloud in the middle right location and also corresponds to partial detections of fluorescence in both channels, but a maximum in channel 1 (X-axis). Nucleotide “C” 520 is labelled with fluorescent dyes ffC-NR455BoC and ffC-NR550S0. The bottom left panel shows nucleotide “A” 530 as a cloud in the bottom center location and corresponds to partial detections of fluorescence in channel 1 (X-axis). Nucleotide “A” 530 is labelled with fluorescent dyes ffA-BL-NR455BoC and ffA-BL-NR650C5. The bottom left panel shows nucleotide “G” 540 as a cloud in the middle-left location and corresponds to partial detections of fluorescence in channel 2 (Y-axis). Nucleotide “G” 540 is labelled with fluorescent dyes ffG-dark and ffG-PEG12- ATTO532. In some embodiments, any of the nucleotide may be switched, such that nucleotide “T” is in the position of nucleotide “C” or vice versa. Note that the top right panel has the clouds for nucleotides “A” and “C’ switched relative to the lower left panel nucleotides “A” 530 and “C’ 520.
[0139] Fig. 6 is one embodiment of the diamond configuration that avoids a ‘dark’ base. Note that none of the X-Y scatterplots include a large cloud at the origin of the scatterplot, which indicates an unlabeled base. However, in some embodiments there may be a fifth cloud thatcorresponds to empty or otherwise unlabeled clusters on a flowcell. Because the emissions of all four nucleotides are removed from the origin, when there is a fifth cloud derived from empty wells, these wells could be detected as empty instead of being labeled as “G’ with high confidence.
[0140] Fig. 6 demonstrates an embodiment of a 2Ex-2Ch system and method encompassed by a method of sequencing polynucleotides bound to a flowcell, comprising: detecting fluorescent emissions from a first labeled nucleotide at a first wavelength and a first intensity; detecting fluorescent emissions from a second labeled nucleotide at a second wavelength and a second intensity, wherein the first wavelength is different from the second wavelength; detecting fluorescent emissions from a third labeled nucleotide at the first and second wavelengths at third and fourth intensities; detecting fluorescent emissions from a fourth labeled nucleotide at the first and second wavelengths at fifth and sixth intensities; and determining the sequence of the polynucleotides based on the detected fluorescent emissions and intensities.
[0141] Specifically, nucleotide “G” 540 in Fig. 6 may correspond to a first labeled nucleotide with emissions at a first wavelength and a first intensity, where the first wavelength is an emission wavelength within a green channel and the first intensity is approximately 33% of a maximum intensity in the green channel (with arbitrary units). A first labeled nucleotide that corresponds nucleotide “G” 540 in Fig.6 may be labeled with a single partially labeled dye that is diluted with a non-fluorescent label or a combination of two dyes. In some embodiments the emissions may be replicated with any of a single dye (with crosstalk), a single partially labeled dye that is diluted with a non-fluorescent label, and a combination of two dyes.
[0142] Nucleotide “A” 530 in Fig. 6 may correspond to a second labeled nucleotide with emissions at a second wavelength and a second intensity, wherein the first wavelength is different from the second wavelength, where the second wavelength is within a blue channel, and the second intensity is approximately 33% of a maximum intensity in the blue channel (with arbitrary units). The emissions of Nucleotide “A” 530 illustrated in Fig. 6 may be replicated with a single partially labeled dye that is diluted with a non-fluorescent label or a combination of two dyes. In some embodiments the emissions may be replicated with any of a single dye (with crosstalk), a single partially labeled dye that is diluted with a non-fluorescent label, and a combination of two dyes.
[0143] Nucleotide “T” 510 in Fig.6 may correspond to a third labeled nucleotide with emissions at the first and second wavelengths at third and fourth intensities, wherein the thirdintensity is approximately 66% of a maximum intensity in the green channel, and the fourth intensity is approximately 33% of a maximum intensity in the blue channel. Nucleotide “C” 520 in Fig.6 may correspond to a fourth labeled nucleotide at the first and second wavelengths at fifth and sixth intensities, wherein the fifth intensity is approximately 33% of a maximum intensity in the green channel, and the sixth intensity is approximately 66% of a maximum intensity in the blue channel. The emissions of Nucleotide “C” or “T” illustrated in Fig. 6 may be each replicated with any of a single dye, a single partially labeled dye that is diluted with a non-fluorescent label, and a combination of two dyes.
[0144] In each of the three X-Y scatterplots of Fig.6, there is a dot in the middle of the clouds 510, 520, 530, and 540, where the dot corresponds to the approximate center of each cloud and provides a point of reference for a real time analysis algorithm to assign each data point into a base call. An aspect of the disclosure is directed to accurate base calling for non-standard scatterplot configurations. Some underlying assumptions including “base call diversity” may be included in traditional RTA methods– meaning at any given cycle for any given population of clusters, the base representation is roughly equal in ratio. While this is true for naturally occurring DNA bases, empty wells (or additional modified bases) do not necessarily satisfy the equal diversity assumption. In some embodiments, an RTA algorithm may be trained to mitigate potential issues with assigning empty wells that have a low probability of occurring.
[0145] The three X-Y scatterplots of Fig. 6 show embodiments of an experimental RTA performed on an alternative scatterplot shape. The lower left panel and the top right panel show successful RTA performed on a diamond shaped scatterplot with good separation between the clouds and high quality base calls (Q30) for the four nucleotides. Note that an untrained RTA model may assign parts of all four clouds, accordingly, as a comparison, a naïve RTA was performed on the lower right panel and each cloud is assigned an equal likelihood for each of the base calls.
[0146] In some embodiments, software methods may remove empty wells by filtering on any cluster with 9 or 10 G nucleotides in the first 10 cycles, when the G nucleotides are unlabeled. However, some embodiments with filters may also remove any cluster whose true sequence was 9 or 10 Gs in the first 10 cycles, so there may be some false detection to the extent that sequence context is present in the samples and it starts at cycle 1. Furthermore, when clusters sequence the whole insert and are no longer incorporating, poly G will be called for the remainderof the read. This may be trimmed in adapter trimming by specifying poly G as an adapter sequence during fastq generation. By partially labeling the G, a fifth cloud may be present in a two- dimensional plot of intensities in 2-channel SBS.
[0147] In some embodiments, traditional 2-dimensional data could still be modeled using four clouds and the empty wells would likely be called G. The detection of poly G could still be done with 9 or 10 Gs in the first 10 cycles. In addition, or in the alternative, in some embodiments, disclosed methods may take the population of the clusters called G and segment them, using binary segmentation such as Otsu’s method, in the y dimension into two populations, empty and occupied G. This may be done for anywhere from one or more cycles to build confidence in the wells being assigned as empty. In some embodiments, the disclosed methods may modify the model for the data to include five populations, the four for A, C, G, and T and one more for the empty wells at the origin and then perform expectation maximization to learn the maximum likelihood parameters of the 5-cloud model. In some embodiments, other segmentation algorithms such as k-means could also be used in place of expectation maximization. The disclosure provides for several permutations on this method. For example, one method may only use the 5-cloud model for a single cycle, a few early cycles, or every cycle in the sequencing run. In some embodiments, the fifth cloud may be assigned as “N” until the cluster is filtered out after some number of early cycles. In some embodiments, this method may be present in the model for every cycle. In some embodiments, disclosed methods may be employed for detecting that an insert had been completely sequenced because it returned an “empty” well. The following non- limiting examples demonstrate methods for mitigating potential issues with assigning bases that have low base diversity. One mitigation may include introducing a known sequence that includes multiples of a less abundant fifth base and allow the RTA to train on that data set. In other examples, such as biotechnological applications like a tumor selection, methods may first artificially increase the diversity of base pairs by adding in unnatural base pairs and then sequencing with more than four base pairs.
[0148] An aspect of the disclosure is directed to effective solutions where, in some embodiments, the disclosed methods provide a sufficient signal-to-noise ratio to distinguish empty cells and partially labeled G bases. If there is an overlap between empty clouds and G clouds, then a new error mode will be introduced of miscalling G as empty and vice versa. In some embodiments, the G cloud may be moved far enough away from the empty cloud so they do notoverlap. In some embodiments, the cloud for a different base than G may be replaced for G, and the same sequencing may be run where G is no longer in the "empty" location, and thereby G may be moved far enough away from the empty cloud, so they do not overlap. In some embodiments, the disclosure provides methods to address an issue where the effectiveness of procedures depends on the occupancy and the population of the empty cloud.
[0149] Some disclosed solutions may be robust across an extensive range of loading concentrations and may yield differences in the population of the fifth cloud. Some embodiments may adjust dye concentrations, ratios, or wavelengths such that the clouds for each base are not aligned with an axis. Some embodiments may control the degree of labeling to maximize the distance between bases. Some embodiments may locate the G cloud by, for example, adjusting ratios, dyes, and wavelengths to minimize the quenching of other adjacent clouds or subsequent bases in the sequences. Amplitude Multiplexing with Labeled “G” Nucleotide
[0150] An aspect of the disclosure is directed to 2Ex-2Ch systems and methods that label at least four bases. Compared to typical methods and systems that use a “dark base” to identify a fourth nucleotide, usually “G”, some embodiments label the four main nucleotides. As described above, labeling “G” instead of relying on the lack of detection to identify a nucleotide will separate the detection of the “G” nucleotide from empty clusters. Also, by labeling “G” different levels of labeling or brighter nucleotides could be used to optimize the separation of clouds in a scatterplot, thereby improving the reliability of base calling and boosting the signal to noise ratio.
[0151] The next series of figures demonstrate embodiments where the “G” nucleotide is labeled in the green channel. Figs. 7A-7C show three panels, each showing a scatterplot with four clouds corresponding to the fluorescence emissions of the dyes labelled to each nucleotide. Fig. 7A shows an X-Y scatterplot of the fluorescent emissions of a system where all four nucleotides are labeled. Here, the Y axis corresponds to intensity of fluorescent emissions in a green channel, and the X axis corresponds to the intensity of fluorescent emissions in a blue channel. The four circles in Fig. 7A correspond to the fluorescence emissions of nucleotide “T” 710, nucleotide “C” 720, nucleotide “A” 730, and nucleotide “G” 740. The “G” nucleotide 740 is shown as a circle with three quarters shaded black and the remaining quarter shaded green, andillustrates that “G” nucleotide 740 may be partially labelled with a dye that fluoresces in the same channel as the “T” nucleotide 710, but at a lower intensity than the “T” nucleotide 710. The partial labeling of “G” nucleotide 740 is highlighted by the green arrow underneath the cloud of G” nucleotide 740.
[0152] In some embodiments, the “G” nucleotide 740 may be labeled with the same fluorescent dye as the “T” nucleotide 710, but a fraction of the “G” nucleotide 740 may be labeled with a non-fluorescent label such that the intensity in the green channel is lower than for a cluster with a base call corresponding to “T” nucleotide 710. In some embodiments, the intensity of emissions for “G” nucleotide 740 may be quantified as a percentage of the intensity of emissions for “T” nucleotide 710, wherein the percentage may be one of 90%, 80%, 70%, 60%, 50%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, and 5% of the intensity of emissions of “T” nucleotide 710 in the green channel.
[0153] The intensity of the fluorescent emissions of the “T” nucleotide 710 in the green channel is shown as approximately the same as the intensity of the fluorescent emissions of the “C” nucleotide 720 in the green channel. The “C” nucleotide 720 is shown as a circle with halves shaded in green and blue, corresponding to the labeling of the fluorescent dyes that emit in the green and blue channels. The “C” nucleotide 720 is shown in the top right corner of the scatterplot and illustrates that a nucleotide may be labeled in two channels. The “A” nucleotide 730 is shown as a circle with one half shaded black and the remaining half shaded blue, and illustrates that “A” nucleotide 730 may be partially labelled with a dye that fluoresces in the same channel and intensity as the “C” nucleotide 720.
[0154] Fig. 7B shows an X-Y scatterplot of the fluorescent emissions of a system where all four nucleotides are labeled, and where all four nucleotides are well resolved from dark or otherwise empty clusters. Again, the Y axis corresponds to intensity of fluorescent emissions in a green channel, and the X axis corresponds to the intensity of fluorescent emissions in a blue channel. The four circles in Fig. 7B correspond to the fluorescence emissions of nucleotide “T” 712, nucleotide “C” 722, nucleotide “A” 732, nucleotide “G” 742, and empty clusters 743. By labeling nucleotide “G” the positions of the other three nucleotides may be rearranged as well. Fig. 7B shows the different levels of labeling and / or brighter dyes on nucleotides that may be used to optimize the separation of clouds in a scatterplot. Nucleotide “T” 712 is shown as a green circle in the top left of Fig.7B, and illustrates that a bright dye or higher level of labeling with a bright dyemay be used to increase the intensity of fluorescent emissions associated with clusters corresponding to nucleotide “T” 712. Nucleotide “T” 712 is shown with an intensity in the green channel that is greater than the intensity in the green channel for nucleotide “C” 722.
[0155] The “G” nucleotide 742 is shown as a green circle, and illustrates that “G” nucleotide 742 may be substantially labelled with a dye that fluoresces in the same channel as the “T” nucleotide 712, but at a lower intensity than the “T” nucleotide 712. For example, “G” nucleotide 742 may be labeled with the dye corresponding to “T” nucleotide 710 in Fig. 7A, and “T” nucleotide 712 may be labeled with an even brighter dye. To conserve the separation between the nucleotides in the scatterplot, the increase in intensity of “G” nucleotide 742 may be accompanied by an increase in intensity of the “T” nucleotide 712. The increased intensity of both the “T” nucleotide 712 and the “G” nucleotide 742 is highlighted by the green arrows underneath the cloud of nucleotides 712 and 742. The lack of emissions corresponds to empty wells 743, which is shown as well resolved from the other nucleotides 712-742.
[0156] Fig. 7B shows nucleotide “C” 722 as a circle in the right side of the scatterplot that is shaded half green and half blue, which illustrates that a nucleotide may be labeled with one or two dyes that emits in both the blue and green channels. Note that a single bright dye that has a broad range of emission wavelengths may emit in both the green and blue channels. In some embodiments, nucleotide “C” 722 may be labeled with two dyes each emitting separately in the green or blue channels. If nucleotide “C” 722 is half-labeled with the same green emitting dye as for nucleotide “T” 712, then the maximum intensity of nucleotide “C” 722 in the green channel will be approximately half the intensity of nucleotide “T” 712. The “A” nucleotide 732 is shown as a circle with one half shaded black and the remaining half shaded blue, and illustrates that “A” nucleotide 732 may be partially labelled with a dye that fluoresces in the same channel and intensity as the “C” nucleotide 722. Unlike “T” nucleotide 712, “A” nucleotide 732 is shown to match the approximate intensity of nucleotide “C” in its respective channel. In some embodiments, matching the intensity of fluorescent emissions between two nucleotides in one channel, while being distinguishable in the other channel, may be advantageous for signal processing and improved signal to noise ratio. Accordingly, in some embodiments, the intensity of green emissions may be the same between nucleotide “C” 722 and nucleotide “G” 742, and the intensity of blue emissions may be the same between nucleotide “C” 722 and nucleotide “A” 732.
[0157] Fig. 7C shows an X-Y scatterplot of the fluorescent emissions of a system where all four nucleotides are labeled, where nucleotide “T” and nucleotide “G” are separated by a partial labeling of a green dye, similar to Fig.7A, but where the dye for nucleotide “T” is brighter. The brighter nucleotide “T” 715 is highlighted by green arrow. Again, the Y axis corresponds to intensity of fluorescent emissions in a green channel, and the X axis corresponds to the intensity of fluorescent emissions in a blue channel. The four circles in Fig. 7B correspond to the fluorescence emissions of nucleotide “T” 715, nucleotide “C” 725, nucleotide “A” 735, and nucleotide “G” 745. Fig. 7C shows nucleotide “C” 725 and nucleotide “A” 735 in the same approximate positions as in Fig. 7A and 7B for nucleotide “C” and nucleotide “A,” respectively. The “G” nucleotide 745 is shown as a circle with three quarters shaded black and the remaining quarter shaded green, and illustrates that “G” nucleotide 745 may be partially labelled with a dye that fluoresces in the same channel as the “T” nucleotide 715, but at a lower intensity than the “T” nucleotide 715. The partial labeling of “G” nucleotide 745 is highlighted by the green arrow underneath the cloud of G” nucleotide 745.
[0158] In some embodiments, the “G” nucleotide 745 may be labeled with the same fluorescent dye as the “T” nucleotide 710 in Fig. 7A, (or the brighter dye for the “T” nucleotide 712 in Fig.7B), but with a fraction of the “G” nucleotide 745 labeled with a non-fluorescent label such that the intensity in the green channel is lower than for a cluster with a base call corresponding to “T” nucleotide 715. In some embodiments, the intensity of emissions for “G” nucleotide 745 may be quantified as a percentage of the intensity of emissions for “T” nucleotide 715, wherein the percentage may be one of 90%, 80%, 70%, 60%, 50%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, and 5% of the intensity of emissions of “T” nucleotide 715 in the green channel.
[0159] Figs. 7A, 7B and 7C demonstrate embodiments encompassed by a method of sequencing polynucleotides bound to a flowcell, comprising: detecting fluorescent emissions from a first labeled nucleotide at a first wavelength and a first intensity; detecting fluorescent emissions from a second labeled nucleotide at a second wavelength and a second intensity, wherein the first wavelength is different from the second wavelength; detecting fluorescent emissions from a third labeled nucleotide at the first and second wavelengths at third and fourth intensities; detecting fluorescent emissions from a fourth labeled nucleotide at the first and second wavelengths at fifth and sixth intensities; and determining the sequence of the polynucleotides based on the detected fluorescent emissions and intensities.
[0160] Specifically, nucleotide “G” 740 in Fig. 7A and nucleotide “G” 745 in Fig. 7C may correspond to a first labeled nucleotide with emissions at a first wavelength and a first intensity, where the first wavelength is an emission wavelength within a green channel and the first intensity is approximately 25% of a maximum intensity in the green channel (with arbitrary units). Similarly, nucleotide “G” 742 may correspond to a first labeled nucleotide with emissions at a first wavelength and a first intensity, where the first wavelength is an emission wavelength within a green channel and the first intensity is approximately 50% of a maximum intensity in the green channel, but otherwise approximately the same intensity in the green channel as for nucleotide “C” 722. Nucleotide “G” in Figs. 7A, 7B, and 7C may be labeled with any of a single dye, a single partially labeled dye that is diluted with a non-fluorescent label, and two dyes.
[0161] Nucleotide “A” 730 (or 732 and 735) in Fig. 7A may correspond to a second labeled nucleotide with emissions at a second wavelength and a second intensity, wherein the first wavelength is different from the second wavelength, where the second wavelength is within a blue channel, and the second intensity is approximately 50% of a maximum intensity in the blue channel (with arbitrary units). The emissions of Nucleotide “A” 730 illustrated in Fig.7A may be replicated with any of a single dye, a single partially labeled dye that is diluted with a non-fluorescent label, and a combination of two dyes.
[0162] Nucleotide “C” 730 (or 732 or 735) in Fig. 7A may correspond to a fourth labeled nucleotide at the first and second wavelengths at fifth and sixth intensities, wherein the fifth intensity is approximately 50% of a maximum intensity in the green channel, and the sixth intensity is approximately 50% of a maximum intensity in the blue channel. The emissions of Nucleotide “C” or “T” illustrated in Fig. 7A may be each replicated with any of a single dye, a single partially labeled dye that is diluted with a non-fluorescent label, and a combination of two dyes. Nucleotide “T” 710 (or 712 or 715) in Fig.7A may correspond to a third labeled nucleotide with emissions at the first and second wavelengths at third and fourth intensities, wherein the third intensity is approximately 100% of a maximum intensity in the green channel, and the fourth intensity is approximately zero intensity in the blue channel.
[0163] Alternative scatterplots that change the intensity of certain base calls can have implications for the accuracy and quality of polynucleotide sequencing data. In traditional scatterplots for two-channel detection in polynucleotide sequencing, the intensity of the signal in each channel is used to identify the base call for each nucleotide. However, in some cases, thesignal intensity for a particular base call may be too low, which can result in errors or ambiguities in the sequencing data.
[0164] By changing the intensity of certain base calls in the examples of alternative scatterplots, the disclosed methods and systems may improve the accuracy and reliability of sequencing data. For example, if the signal intensity for a particular base call is consistently low, an alternative scatterplot may be used to increase the intensity of that base call in the scatterplot, which could help to distinguish it from other base calls with nearby signal intensities in the X-Y plane. Alternatively, if a particular base call is consistently associated with a high level of noise or interference, the researcher may choose to decrease the intensity of that base call in the scatterplot and separate the nucleotide from other bases, which could help to reduce the level of noise and improve the accuracy of the sequencing data. By labeling nucleotide “G” the positions of the other three nucleotides may be rearranged as well. Fig. 7A-7C shows the different levels of labeling and / or brighter dyes on nucleotides that may be used to optimize the separation of clouds in a scatterplot. Amplitude Multiplexing with Unlabeled “G”
[0165] An aspect of the disclosure is directed to alternative scatterplots that may encode more than one nucleotide per signal state in a channel. Such systems and methods may still use an unlabeled or partially labeled base (here, nucleotide “G” is used as a nonlimiting example). Fig.8 illustrates two-channel systems and methods that use two different dyes in one same channel, but where the different dyes have differing intensities in that channel. If a typical two channel scatterplot is referred to as a square scatterplot, Fig. 8 displays an embodiment with clouds arranged in an L-Shape on a scatterplot.
[0166] Fig.8 shows an X-Y scatterplot of the fluorescent emissions of a system where three nucleotides are labeled, and “G” nucleotide is unlabeled. The Y axis corresponds to intensity of fluorescent emissions in a green channel, and the X axis corresponds to the intensity of fluorescent emissions in a blue channel. The four circles in Fig. 8 correspond to the fluorescence emissions of nucleotide “T” 810, nucleotide “A” 820, nucleotide “C” 830, and nucleotide “G” 840. Nucleotide “T” 810 and nucleotide “A” 820 are shown as solid circles and are both shown to emit in the green channel.
[0167] Fig. 8 shows both nucleotide “T” 810 and nucleotide “A” 820 as solid circles; however, the nucleotides may be partially labelled. In some embodiments, the “A” nucleotide 820 may be labeled with a different fluorescent dye as the “T” nucleotide 810. In some embodiments, the “A” nucleotide 820 may be labeled with the same fluorescent dye as the “T” nucleotide 810, but with a fraction of the “A” nucleotide 820 labeled with a non-fluorescent label such that the intensity in the green channel is lower than for a cluster with a base call corresponding to “T” nucleotide 810. In some embodiments, the intensity of emissions for “A” nucleotide 820 may be quantified as a percentage of the intensity of emissions for “T” nucleotide 810, wherein the percentage may be one of 90%, 80%, 70%, 60%, 50%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, and 5% of the intensity of emissions of “T” nucleotide 810 in the green channel.
[0168] Nucleotide “G” 840 is shown as a solid black circle and illustrates a nucleotide identified by the lack of fluorescent emissions. However, in some embodiments, the nucleotide “G” may be partially labeled as shown in Fig. 7A, for example, and still be in an L-shaped scatterplot. Nucleotide “C” 830 is shown as a solid blue circle and is shown to emit in the blue channel. Nucleotide “C” 830 is shown as the only nucleotide emitting in the blue channel, and the other labeled nucleotides also only do not need to labeled with two dyes—one each for two channels. Accordingly, the dyes used for each channel may be optimized for brightness. For example, the intensity of emissions of nucleotide “C” 830 in the blue channel may be approximately twice as intense as the emissions of nucleotide “C” 720 in Fig.7A.
[0169] Fig. 9A displays an experimental example of an L-shape scatterplot including an X-Y scatterplot of the fluorescent emissions of a two-channel system, where three nucleotides are labeled, and “G” nucleotide is unlabeled. This experimental example provides a practical application of the concepts illustrated in the preceding figures and demonstrates successful application on a polynucleotide sequence. The Y axis 901 corresponds to intensity of fluorescent emissions in a first channel, and the X axis 905 corresponds to the intensity of fluorescent emissions in a second channel. The four clouds in Fig.9A correspond to the fluorescence emissions of nucleotide “T” 910, nucleotide “A” 920, nucleotide “C” 930, and nucleotide “G” 940. Nucleotide “A” 920 and nucleotide “C” 930 are both shown to emit in the second channel. Fig. 9A provides an experimental example of an L-shaped scatterplot with two nucleotides encoded at different intensities of the signal state of channel two.
[0170] Beneath the X-Y scatterplot in Fig.9A, a table lists the fluorescent and / or non- fluorescent labels used for each nucleotide and the respective concentrations. Nucleotide “G” 940 is shown as fully labeled with a non-fluorescent dye. Nucleotide “T” 910 is shown as partially labeled with "ffT-AF550POPOSO," and could potentially have an increased intensity by fully labelling the nucleotides. Nucleotide “A” 920 is listed as fully labeled with ffA-NR455BoC. Nucleotide “C” 930 is listed as partially labeled with ffC-NR455BoC and ffC-SO7181.
[0171] Fig. 9B provides an experimental example of an L-shaped scatterplot with two nucleotides encoded at different intensities of the signal state of channel two. Fig. 9B displays experimental data in an L-shape scatterplot including an X-Y scatterplot of the fluorescent emissions of a two-channel system, where three nucleotides are labeled, and “G” nucleotide is unlabeled. The Y axis 901 corresponds to intensity of fluorescent emissions in a first channel, and the X axis 905 corresponds to the intensity of fluorescent emissions in a second channel. The four clouds in Fig.9A correspond to the fluorescence emissions of nucleotide “T” 911, nucleotide “A” 921, nucleotide “C” 931, and nucleotide “G” 940. Nucleotide “A” 920 and nucleotide “C” 930 are both shown to emit in the second channel.
[0172] Beneath the X-Y scatterplot in Fig.9B, a table lists the fluorescent and / or non- fluorescent labels used for each nucleotide and the respective concentrations. Nucleotide “G” 940 is shown as fully labeled with a non-fluorescent dye. Nucleotide “T” 911 is shown as partially labeled with "ffT-AF550POPOSO," and partially labeled with a dark nucleotide. Accordingly, nucleotide “T” 911 could potentially have an increased intensity by fully labelling the nucleotide with fluorescent dyes. Nucleotide “A” 921 is listed as fully labeled with ffA-NR560A. Nucleotide “C” 930 is listed as fully labeled with ffC-NR455BoC. The concepts illustrated by examples Figs. 9A and 9B may be combined with the other examples disclosed herein. For example, a partially labeled “G’ may be incorporated into either Fig. 9A or 9B where the partially labeled “G” nucleotide is labeled in either the channel that already includes two nucleotides, or in the channel that only includes one nucleotide.
[0173] Figs.9A and 9B are color coded with back, blue and green sections, and several red data points, which corresponds to an automatic real-time analysis that was partially applied to this experimental example. Real-time analysis in Illumina sequencing systems refers to the automated and simultaneous analysis of the sequencing data during the sequencing run itself, as opposed to waiting until the sequencing run is complete before analyzing the data. RTA softwareprocesses these raw reads as they are generated and provides quality metrics and other key information to the researcher in real-time. An aspect of this disclosure is directed to training an RTA software process to accommodate alternative scatterplot shapes.
[0174] RTA software, which is used for the analysis of sequencing data, may be implemented via a machine learning-based algorithm that has been trained using large amounts of data from previous sequencing runs. However, other examples such as supervised learning or encoded training may be used to determine base calls. During the RTA training process, the RTA algorithm may be exposed to a range of sequencing data including the experimental examples in Fig. 9A, and also including different types of samples, sequencing conditions, and experimental designs. The algorithm analyzes this data and learns to recognize patterns and correlations in the sequencing data, such as the location of the nucleotide clouds, which can be used to improve the accuracy and efficiency of the sequencing workflow.
[0175] A training process may involve several steps, including data preprocessing, feature extraction, and model training. During the data preprocessing step, the sequencing data is cleaned, normalized, and transformed into a format that can be used for further analysis. In the feature extraction step, key features of the sequencing data, such as base call quality scores and signal intensities, are extracted and used to train the machine learning model. Finally, in the model training step, the machine learning model is trained using a variety of techniques, such as supervised learning, unsupervised learning, and reinforcement learning, to improve its accuracy and generalization ability.
[0176] Once the RTA algorithm has been trained, the model can be used in real-time during the sequencing run to analyze the sequencing data and provide quality metrics and other key information to the researcher. The algorithm is able to make accurate predictions based on the patterns and correlations it has learned during the training process, allowing it to quickly identify any issues or errors that may be affecting the quality of the sequencing data and make adjustments in real-time.
[0177] Fig.10A displays a panel of the results of five experimental examples resulting in five different X-Y scatterplots of differing degrees of labelled G on an ILLUMINA® NextSeq2K system. The Y axis corresponds to intensity of fluorescent emissions in a first channel, and the X axis corresponds to the intensity of fluorescent emissions in a second channel. The panel progresses from left to right, where the labeling of the “G” nucleotide is increased from a baselineof no labeling to progressively higher labeling of G with a fluorescent dye, and therefore higher emissions in the first channel. The first panel in the top left, corresponding to baseline of unlabeled G, shows four clouds of nucleotides, where each is labeled T, A, C, and G near the origin of the X-Y scatterplot. The next panel corresponds to 30% labeling of the “G” nucleotide with a traditional “T” nucleotide dye—that has been functionalized for “G”--T-AF550POPOS0. After partially labeling the “G” nucleotide, the corresponding cloud is observed to increase slightly along the Y-Axis, to higher intensity values in channel 1. The trend continues in the next panel, which corresponds to 30% labeling of the “G” nucleotide with a different traditional “T” nucleotide dye— that has been functionalized for “G”--T-NR550S0. These two panels demonstrate that the labeling of the “G” nucleotide may be controlled by either the type or quantity of fluorescent dye.
[0178] The next panel of Fig. 10A is labeled T-NR550S050% FFG, and corresponds to 50% labelling of the “G” nucleotide with 50% T-NR550S0. The cloud corresponding to “G” is observed to increase slightly along the Y-Axis, to higher intensity values in channel 1. Good separation between the four clouds is visible. The last panel is labeled T-NR550S0100% FFG, and corresponds to 50% labelling of the “G” nucleotide with 100% T-NR550S0. The cloud corresponding to “G” is observed to be shifted higher on the Y-Axis and is completely separated from the origin. Even with 100% coverage with T-NR550S0, the cloud corresponding to the “G” nucleotide has less intensity in the first channel than the cloud corresponding to the “T” nucleotide, which indicates that the “T” nucleotide uses a brighter dye than the “G” nucleotide. While the last panel of Fig.10A shows an increased separation of the “G” cloud from the origin of the scatterplot, the separation of the four clouds has reduced, such that distinguishing base calls may be more difficult. However, the separation between the “G” cloud and the “C” cloud, for example, can be increased by also increasing the brightness of the dye associated with the “C” nucleotide.
[0179] Fig. 10B shows a graph of the resulting Q30 scores from each of the experiments in Fig. 10A. The Y-Axis corresponds to the percentage of base calls that have a high base calling accuracy as a function of cycles of the system, where each cycle labels and base calls a new nucleotide of the polynucleotide to be sequenced. A next-generation sequencing experiment is made up of discrete steps that each contribute to the overall quality of a data collection in a unique way. Metrics for sequencing quality can give essential information about the correctness of each stage in the process, such as library preparation, base calling, read alignment, and variant calling. The most frequent metric used to quantify the accuracy of a sequencing platform is basecalling accuracy, as defined by the Phred Quality score (Q score). The score denotes the likelihood that a specific base will be called incorrectly by the sequencer. Q scores are defined as a property that is logarithmically related to the base calling error probabilities (P): Q = -10 log10 P. For example, if Phred assigns a Q score of 30 (Q30) to a base, this is equivalent to the probability of an incorrect base call 1 in 1000 times. This means that the base call accuracy (i.e., the probability of a correct base call) is 99.9%.
[0180] Fig. 10B shows that the “NSR baseline” corresponding to an unlabeled “G’ nucleotide demonstrates the highest Q30 levels, with the Q30 levels remaining high for the other levels of labeled “G” with a slight concomitant decrease in Q30 for increasing levels of labeled “G.” The last example with 100% labeled “G” shows a drop in Q30, which is largely due to the reduced separation of the four clouds that is observable in the last panel of Fig.10A. Systems and methods that label the “G” nucleotide will be able to detect empty clusters for the modest tradeoff in reduced Q30 even for 50% labeling. Amplitude Multiplexing Across Two Channels With Labeled G
[0181] An aspect of the disclosure is directed to alternative scatterplots that may encode multiple nucleotides per signal state across two channels. Such systems and methods may use partially labeled bases. The following example provides a demonstration where four nucleotides may be partially labeled, however, in some embodiments, fewer than four nucleotides may be partially labeled. Fig.11 illustrates two-channel systems and methods that use two different dyes to label each of the nucleotides such that each nucleotide is detected in both channels, and determined by differing intensities in the two channels. If a typical two channel scatterplot is referred to as a square scatterplot, Fig. 11 displays an embodiment with clouds arranged in a diagonal shape on a scatterplot.
[0182] Fig.11 shows an X-Y scatterplot of the fluorescent emissions of a system where four nucleotides are labeled, and empty wells are distinguishable from the four labeled nucleotide. The Y axis corresponds to intensity of fluorescent emissions in a green channel, and the X axis corresponds to the intensity of fluorescent emissions in a blue channel. The four circles in Fig.11 correspond to the fluorescence emissions of nucleotide “G” 1110, nucleotide “T” 1120, nucleotide “C” 1130, nucleotide “A” 1140, and dark or otherwise empty wells 1150. Nucleotide “G” 1110 is shown as a solid green circle emitting only in the green channel. . Nucleotide “T” 1120 andnucleotide “C” 1130 are shown as partially shaded circles with one third labeled blue and two thirds labeled green for Nucleotide “T” 1120, and vice versa for Nucleotide “C” 1130. Nucleotide “A” 1140 is shown as a solid blue circle emitting only in the blue channel. In this embodiment, the nucleotide clouds are shown as a diagonal; however, the nucleotides do not need to be strictly along a straight line. For example, nucleotide “T” 1120 is shown to be one third labeled blue and two thirds labeled green, but could be labeled with a different blue dye, such that it emits with two thirds intensity in the green channel and two thirds intensity in the blue channel. Fig. 11 shows both nucleotide “G” 1110 and nucleotide “A” 1140 as solid circles; however, the nucleotides may be partially labelled.
[0183] Fig. 11 demonstrates an embodiment encompassed by a method of sequencing polynucleotides bound to a flowcell, comprising: detecting fluorescent emissions from a first labeled nucleotide at a first wavelength and a first intensity; detecting fluorescent emissions from a second labeled nucleotide at a second wavelength and a second intensity, wherein the first wavelength is different from the second wavelength; detecting fluorescent emissions from a third labeled nucleotide at the first and second wavelengths at third and fourth intensities; detecting fluorescent emissions from a fourth labeled nucleotide at the first and second wavelengths at fifth and sixth intensities; and determining the sequence of the polynucleotides based on the detected fluorescent emissions and intensities. Specifically, nucleotide “G” 1110 in Fig.11 may correspond to a first labeled nucleotide with emissions at a first wavelength and a first intensity, where the first wavelength is an emission wavelength within a green channel and the first intensity is approximately 100% of a maximum intensity in the green channel (with arbitrary units). A first labeled nucleotide that corresponds nucleotide “G” 1110 in Fig. 11 may be labeled with a single dye, a single dye partially labeling the nucleotide that is also diluted with a non-fluorescent label, and a combination of two fluorescent dyes.
[0184] Nucleotide “A” 1140 in Fig.11 may correspond to a second labeled nucleotide with emissions at a second wavelength and a second intensity, wherein the first wavelength is different from the second wavelength, where the second wavelength is within a blue channel, and the second intensity is approximately 100% of a maximum intensity in the blue channel (with arbitrary units). The emissions of Nucleotide “A” 1140 illustrated in Fig. 11 may be replicated with any of a single dye, a single dye partially labeling the nucleotide that is also diluted with a non-fluorescent label, and a combination of two fluorescent dyes.
[0185] Nucleotide “T” 1120 in Fig. 11may correspond to a third labeled nucleotide with emissions at the first and second wavelengths at third and fourth intensities, wherein the third intensity is approximately 66% of a maximum intensity in the green channel, and the fourth intensity is approximately 33% of a maximum intensity in the blue channel. Nucleotide “C” 1130 in Fig. 11 may correspond to a fourth labeled nucleotide with emissions at the first and second wavelengths at fifth and sixth intensities, wherein the fifth intensity is approximately 33% of a maximum intensity in the green channel, and the sixth intensity is approximately 66% of a maximum intensity in the blue channel. The emissions of Nucleotide “C” or “T” illustrated in Fig. 11 may be each replicated with any of a single dye, a single partially labeled dye that is diluted with a non-fluorescent label, and a combination of two dyes.
[0186] In some embodiments, the techniques described herein relate to a method, wherein the fourth labeled nucleotide is a modified nucleotide. Examples of modified nucleotides are includes throughout the disclosure, including the definition section. However, one skilled in the art will understand that new modified nucleotides beyond those currently discovered may be used as one of the modified bases. Systems
[0187] Embodiments of the present disclosure also include a system for analyzing and assembling sequences of polynucleotides. Fig. 12 is a block diagram of an exemplary computing system 1200 that may be used in connection with an illustrative sequencing system. The computing system 1200 may be configured to determine a DNA sequence by using the sequencing and assembly methods disclosed herein. The general architecture of the computing system 1200 depicted in Fig.1A includes an arrangement of computer hardware and software components. The computing system 1200 may include many more (or fewer) elements than those shown in Fig.1A. It is not necessary, however, that all of these generally conventional elements be shown in order to provide an enabling disclosure.
[0188] As illustrated, the computing system 1200 includes a processing unit 1210, a network interface 1220, a computer-readable medium drive 1230, an input / output device interface 1240, a display 1250, and an input device 1260, all of which may communicate with one another by way of a communication bus. The network interface 1270 may provide connectivity to one or more networks or computing systems. The processing unit 1210 may thus receive information andinstructions from other computing systems or services via a network. The processing unit 1210 may also communicate to and from memory 1270 and further provide output information for an optional display 1250 via the input / output device interface 1240. The input / output device interface 1240 may also accept input from the optional input device 1260, such as a keyboard, mouse, digital pen, microphone, touch screen, gesture recognition system, voice recognition system, gamepad, accelerometer, gyroscope, or other input device.
[0189] The memory 1270 may contain computer program instructions (grouped as modules or components in some embodiments) that the processing unit 1210 executes in order to implement one or more embodiments. The memory 1270 generally includes RAM, ROM and / or other persistent, auxiliary or non-transitory computer-readable media. The memory 1270 may store an operating system 1272 that provides computer program instructions for use by the processing unit 1210 in the general administration and operation of the computing device 1200. The memory 1270 may further include computer program instructions and other information for implementing aspects of the present disclosure.
[0190] For example, in one embodiment, the memory 1270 includes a two-channel sequencing module 1274 for analyzing and assembling sequences of polynucleotides. The two- channel sequencing module 1274 can perform the methods disclosed herein, including the method described with respect to the flow diagrams of Fig. 2. In addition, memory 1270 may include or communicate with the data store 1290 and / or one or more other data stores that store one or more inputs, one or more outputs, and / or one or more results (including intermediate results) of determining a DNA sequence and providing an assembly process according to the present disclosure.
[0191] Particular embodiments of the method of sequencing may utilize a one- excitation, two-channel detection system (also known as 1Ex-2Ch, as defined above) or a two- excitation, two-channel detection system (also known as 2Ex-2Ch, as defined above). Detailed disclosures are provided in WO 2018 / 165099 and U.S. Ser. No. 17 / 338590, each of which is incorporated by reference in its entirety. However, 1Ex-2Ch and 2Ex-2Ch systems and methods are not necessarily considered mutually exclusive and can be used in various combinations. For example, some dyes used in 1Ex-2Ch may be used in 2Ex-2Ch configurations.
[0192] In some embodiments, methods according to the disclosure may be performed on an automated sequencing instrument, and wherein the automatic sequencing instrument maycomprise two light sources operating at different wavelengths (e.g., at 350-360 nm (blue), 520- 530 nm (green), 630 nm-670 nm (red)). The incorporation of the first type of the nucleotide conjugates is determined by a signal state in the first imaging event and a dark state in the second imaging event. The incorporation of the second type of the nucleotide conjugates is determined by a dark state in the first imaging event and a signal state in the second imaging event. The incorporation of the third type of the nucleotide conjugates is determined by a signal state in both the first imaging event and the second imaging event. The incorporation of the fourth type of the nucleotide conjugates is determined by a dark state in the first imaging event and a partial signal state in the second imaging event. The incorporation of the fifth type of the nucleotide conjugates is determined by a dark state in both the first imaging event and the second imaging event.
[0193] In some embodiments, the automatic sequencing instrument may comprise a single light source operating with a blue laser at about 350 nm to about 360 nm. The incorporation of the first type of the nucleotide may be determined by detection in the one of the blue or green channel / region (e.g., at a blue region with a wavelength ranging from about 372 to about 520 nm, or at a green region with a wavelength ranging from about 540 nm to about 640nm). The incorporation of the second type of nucleotide is determined by detection in the other one of the blue or green detection channel / region. The incorporation of the third type of nucleotide is determined by detection in both the blue and green channels / regions. The incorporation of the fourth type of nucleotide is determined by a partial detection in the blue channel / region but no detection green channel / region. The incorporation of the fifth type of nucleotide is determined by no detection in either the blue or detection green channels / regions.
[0194] In some embodiments, the disclosed systems and methods may involve approaches for shifting or distributing certain sequence data analysis features and sequence data storage to a cloud computing environment or cloud-based network. User interaction with sequencing data, genome data, or other types of biological data may be mediated via a central hub that stores and controls access to various interactions with the data. In some embodiments, the cloud computing environment may also provide sharing of protocols, analysis methods, libraries, sequence data as well as distributed processing for sequencing, analysis, and reporting. In some embodiments, the cloud computing environment facilitates modification or annotation of sequence data by users. In some embodiments, the systems and methods may be implemented in a computer browser, on-demand or on-line.
[0195] In some embodiments, software written to perform the methods as described herein is stored in some form of computer readable medium, such as memory, CD-ROM, DVD- ROM, memory stick, flash drive, hard drive, SSD hard drive, server, mainframe storage system and the like.
[0196] In some embodiments, the methods may be written in any of various suitable programming languages, for example compiled languages such as C, C#, C++, Fortran, and Java. Other programming languages could be script languages, such as Perl, MatLab, SAS, SPSS, Python, Ruby, Pascal, Delphi, R and PHP. In some embodiments, the methods are written in C, C#, C++, Fortran, Java, Perl, R, Java or Python. In some embodiments, the method may be an independent application with data input and data display modules. Alternatively, the method may be a computer software product and may include classes wherein distributed objects comprise applications including computational methods as described herein.
[0197] In some embodiments, the methods may be incorporated into pre-existing data analysis software, such as that found on sequencing instruments. Software comprising computer implemented methods as described herein are installed either onto a computer system directly, or are indirectly held on a computer readable medium and loaded as needed onto a computer system. Further, the methods may be located on computers that are remote to where the data is being produced, such as software found on servers and the like that are maintained in another location relative to where the data is being produced, such as that provided by a third party service provider.
[0198] An assay instrument, desktop computer, laptop computer, or server which may contain a processor in operational communication with accessible memory comprising instructions for implementation of systems and methods. In some embodiments, a desktop computer or a laptop computer is in operational communication with one or more computer readable storage media or devices and / or outputting devices. An assay instrument, desktop computer and a laptop computer may operate under a number of different computer based operational languages, such as those utilized by Apple based computer systems or PC based computer systems. An assay instrument, desktop and / or laptop computers and / or server system may further provide a computer interface for creating or modifying experimental definitions and / or conditions, viewing data results and monitoring experimental progress. In some embodiments, an outputting device may be a graphic user interface such as a computer monitor or a computer screen, a printer, a hand-held device suchas a personal digital assistant (i.e., PDA, Blackberry, iPhone), a tablet computer (for example, iPAD), a hard drive, a server, a memory stick, a flash drive and the like.
[0199] A computer readable storage device or medium may be any device such as a server, a mainframe, a supercomputer, a magnetic tape system and the like. In some embodiments, a storage device may be located onsite in a location proximate to the assay instrument, for example adjacent to or in close proximity to, an assay instrument. For example, a storage device may be located in the same room, in the same building, in an adjacent building, on the same floor in a building, on different floors in a building, etc. in relation to the assay instrument. In some embodiments, a storage device may be located off-site, or distal, to the assay instrument. For example, a storage device may be located in a different part of a city, in a different city, in a different state, in a different country, etc. relative to the assay instrument. In embodiments where a storage device is located distal to the assay instrument, communication between the assay instrument and one or more of a desktop, laptop, or server is typically via Internet connection, either wireless or by a network cable through an access point. In some embodiments, a storage device may be maintained and managed by the individual or entity directly associated with an assay instrument, whereas in other embodiments a storage device may be maintained and managed by a third party, typically at a distal location to the individual or entity associated with an assay instrument. In embodiments as described herein, an outputting device may be any device for visualizing data.
[0200] An assay instrument, desktop, laptop and / or server system may be used itself to store and / or retrieve computer implemented software programs incorporating computer code for performing and implementing computational methods as described herein, data for use in the implementation of the computational methods, and the like. One or more of an assay instrument, desktop, laptop and / or server may comprise one or more computer readable storage media for storing and / or retrieving software programs incorporating computer code for performing and implementing computational methods as described herein, data for use in the implementation of the computational methods, and the like. Computer readable storage media may include, but is not limited to, one or more of a hard drive, a SSD hard drive, a CD-ROM drive, a DVD-ROM drive, a floppy disk, a tape, a flash memory stick or card, and the like. Further, a network including the Internet may be the computer readable storage media. In some embodiments, computer readable storage media refers to computational resource storage accessible by a computer network via theInternet or a company network offered by a service provider rather than, for example, from a local desktop or laptop computer at a distal location to the assay instrument.
[0201] In some embodiments, computer readable storage media for storing and / or retrieving computer implemented software programs incorporating computer code for performing and implementing computational methods as described herein, data for use in the implementation of the computational methods, and the like, is operated and maintained by a service provider in operational communication with an assay instrument, desktop, laptop and / or server system via an Internet connection or network connection.
[0202] In some embodiments, a hardware platform for providing a computational environment comprises a processor (i.e., CPU) wherein processor time and memory layout such as random access memory (i.e., RAM) are systems considerations. For example, smaller computer systems offer inexpensive, fast processors and large memory and storage capabilities. In some embodiments, graphics processing units (GPUs) can be used. In some embodiments, hardware platforms for performing computational methods as described herein comprise one or more computer systems with one or more processors. In some embodiments, smaller computer are clustered together to yield a supercomputer network.
[0203] In some embodiments, computational methods as described herein are carried out on a collection of inter- or intra-connected computer systems (i.e., grid technology) which may run a variety of operating systems in a coordinated manner. For example, the CONDOR framework (University of Wisconsin-Madison) and systems available through United Devices are exemplary of the coordination of multiple stand-alone computer systems for the purpose dealing with large amounts of data. These systems may offer Perl interfaces to submit, monitor and manage large sequence analysis jobs on a cluster in serial or parallel configurations. One aspect of the disclosure is directed to a workflow module that may be integrated into existing workflows. In some embodiments, a workflow module may be a two-channel sequencing module and may be integrated into a NGS sequence analysis platform, for example the DRAGEN™ Bio-ID platform from Illumina. Samples
[0204] In some embodiments, the sample comprises or consists of a purified or isolated polynucleotide derived from a tissue sample, a biological fluid sample, a cell sample, and the like. Suitable biological fluid samples include, but are not limited to blood, plasma, serum, sweat, tears,sputum, urine, sputum, ear flow, lymph, saliva, cerebrospinal fluid, ravages, bone marrow suspension, vaginal flow, trans-cervical lavage, brain fluid, ascites, milk, secretions of the respiratory, intestinal and genitourinary tracts, amniotic fluid, milk, and leukophoresis samples. In some embodiments, the sample is a sample that is easily obtainable by non-invasive procedures, e.g., blood, plasma, serum, sweat, tears, sputum, urine, sputum, ear flow, saliva or feces. In certain embodiments the sample is a peripheral blood sample, or the plasma and / or serum fractions of a peripheral blood sample. In other embodiments, the biological sample is a swab or smear, a biopsy specimen, or a cell culture. In another embodiment, the sample is a mixture of two or more biological samples, e.g., a biological sample can comprise two or more of a biological fluid sample, a tissue sample, and a cell culture sample. As used herein, the terms “blood,” “plasma” and “serum” expressly encompass fractions or processed portions thereof. Similarly, where a sample is taken from a biopsy, swab, smear, etc., the “sample” expressly encompasses a processed fraction or portion derived from the biopsy, swab, smear, etc.
[0205] In certain embodiments, samples can be obtained from sources, including, but not limited to, samples from different individuals, samples from different developmental stages of the same or different individuals, samples from different diseased individuals (e.g., individuals with cancer or suspected of having a genetic disorder), normal individuals, samples obtained at different stages of a disease in an individual, samples obtained from an individual subjected to different treatments for a disease, samples from individuals subjected to different environmental factors, samples from individuals with predisposition to a pathology, samples individuals with exposure to an infectious disease agent, and the like.
[0206] In one illustrative, but non-limiting embodiment, the sample is a maternal sample that is obtained from a pregnant female, for example a pregnant woman. The maternal sample can be a tissue sample, a biological fluid sample, or a cell sample. In another illustrative, but non-limiting embodiment, the maternal sample is a mixture of two or more biological samples, e.g., the biological sample can comprise two or more of a biological fluid sample, a tissue sample, and a cell culture sample.
[0207] In certain embodiments samples can also be obtained from in vitro cultured tissues, cells, or other polynucleotide-containing sources. The cultured samples can be taken from sources including, but not limited to, cultures (e.g., tissue or cells) maintained in different media and conditions (e.g., pH, pressure, or temperature), cultures (e.g., tissue or cells) maintained fordifferent periods of length, cultures (e.g., tissue or cells) treated with different factors or reagents (e.g., a drug candidate, or a modulator), or cultures of different types of tissue and / or cells.
[0208] In some embodiments, the use of the disclosed sequencing technology does not involve the preparation of sequencing libraries. In other embodiments, the sequencing technology contemplated herein involve the preparation of sequencing libraries. In one illustrative approach, sequencing library preparation involves the production of a random collection of adapter-modified DNA fragments (e.g., polynucleotides) that are ready to be sequenced.
[0209] Sequencing libraries of polynucleotides can be prepared from DNA or RNA, including equivalents, analogs of either DNA or cDNA, for example, DNA or cDNA that is complementary or copy DNA produced from an RNA template, by the action of reverse transcriptase. The polynucleotides may originate in double-stranded form (e.g., dsDNA such as genomic DNA fragments, cDNA, PCR amplification products, and the like) or, in certain embodiments, the polynucleotides may originated in single-stranded form (e.g., ssDNA, RNA, etc.) and have been converted to dsDNA form. By way of illustration, in certain embodiments, single stranded mRNA molecules may be copied into double-stranded cDNAs suitable for use in preparing a sequencing library. The precise sequence of the primary polynucleotide molecules is generally not material to the method of library preparation, and may be known or unknown. In one embodiment, the polynucleotide molecules are DNA molecules. More particularly, in certain embodiments, the polynucleotide molecules represent the entire genetic complement of an organism or substantially the entire genetic complement of an organism, and are genomic DNA molecules (e.g., cellular DNA, cell free DNA (cfDNA), etc.), that typically include both intron sequence and exon sequence (coding sequence), as well as non-coding regulatory sequences such as promoter and enhancer sequences. In certain embodiments, the primary polynucleotide molecules comprise human genomic DNA molecules, e.g., cfDNA molecules present in peripheral blood of a pregnant subject.
[0210] Methods of isolating nucleic acids from biological sources may differ depending upon the nature of the source. One of skill in the art can readily isolate nucleic acids from a source as needed for the method described herein. In some instances, it can be advantageous to fragment large nucleic acid molecules (e.g. cellular genomic DNA) in the nucleic acid sample to obtain polynucleotides in the desired size range. Fragmentation can be random, or it can be specific, as achieved, for example, using restriction endonuclease digestion. Methods for randomfragmentation may include, for example, limited DNase digestion, alkali treatment and physical shearing. Fragmentation can also be achieved by any of a number of methods known to those of skill in the art. For example, fragmentation can be achieved by mechanical means including, but not limited to nebulization, sonication and hydroshear.
[0211] In some embodiments, sample nucleic acids are obtained from as cfDNA, which is not subjected to fragmentation. For example, cfDNA, typically exists as fragments of less than about 300 base pairs and consequently, fragmentation is not typically necessary for generating a sequencing library using cfDNA samples.
[0212] Typically, whether polynucleotides are forcibly fragmented (e.g., fragmented in vitro), or naturally exist as fragments, they are converted to blunt-ended DNA having 5’- phosphates and 3’-hydroxyl. Standard protocols, e.g., protocols for sequencing using, for example, the Illumina platform, instruct users to end-repair sample DNA, to purify the end-repaired products prior to dA-tailing, and to purify the dA-tailing products prior to the adaptor-ligating steps of the library preparation.
[0213] In various embodiments, verification of the integrity of the samples and sample tracking can be accomplished by sequencing mixtures of sample genomic nucleic acids, e.g., cfDNA, and accompanying marker nucleic acids that have been introduced into the samples, e.g., prior to processing. Sequencing Techniques
[0214] The disclosed sequencing systems and methods may be compatible with any sequencing techniques based on optical detection, for example, next-generation sequencing (NGS), fluorescent in situ sequencing (FISSEQ), and Massively Parallel Signature Sequencing (MPSS). In one embodiment, the disclosed systems and methods may be compatible with NGS technologies that allow multiple samples to be sequenced individually as genomic molecules (i.e., singleplex sequencing) or as pooled samples comprising indexed genomic molecules (e.g., multiplex sequencing) on a single sequencing run. These methods can generate up to several hundred million reads of DNA sequences.
[0215] The disclosed technology may implement sequencing reactions such as those incorporating sequencing-by-synthesis methods described in U.S. Patent Application Publication Numbers 2007 / 0166705, 2006 / 0188901, 2006 / 0240439, 2006 / 0281109, 2005 / 0100900, U.S.Patent Number 7,057,026, PCT Application Publication Numbers WO 2005 / 065814, WO 2006 / 064199, and WO 2007 / 010251, the disclosures of which are incorporated herein by reference in their entireties. In some embodiments, the sequencers may implement sequencing-by-synthesis methods similar to those used in the HiSeq, MiSeq, or HiScanSQ systems from Illumina (San Diego, Calif.).
[0216] Alternatively, sequencing by ligation techniques may be used in the disclosed technology, such as described in U.S. Patent Numbers 6,969,488, 6,172,218, and 6,306,597, the disclosures of which are incorporated herein by reference in their entireties. Sequencing by ligation techniques use DNA ligase to incorporate oligonucleotides and identify the incorporation of such oligonucleotides.
[0217] The disclosed technology may be implemented in some sequencing techniques which are available commercially, such as the sequencing-by-hybridization platform from Affymetrix Inc. (Sunnyvale, CA) and the sequencing-by-synthesis platforms from 454 Life Sciences (Bradford, CT) and Helicos Biosciences (Cambridge, MA), the sequencing-by-ligation platform from Applied Biosystems (Foster City, CA), or the SMRT technology of Pacific Biosciences.
[0218] In one illustrative, but non-limiting, embodiment, the methods described herein comprise obtaining sequence information for the nucleic acids in a sample using Illumina’s sequencing-by-synthesis and reversible terminator-based sequencing chemistry (e.g. as described in Bentley et al., Nature 6:53-59
[2009] ). Illumina’s sequencing technology may include the attachment of fragmented genomic DNA to a planar, optically transparent surface on which oligonucleotide anchors are bound. For example, template DNA is end-repaired to generate 5’- phosphorylated blunt ends, and the polymerase activity of Klenow fragment is used to add a single A base to the 3’ end of the blunt phosphorylated DNA fragments. This addition prepares the DNA fragments for ligation to oligonucleotide adapters, which have an overhang of a single T base at their 3’ end to increase ligation efficiency. The adapter oligonucleotides are complementary to the flowcell anchor oligos. Under limiting-dilution conditions, adapter-modified, single-stranded template DNA is added to the flowcell and immobilized by hybridization to the anchor oligos. Attached DNA fragments are extended and bridge amplified to create an ultra-high density sequencing flowcell with hundreds of millions of clusters, each containing about 1,000 copies of the same template. In one embodiment, the randomly fragmented genomic DNA is amplified usingPCR before it is subjected to cluster amplification. Alternatively, an amplification-free (e.g., PCR free) genomic library preparation is used, and the randomly fragmented genomic DNA is enriched using the cluster amplification alone (Kozarewa et al., Nature Methods 6:291-295
[2009] ). The sequencing-by-synthesis reaction may employ reversible terminators with removable fluorescent dyes. Short sequence reads of about tens to a few hundred base pairs are aligned against a reference genome and unique mapping of the short sequence reads to the reference genome are identified. After completion of the first read, the templates can be regenerated in situ to enable a second read from the opposite end of the fragments. Thus, either single-end or paired end sequencing of the DNA fragments can be used. Detailed information about paired end sequencing can be found in US Patent No. 7601499 and US Patent Publication No. 2012 / 0,053,063, which are incorporated by reference.
[0219] In some embodiments, the sequencing by synthesis platform by Illumina involves clustering fragments. Clustering is a process in which each fragment molecule is isothermally amplified. In some embodiments, the fragment has two different adaptors attached to the two ends of the fragment, the adaptors allowing the fragment to hybridize with the two different oligos on the surface of a flowcell lane. The fragment further includes or is connected to two index sequences at two ends of the fragment, where index sequences provide labels to identify different samples in multiplex sequencing.
[0220] In some implementation, a flowcell for clustering in the Illumina platform is a glass slide with lanes. Each lane is a glass channel coated with a lawn of two types of oligos. Hybridization is enabled by the first of the two types of oligos on the surface. This oligo is complementary to a first adapter on one end of the fragment. A polymerase creates a compliment strand of the hybridized fragment. The double-stranded molecule is denatured, and the original template strand is washed away. The remaining strand, in parallel with many other remaining strands, is clonally amplified through bridge application.
[0221] In bridge amplification, a strand folds over, and a second adapter region on a second end of the strand hybridizes with the second type of oligos on the flowcell surface. A polymerase generates a complimentary strand, forming a double-stranded bridge molecule. This double-stranded molecule is denatured resulting in two single-stranded molecules tethered to the flowcell through two different oligos. The process is then repeated over and over, and occurs simultaneously for millions of clusters resulting in clonal amplification of all the fragments. Afterbridge amplification, the reverse strands are cleaved and washed off, leaving only the forward strands. The 3’ ends are blocked to prevent unwanted priming.
[0222] After clustering, sequencing starts with extending a first sequencing primer to generate the first read. With each cycle, fluorescently tagged nucleotides compete for addition to the growing chain. Only one is incorporated based on the sequence of the template. After the addition of each nucleotide, the cluster is excited by a light source, and a characteristic fluorescent signal is emitted. The number of cycles determines the length of the read. The emission wavelength and the signal intensity determine the base call. For a given cluster all identical strands are read simultaneously. Hundreds of millions of clusters, or thousands to tens of thousands of millions of clusters, are sequenced in a massively parallel manner. At the completion of the first read, the read product is washed away.
[0223] In processes involving two index primers, an index 1 primer is introduced and hybridized to an index 1 region on the template. Index regions provide identification of fragments, which is useful for de-multiplexing samples in a multiplex sequencing process. The index 1 read is generated similar to the first read. After completion of the index 1 read, the read product is washed away and the 3’ end of the strand is de-protected. The template strand then folds over and binds to a second oligo on the flowcell. An index 2 sequence is read in the same manner as index 1. Then an index 2 read product is washed off at the completion of the step.
[0224] After reading two indices, read 2 initiates by using polymerases to extend the second flowcell oligos, forming a double-stranded bridge. This double-stranded DNA is denatured, and the 3’ end is blocked. The original forward strand is cleaved off and washed away, leaving the reverse strand. Read 2 begins with the introduction of a read 2 sequencing primer. As with read 1, the sequencing steps are repeated until the desired length is achieved. The read 2 product is washed away. This entire process generates millions of reads, representing all the fragments. Sequences from pooled sample libraries are separated based on the unique indices introduced during sample preparation. For each sample, reads of similar stretches of base calls are locally clustered. Forward and reversed reads are paired creating contiguous sequences. These contiguous sequences are aligned to the reference genome for variant identification.Systems and Instruments
[0225] In some embodiments, the methods may be written in any of various suitable programming languages, for example compiled languages such as C, C#, C++, Fortran, and Java. Other programming languages could be script languages, such as Perl, MatLab, SAS, SPSS, Python, Ruby, Pascal, Delphi, R and PHP. In some embodiments, the methods are written in C, C#, C++, Fortran, Java, Perl, R, Java or Python. In some embodiments, the method may be an independent application with data input and data display modules. Alternatively, the method may be a computer software product and may include classes wherein distributed objects comprise applications including computational methods as described herein.
[0226] In some embodiments, the methods may be incorporated into pre-existing data analysis software, such as that found on sequencing instruments. Software comprising computer implemented methods as described herein are installed either onto a computer system directly, or are indirectly held on a computer readable medium and loaded as needed onto a computer system. Further, the methods may be located on computers that are remote to where the data is being produced, such as software found on servers and the like that are maintained in another location relative to where the data is being produced, such as that provided by a third party service provider.
[0227] An assay instrument, desktop computer, laptop computer, or server which may contain a processor in operational communication with accessible memory comprising instructions for implementation of systems and methods. In some embodiments, a desktop computer or a laptop computer is in operational communication with one or more computer readable storage media or devices and / or outputting devices. An assay instrument, desktop computer and a laptop computer may operate under a number of different computer based operational languages, such as those utilized by Apple based computer systems or PC based computer systems. An assay instrument, desktop and / or laptop computers and / or server system may further provide a computer interface for creating or modifying experimental definitions and / or conditions, viewing data results and monitoring experimental progress. In some embodiments, an outputting device may be a graphic user interface such as a computer monitor or a computer screen, a printer, a hand-held device such as a personal digital assistant (i.e., PDA, Blackberry, iPhone), a tablet computer (for example, iPAD), a hard drive, a server, a memory stick, a flash drive and the like.
[0228] A computer readable storage device or medium may be any device such as a server, a mainframe, a supercomputer, a magnetic tape system and the like. In some embodiments,a storage device may be located onsite in a location proximate to the assay instrument, for example adjacent to or in close proximity to, an assay instrument. For example, a storage device may be located in the same room, in the same building, in an adjacent building, on the same floor in a building, on different floors in a building, etc. in relation to the assay instrument. In some embodiments, a storage device may be located off-site, or distal, to the assay instrument. For example, a storage device may be located in a different part of a city, in a different city, in a different state, in a different country, etc. relative to the assay instrument. In embodiments where a storage device is located distal to the assay instrument, communication between the assay instrument and one or more of a desktop, laptop, or server is typically via Internet connection, either wireless or by a network cable through an access point. In some embodiments, a storage device may be maintained and managed by the individual or entity directly associated with an assay instrument, whereas in other embodiments a storage device may be maintained and managed by a third party, typically at a distal location to the individual or entity associated with an assay instrument. In embodiments as described herein, an outputting device may be any device for visualizing data.
[0229] An assay instrument, desktop, laptop and / or server system may be used itself to store and / or retrieve computer implemented software programs incorporating computer code for performing and implementing computational methods as described herein, data for use in the implementation of the computational methods, and the like. One or more of an assay instrument, desktop, laptop and / or server may comprise one or more computer readable storage media for storing and / or retrieving software programs incorporating computer code for performing and implementing computational methods as described herein, data for use in the implementation of the computational methods, and the like. Computer readable storage media may include, but is not limited to, one or more of a hard drive, a SSD hard drive, a CD-ROM drive, a DVD-ROM drive, a floppy disk, a tape, a flash memory stick or card, and the like. Further, a network including the Internet may be the computer readable storage media. In some embodiments, computer readable storage media refers to computational resource storage accessible by a computer network via the Internet or a company network offered by a service provider rather than, for example, from a local desktop or laptop computer at a distal location to the assay instrument.
[0230] In some embodiments, computer readable storage media for storing and / or retrieving computer implemented software programs incorporating computer code for performingand implementing computational methods as described herein, data for use in the implementation of the computational methods, and the like, is operated and maintained by a service provider in operational communication with an assay instrument, desktop, laptop and / or server system via an Internet connection or network connection.
[0231] In some embodiments, a hardware platform for providing a computational environment comprises a processor (i.e., CPU) wherein processor time and memory layout such as random access memory (i.e., RAM) are systems considerations. For example, smaller computer systems offer inexpensive, fast processors and large memory and storage capabilities. In some embodiments, graphics processing units (GPUs) can be used. In some embodiments, hardware platforms for performing computational methods as described herein comprise one or more computer systems with one or more processors. In some embodiments, smaller computer are clustered together to yield a supercomputer network.
[0232] In some embodiments, computational methods as described herein are carried out on a collection of inter- or intra-connected computer systems (i.e., grid technology) which may run a variety of operating systems in a coordinated manner. For example, the CONDOR framework (University of Wisconsin-Madison) and systems available through United Devices are exemplary of the coordination of multiple stand-alone computer systems for the purpose dealing with large amounts of data. These systems may offer Perl interfaces to submit, monitor and manage large sequence analysis jobs on a cluster in serial or parallel configurations. One aspect of the disclosure is directed to a workflow module that may be integrated into existing workflows. In some embodiments, a workflow module may be a two-channel sequencing module and may be integrated into a NGS sequence analysis platform, for example the DRAGEN™ Bio-ID platform from Illumina.
Claims
WHAT IS CLAIMED IS:
1. A method of sequencing polynucleotides bound to a flowcell, comprising: detecting fluorescent emissions from a first labeled nucleotide at a first wavelength and a first intensity; detecting fluorescent emissions from a second labeled nucleotide at a second wavelength and a second intensity, wherein the first wavelength is different from the second wavelength; detecting fluorescent emissions from a third labeled nucleotide at the first and second wavelengths at third and fourth intensities; detecting fluorescent emissions from a fourth labeled nucleotide at the first and second wavelengths at fifth and sixth intensities; and determining the sequence of the polynucleotides based on the detected fluorescent emissions and intensities.
2. The method of claim 1, wherein the first labeled nucleotide and the second labeled nucleotide are labeled with fluorescent dyes having the same excitation wavelength.
3. The method of claim 1 or 2, wherein the first labeled nucleotide and the second labeled nucleotide are labeled with fluorescent dyes having different stokes shifts.
4. The method of claim 1, wherein at least two of the first through sixth intensities are greater than zero.
5. The method of claim 1, wherein at least three of the first through sixth intensities are greater than zero.
6. The method of claim 1, wherein at least four of the first through sixth intensities are greater than zero.
7. The method of claim 1, wherein at least five of the first through sixth intensities are greater than zero.
8. The method of claim 1, wherein all of the first through sixth intensities are greater than zero.
9. The method of claim 8, wherein the ratio of the third intensity to the fourth intensity is approximately equal to the ratio of the sixth intensity to the fifth intensity.
10. The method of claim 9, wherein the first intensity is greater than the third intensity and the fifth intensity; further wherein the second intensity is greater than the fourth intensity and the sixth intensity.
11. The method of claim 9, wherein the first intensity is less than the third intensity; further wherein the second intensity is less than the sixth intensity; and wherein four labeled nucleotides form a diamond configuration.
12. The method of claim 1, further comprising determining the presence of an empty well based on a substantial absence of the detected fluorescent emissions and intensities.
13. The method of claim 1, further comprising detecting fluorescent emissions from a fifth labeled nucleotide at the first and second wavelengths at seventh and eighth intensities.
14. A kit for determining the sequence of a polynucleotide, comprising: a first mixture of a first nucleotide-first fluorescent dye conjugate detectable in a first wavelength channel and a first nucleotide-second fluorescent dye conjugate detectable in a second wavelength channel; a second mixture of a second nucleotide-first fluorescent dye conjugate detectable in the first wavelength channel and a second nucleotide-second fluorescent dye conjugate detectable in the second wavelength channel; a third mixture of a third nucleotide-first fluorescent dye conjugate detectable in the first wavelength channel and a third nucleotide-second fluorescent dye conjugate detectable in the second wavelength channel; and a fourth mixture of a fourth nucleotide-first fluorescent dye conjugate detectable in the first wavelength channel and a fourth nucleotide-second fluorescent dye conjugate detectable in the second wavelength channel.
15. The kit of claim 14, wherein the ratio of intensity of emissions in the first channel versus the second channel for the third mixture is approximately equal to the ratio of intensity of emissions in the second channel versus the first channel for the fourth mixture.
16. A system for determining the sequence of a polynucleotide, comprising: a machine-readable memory; and a processor configured to execute machine-readable instructions, which, when executed by the processor, cause the system to perform steps including:detecting fluorescent emissions from a first labeled nucleotide at a first wavelength and a first intensity; detecting fluorescent emissions from a second labeled nucleotide at a second wavelength and a second intensity, wherein the first wavelength is different from the second wavelength; detecting fluorescent emissions from a third labeled nucleotide at the first and second wavelengths at third and fourth intensities; detecting fluorescent emissions from a fourth labeled nucleotide at the first and second wavelengths at fifth and sixth intensities; and determining the sequence of the polynucleotides based on the detected fluorescent emissions and intensities.
17. The system of claim 16, further comprising: a first mixture of a first nucleotide-first fluorescent dye conjugate detectable in a first wavelength channel and a first nucleotide-second fluorescent dye conjugate detectable in a second wavelength channel; a second mixture of a second nucleotide-first fluorescent dye conjugate detectable in the first wavelength channel and a second nucleotide-second fluorescent dye conjugate detectable in the second wavelength channel; and a third mixture of a third nucleotide-first fluorescent dye conjugate detectable in the first wavelength channel and a third nucleotide-second fluorescent dye conjugate detectable in the second wavelength channel.
18. The system of claim 16, further comprising, a fourth mixture of a fourth nucleotide- first fluorescent dye conjugate detectable in the first wavelength channel and a fourth nucleotide- second fluorescent dye conjugate detectable in the second wavelength channel.
19. The system of claim 18, further comprising, a fifth mixture of a fifth nucleotide-first fluorescent dye conjugate detectable in the first wavelength channel and a fifth nucleotide-second fluorescent dye conjugate detectable in the second wavelength channel.
20. The system of claim 19, wherein all of the mixtures have different ratios of first fluorescent dye conjugate to second fluorescent dye conjugate.
21. The system of any of claims 16 to 20, wherein all of the mixtures have different amounts of the first fluorescent dye conjugate.
22. The system of any of claims 16 to 21, wherein all of the mixtures have different amounts of the second fluorescent dye conjugate.
23. The system of claim 16, comprising the additional step of determining the presence of an empty well from a dark state.