Amplitude modulation for accelerated base calling
Patent Information
- Application Number
- JP2023575840
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-09-01
- Filing Date
- 2022-08-23
- Publication Date
- 2025-08-14
AI Technical Summary
Existing DNA sequencing systems require multiple excitation light sources and detectors, which increase sequencing time and expose samples to unnecessary light, leading to DNA damage and reduced sequencing speed.
A nucleic acid sequencing system utilizing a single light source and a single detector, employing amplitude modulation to encode nucleobases with different fluorescence intensities, allowing for faster sequencing and reduced light exposure.
This approach doubles the information per imaging cycle, halves the required imaging time, reduces equipment complexity and cost, and minimizes light-induced DNA damage, while maintaining accurate base calling.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Background technology]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Patent Application No. 63 / 260,838, filed September 1, 2021, the contents of which are incorporated herein by reference in their entirety.
[0002] In some types of next-generation sequencing technology, DNA clusters are created on a flow cell after amplification of target polynucleotide. Existing DNA sequencing systems and methods, such as existing sequencing platforms using two-channel or four-channel sequencing chemistry, can utilize two or more excitation light sources to excite deoxyribonucleic acid analogs conjugated with fluorescent labels in target polynucleotides and can utilize two or more detectors to capture fluorescent images in two or more optical channels. Reducing the number of excitation light sources and / or the number of detectors can increase the sequencing speed of such systems. In addition, reducing the number of excitation light sources can reduce unnecessary exposure of samples to light, thus reducing light-induced DNA damage. Summary of the Invention [Means for solving the problem]
[0003] The disclosed technology relates to the field of nucleic acid sequencing, and more particularly to systems and methods for nucleic acid sequencing that utilize a single light source and a single detector.
[0004] In some aspects, the disclosed technology relates to a system and method for identifying nucleobases in a polynucleotide bound to a substrate. The disclosed system may include a first detector configured to detect the intensity of light within a first detection wavelength range. The disclosed system may further include a first light source configured to output light at a first excitation wavelength. The disclosed system may further include a processor configured to control the first light source to generate light at the first excitation wavelength to stimulate emission from the polynucleotide bound to the substrate, and identify the nucleobases in the polynucleotide based on the intensity of the emission received by the first detector. For example, the first nucleobase is identified based on receiving a full intensity emission by the first detector, and the second nucleobase is identified by receiving an emission that is less than the full intensity emission. In some embodiments, at least four types of nucleobases can be identified from the image captured by the first detector. In various embodiments, the disclosed systems may utilize four types of nucleotide analogs that emit light at four distinguishable levels when excited by a light source, and the four types of nucleotide analogs may be bound to different fluorophores or to the same fluorophore with different probabilities or copy numbers.
[0005] In some embodiments, the disclosed system may further include a second detector configured to detect light within a second detection wavelength range and / or a second light source configured to output light at a second excitation wavelength. The processor may be further configured to control the second light source to generate light at the second excitation wavelength to stimulate emission from the polynucleotide and / or identify the nucleobase in the polynucleotide based on the intensity of the emission received by the first detector and the second detector. In some embodiments, the processor is further configured to determine the quality of identifying the nucleobase and / or the error rate of identifying the nucleobase and / or the signal-to-noise ratio of the emission from the plurality of polynucleotides bound to the substrate. In response to the determined quality and / or the determined error rate and / or the determined signal-to-noise ratio, the processor may activate the second detector, the second light source, or both, and / or switch from a first mode of identifying the nucleobase based on the intensity of the emission received by the first detector to a second mode of identifying the nucleobase based on the intensity of the emission received by the first detector and the second detector. In some embodiments, the processor may activate a second detector, a second light source, or both, after a predetermined number of cycles of identifying a nucleobase in the polynucleotide, and / or may switch from a first mode of identifying a nucleobase based on the intensity of luminescence received by the first detector to a second mode of identifying a nucleobase based on the intensity of luminescence received by the first detector and the second detector. In some embodiments, the processor is further configured to control the fluidic device to deliver an alternative set of nucleotide analogs (e.g., as part of an alternative sequencing reaction mix) to the polynucleotide based on a determined quality of identifying the nucleobase, based on a determined error rate of identifying the nucleobase, based on a determined signal-to-noise ratio, or after a predetermined number of cycles of identifying the nucleobase.
[0006] The systems, devices, kits, and methods disclosed herein each have several aspects, no one of which is solely responsible for their desirable attributes. Numerous other embodiments are contemplated, including embodiments having fewer, additional, and / or different components, steps, features, objects, benefits, and advantages. The components, aspects, and steps may be arranged and ordered differently. After considering this discussion, and particularly after reading the section entitled "Detailed Description of the Invention," one will appreciate how the features of the devices and methods disclosed herein provide advantages over other known devices and methods.
[0007] It is understood that any features of the systems disclosed herein can be combined in any desired manner and / or configuration. It is further understood that any features of the methods disclosed herein can be combined in any desired manner. It is further understood that any combination of method and / or system features can be used together and / or combined with any of the embodiments disclosed herein. It is understood that all combinations of the foregoing concepts and additional concepts discussed in more detail below are considered to be part of the inventive subject matter disclosed herein and can be used to achieve the benefits and advantages described herein. [Brief description of the drawings]
[0008] Features of examples of the present disclosure will become apparent by reference to the following detailed description and the drawings in which like reference numbers correspond to similar, though perhaps not identical, components. For purposes of brevity, reference numbers or features having previously described functions may or may not be described in conjunction with the other drawings in which they appear.
[0009] [Figure 1A] 1 illustrates generally an exemplary sequencing system in which embodiments of the disclosed sequencing techniques can be implemented.
[0010] [Figure 1B] 1 illustrates generally an exemplary imaging system for use in embodiments of the disclosed sequencing techniques.
[0011] [Figure 1C] 1 illustrates generally another exemplary imaging system for use in embodiments of the disclosed sequencing techniques.
[0012] [Diagram 2] FIG. 1B shows a functional block diagram of an exemplary computer system used in the sequencing system shown in FIG. 1A.
[0013] [Diagram 3] 1 shows an exemplary two-dye labeling scheme according to an embodiment of the disclosed sequencing technology.
[0014] [Figure 4A] 1 is a line graph showing expected fluorescence detection results from a prophetic experiment performed by one embodiment of the disclosed sequencing technology.
[0015] [Figure 4B] FIG. 1 is a scatter plot showing predicted fluorescence detection clouds obtained from prophetic experiments performed by one embodiment of the disclosed sequencing technology.
[0016] [Diagram 5] 1 shows exemplary emission spectra of a collection of nucleotide analogs used in a dye labeling scheme according to one embodiment of the disclosed technology.
[0017] [Figure 6] 1 is a line graph illustrating the signal-to-noise ratio requirements in support of the sequencing technique of the present disclosure compared to prior art two-channel sequencing techniques.
[0018] [Figure 7] FIG. 1 is a flow diagram illustrating one embodiment of a method for sequencing a polynucleotide according to one embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0019] All patents, patent applications, and other publications, including all sequences disclosed therein and referred to herein, are expressly incorporated herein by reference to the same extent as if each publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. All cited documents are, in relevant part, incorporated herein by reference in their entirety for the purposes indicated by the context of the citation herein. However, the citation of any document should not be construed as an admission that it is prior art to the present disclosure.
[0020] Introduction An embodiment of the disclosed technology relates to a next-generation sequencing system and method that can identify four nucleotide bases using a single excitation light source and a single optical channel for detection. The disclosed sequencing technology can utilize a sequencing-by-synthesis process. During each sequencing cycle, four types of nucleotide analogs can be incorporated onto a growing primer that hybridizes to a polynucleotide to be sequenced. The four types of nucleotide analogs can be attached to different fluorophores or to fluorophores of the same type but with different fluorescence emission intensities. The same type of fluorophore can also be attached to the four types of nucleotide analogs, but to analogs at different percentages of the total nucleotide population in the reaction mix. For example, for A, only 25% of the A nucleotides in the population of A nucleotides are labeled with a fluorophore. For G, only 50% of the G nucleotides in the population of G nucleotides are labeled with a fluorophore. For example, for C, 75% of the C nucleotides in the population of G nucleotides are labeled with a fluorophore. For T, 100% of the T nucleotides are labeled with a fluorophore. In each of the above scenarios, the four types of nucleotide analogs can emit light at four distinct emission levels or intensities when excited by a light source. The nucleobases in a polynucleotide can then be identified based on the distinct levels or intensities of fluorescent emission.
[0021] The disclosed technology may be viewed as an encoding mechanism that uses amplitude modulation of emission wavelength in a "one-color" system, where each base is encoded as having a different intensity in a single optical channel. For example, G may be encoded as OFF, T may be encoded as 1 / 3 amplitude, C may be encoded as 2 / 3 amplitude, and A may be encoded as full amplitude. In some embodiments, a single-color, single-image sequencing approach may be implemented by using different dyes with different fluorescent intensities but the same emission wavelength for different base types. In some embodiments, a single-color, single-image sequencing approach may be implemented by varying the ratio of fluorescently tagged nucleotide analogs to non-tagged nucleotide analogs for each nucleotide base type. Improvements in the real-time sequence analysis process, such as adaptive equalization, adaptive updating, time tracking, offset removal estimation, gain estimation, channel estimation per cluster, and fading correction per cluster, may improve the signal-to-noise ratio of each DNA cluster and allow more accurate tracking of DNA cluster amplitudes, thus facilitating the implementation of this "one-color" encoding scheme. It should be appreciated that, particularly during the initial cycles of reading, the SNR is relatively high, since fading and other issues have not yet begun to degrade the SNR of the system. The approach described herein may also be particularly applicable in situations where the read length is generally relatively short, such as RNA reads. Ensuring a sufficiently high signal-to-noise ratio may allow the "one-color" system to control or maintain an acceptable basecalling error rate during long sequence read applications.
[0022] In a sense, base calling can be interpreted as a communication system that uses fluorescence intensity modulation to emit different intensity levels for each labeled nucleotide and then detects those intensity levels directly. Base calling essentially detects a signal corrupted by noise, where the signal intensity is modulated by the amount of fluorescence emission reflecting the base, and the noise can result from intensity fluctuations, such as laser intensity fluctuations, fading, shot noise, cluster intensity fluctuations, etc. The raw intensity measurements in a given cycle can be used to represent a symbol. When the signaling rate is limited and the signal-to-noise ratio is high, the communication system can use advanced modulation techniques that trade signal-to-noise ratio for throughput, thus reducing the symbol rate. The disclosed "one color" (or sometimes referred to as "Pulse Amplitude Modulation 4" or "PAM4") system can encode more bits per symbol and transmit more information in each cycle using a higher signal-to-noise ratio. In this sense, four amplitude levels are encoded per symbol, and thus a symbol is equivalent to two bits. The disclosed "one-color" system can double the information per imaging cycle, thus halving the number of images obtained and halving the imaging time required. Primary analysis computational requirements can also be reduced by up to 50% using this system. In some embodiments, the disclosed "one-color" system can be implemented by modifying the decision logic in an existing real-time sequence analysis process, while the equalization, phasing, and parameter estimation processes can remain the same and no extra computational or storage costs can be required. Furthermore, this approach can reduce the amount of time a particular sample is exposed to the laser, thereby reducing laser damage and increasing potential read lengths.
[0023] The disclosed "one-color" sequencing technology adds a new design dimension that can be used to trade signal-to-noise ratio against throughput and complexity / cost of instrument design. For example, the disclosed technology can trade signal-to-noise ratio against imaging time. The disclosed technology can halve the imaging time for certain applications, or double the throughput at the same imaging speed. The disclosed technology can reduce the instrument cost of the sold product and / or system intensity costs. The disclosed technology can be well suited for cost-sensitive short-read applications. The disclosed technology can obtain signal-to-noise ratio gains from improvements in real-time analysis processes (e.g., equalization or cluster-wise parameter estimation processes). The disclosed technology can allow for reduced imaging times, thus making more time available for sequencing chemistry to act on the polynucleotide being sequenced. The disclosed technology can allow for better instrument intensity control, simplified optics, and higher signal intensity. The disclosed technology can apply knowledge of the noise shape to better identify bases, for example, using Poisson-shaped counting noise to adjust signaling levels accordingly. The disclosed technology can be expanded to encode five or six different nucleic acid bases as different levels of signal. The disclosed "one-color" sequencing technology allows base calling from a single image obtained with a single optical channel in each sequencing cycle. The disclosed sequencing technology allows for reduced average signal intensity, reduced signal-to-noise ratio, simplified laser requirements, simplified optical filter set, fewer components and lower cost of sequencing equipment, reduced sequencing time, reduced data storage requirements and memory size, and fewer images to be processed, resulting in reduced computational requirements.
[0024] Example Sequencer 1A illustrates an exemplary sequencing system 100 that can implement the disclosed sequencing techniques. The sequencing system 100 can be configured to utilize the disclosed sequencing methods based on a single optical excitation and single detection channel. Non-limiting examples of sequencing reactions utilized can include variations of sequencing-by-synthesis processes such as those used in Illumina® dye sequencing or HeliScope® single molecule sequencing.
[0025] The sequencing system 100 may include an optical system 102 configured to generate raw sequencing data using sequencing reagents provided by a fluidic system 104 that is part of the sequencing system 100. The raw sequencing data may include fluorescent images captured by the optical system 102. The sequencing system 100 may further include a computer system 106 that may be configured to control the optical system 102 and the fluidic system 104 via communication channels 108a and 108b. For example, a computer interface 110 of the optical system 102 may be configured to communicate with the computer system 106 via communication channel 108a.
[0026] During a sequencing reaction, the fluidic system 104 can direct the flow of reagents through one or more reagent tubes 112 to and from a flow cell 114 positioned on a mounting stage 116. The reagents can include, for example, fluorescently labeled nucleotides, buffers, enzymes, and cleavage reagents. The flow cell 114 can include at least one fluidic channel. The flow cell 114 can be a patterned array flow cell or a random array flow cell. The flow cell 114 can include multiple clusters of single-stranded polynucleotides to be sequenced in at least one fluidic channel. The length of the polynucleotides can vary, for example, from about 50 bases, 100 bases, 150 bases, 200 bases, 300 bases, 500 bases to about 1000 bases. The polynucleotides can be attached to one or more fluidic channels of the flow cell 114. In some embodiments, the flow cell 114 can include multiple wells, each well containing multiple copies of a target polynucleotide to be sequenced. The mounting stage 116 can be configured to allow for proper alignment and movement of the flow cell 114 relative to other components of the optical system 102. In one embodiment, the mounting stage 116 can be used to align the flow cell 114 with the lens 118.
[0027] The optical system 102 can include a single light source 120, such as a single laser or a single LED, configured to generate light having a predetermined wavelength, e.g., a narrowly distributed wavelength around 455 nm. In some embodiments, the predetermined wavelength is in the range of 405 nm to 460 nm. However, the embodiments are not limited to any particular wavelength of light. The light source need only be configured to generate the correct wavelength of light that excites the fluorescent labels attached to the nucleotides on the flow cell.
[0028] Light generated by the light source 120 can pass through a fiber optic cable 122 to excite the fluorescent labels in the flow cell 114. A lens 118 attached to a focuser 124 can move along the z-axis. The focused fluorescent emission can be detected by a detector 126, such as a charge-coupled device (CCD) sensor or a complementary metal oxide semiconductor (CMOS) sensor. In some embodiments, nucleotide incorporation can be detected using a zero-mode waveguide, for example, as described in Levene et al. Science 299, 682-686 (2003), Lundquist et al. Opt. Lett. 33, 1026-1028 (2008), and Korlach et al. Proc. Natl. Acad. Sci. USA 105, 1176-1181 (2008), the disclosures of which are incorporated herein by reference in their entirety.
[0029] The filter assembly 128 of the optical system 102 can be configured to filter the fluorescent emission from the fluorescent label in the flow cell 114. The filter assembly 128 can include multiple optical filters, and the correct filter can be selected depending on the particular fluorophore used in the sequencing reaction. In an alternative embodiment, the computer system 106 may automatically determine which optical filter should be used in the sequencing reaction, for example, by scanning a label and / or barcode attached to the sample vial and determining the particular fluorophore used in the sequencing reaction based on the label and / or barcode, or by retrieving information stored in memory related to a previous sequencing reaction, and then control the filter assembly 128 to select and use the desired optical filter. The selected filter can be a long pass filter, a short pass filter, a band stop filter, or a band pass filter depending on the type of fluorescent molecule used in the system. For example, the selected filter can be a band pass filter selected to match a peak in the emission spectrum of a particular fluorescent label.
[0030] In some embodiments, the detector 126 includes one sub-detector, but the filter of the filter assembly 128 may be mechanically switched or rotated in front of the sub-detector, so that different filtered images can be captured sequentially by the sub-detector. In some embodiments, the detector 126 includes one sub-detector, and the filter assembly 128 may include at least one layer of a switchable material having an optical transmittance that is variable upon application of a stimulus, which may be optical, electrical, thermal, or any combination thereof. As a result, the filter assembly 128 may provide multiple optical filters such that different filtered images can be captured sequentially by the sub-detector. In some embodiments, the detector 126 includes one sub-detector, and the filter assembly 128 may include one or more switchable filters based on micro-electromechanical systems technology, so that different filtered images can be captured sequentially by the sub-detector.
[0031] In some embodiments, the detector 126 can include two or more sub-detectors selected according to the set of fluorophores used, for example, a first detector coupled with a first filter and a second detector coupled with a second filter. In some embodiments, the optical system 102 can include two or more dichroic mirrors / beam splitters configured to split the fluorescent emission, such that after splitting the fluorescent emission with the dichroic mirror, the detector 126 can obtain two differently filtered images simultaneously (or closely in time) using two sub-detectors coupled with two different filters. In some embodiments, the detector 126 can include two or more sub-detectors stacked along the direction of incidence of the fluorescent emission. Different wavelengths of fluorescent emission can be differentially attenuated or differentially absorbed along the direction of incidence, such that sub-detectors at different positions along the direction of incidence can be selected according to the set of fluorophores used or configured to capture differently filtered images simultaneously (or closely in time).
[0032] In use, a sample having polynucleotides to be sequenced can be loaded into the flow cell 114 and placed on the mounting stage 116. The computer system 106 can then operate the fluidic system 104 to initiate a sequencing cycle. During a sequencing reaction, the computer system 106 can instruct the fluidic system 104 via the communication interface 108b to provide reagents, e.g., labeled nucleotide analogs, to the flow cell 114. Through the communication interface 108a and the computer interface 110, the computer system 106 can control the light source 120 of the optical system 102 to generate light near a predetermined wavelength, e.g., to excite nucleotide analogs incorporated into a growing primer hybridized to the polynucleotide being sequenced. The computer system 106 can control the detector 126 of the optical system 102 to capture images of diffraction-limited spots of DNA clusters having fluorescently labeled nucleotide analogs. The computer system 106 can receive fluorescent images from the detector 126 and process the received fluorescent images to determine the nucleotide sequence of the polynucleotide being sequenced.
[0033] FIG. 1B illustrates an example of an imaging system 10000 used in the disclosed sequencing techniques. For example, the imaging system 10000 may be used in the exemplary sequencing system 100 illustrated in FIG. 1A. The imaging system 10000 may include a light source 11000 capable of providing light for exciting fluorophores at target points on a sample. The light source 11000 may include one or more lasers, light emitting diodes, or other light sources such that the light source 11000 may provide light of various wavelengths. In some embodiments, the light source 11000 may be configured to selectively provide light having a predetermined range of wavelengths tuned to the set of fluorophores being used. In some embodiments, the light source 11000 may be configured to output light at optical frequencies corresponding to wavelengths within a predetermined range of wavelengths of light. In some embodiments, a user of the disclosed sequencing system may select a particular optical frequency output from the light source 11000 depending on the particular fluorophores used in the sequencing reaction. In an alternative embodiment, the computer system 106 may automatically determine which light frequencies should be output from the light source 11000, for example, by scanning a label and / or barcode attached to a sample vial and determining the particular fluorophore used in the sequencing reaction based on the label and / or barcode, or by searching for information stored in memory related to a previous sequencing reaction, and then control the light source 11000 to select and output the desired light frequencies.
[0034] The imaging system 10000 may include an optical path 12000 from the light source 11000 to the sample 13000, e.g., a microfluidic device including one or more flow chambers in which one or more sequencing reactions occur. In some embodiments, the optical path 12000 may include one or more combinations of mirrors, lenses, prisms, quarter wave plates, half wave plates, polarizers, filters, dichroic mirrors, beam splitters, beam combiners, objective lenses, wide field optics configured to spread light from the light source over a relatively large area of the sample, and the like. The optical path 12000 may be configured to direct light from the light source 11000 to the sample 13000. Additionally, the optical path 12000 may include optical components that may be configured to direct light emitted from 13000 to an integrated detection system 15000. In some embodiments, a portion of the optical elements used to direct light from light source 11000 to sample 13000 are also used to direct light from sample 13000 to integral detection system 15000. Further examples of light paths and optical systems can be found in U.S. Patent No. 7,589,315, U.S. Patent No. 8,951,781, or U.S. Patent No. 9,193,996, each of which is incorporated herein by reference in its entirety.
[0035] The imaging system 10000 can include a scanning system 14000 that effectively moves light relative to the sample 13000 to scan the sample and generate an image. In some embodiments, the scanning system 14000 can be implemented in the optical path 12000. For example, the scanning system 14000 can include one or more scanning mirrors that move relative to each other in the optical path 12000 to effectively move the light from the light source 11000 across the sample. In some embodiments, the scanning system 14000 can be implemented as a mechanical system that physically moves the sample 13000 such that the sample moves relative to the light from the light source 11000. In some embodiments, the scanning system 14000 can be a combination of optical components in the optical path 12000 and a mechanical system for physically moving the sample 13000 such that the light from the light source 11000 and the sample 13000 move relative to each other.
[0036] The imaging system 10000 may include an integral detection system 15000 including one or more photodetectors and associated electronic circuitry, processors, data storage, memory, etc., for acquiring and processing image data of the sample 13000. In some embodiments, the integral detection system 15000 may include photomultiplier tubes, avalanche photodiodes, image sensors (e.g., CCD, CMOS sensors, etc.), and the like. In some embodiments, the photodetectors of the integral detection system 15000 may include components for amplifying the optical signal and may be sensitive to single photons. In some embodiments, the photodetectors of the integral detection system 15000 may have multiple channels or pixels. The integral detection system 15000 may capture one or more images based on the light detected from the sample 13000.
[0037] In some embodiments, the light path 12000 can include an array generator 12100 that can generate multiple exposure areas on the sample 13000. In some embodiments, the array generator 12100 can generate a specific exposure pattern on the sample 13000. These exposure areas can be scanned over the sample 13000 using the scanning system 14000 to selectively illuminate areas of the sample 13000 for imaging. The integral detection system 15000 can integrate signals corresponding to specific points on the sample 13000 as multiple exposure areas are scanned over the sample 13000. For example, for individual points on the sample 13000, the integral detection system 1500 can selectively aggregate detection signals corresponding to individual points when the individual points are illuminated at different times by different exposure areas. In some embodiments, the combination of the array generator 12100 and the integral detection system 15000 can detect light from multiple points on the sample 13000 simultaneously or near simultaneously. In some embodiments, the combination of the array generator 12100 and the integrating detection system 15000 can integrate detected light from multiple points on the sample over time.
[0038] In some embodiments, multiple sequencing reactions can be performed in parallel in multiple flow chambers of the sample 13000. For example, multiple sequencing reactions can be performed on multiple biological specimens. In some embodiments, the multiple sequencing reactions can use different sets of fluorophores. In some embodiments, the light source 11000, the array generator 12100, and the scanning system 14000 can be configured to selectively illuminate different regions of the sample 13000 with different light frequencies depending on the different sets of fluorophores used for the sequencing reactions occurring in the different regions of the sample 13000.
[0039] FIG. 1C illustrates another example of an imaging system 1500 for use in the disclosed sequencing techniques. For example, the imaging system 1500 may be used in the exemplary sequencing system 100 illustrated in FIG. 1A. The imaging system 1500 may be used to image a flow cell 1600 having an upper layer 1671 and a lower layer 1673, which may be separated by a fluid-filled channel 1675. In the illustrated configuration, the upper layer 1671 may be optically transparent, and light from the imaging system 1500 may be focused on an area 1676 on an inner surface 1672 of the upper layer 1671. In an alternative configuration, light from the imaging system 1500 may be focused on an inner surface 1674 of the lower layer 1673. One or both of the surfaces may include array features containing polynucleotides and sequencing reactions to be detected by the imaging system 1500.
[0040] The imaging system 1500 can include an objective lens 1501 configured to direct excitation light from a light source 1502 to a flow cell 1600 and emission light from the flow cell 1600 to a detector 1508. In an exemplary layout, the excitation light from the light source 1502 passes through a lens 1505, then through a beam splitter 1506, and then through the objective lens 1501 on the way to the flow cell 1600. In some embodiments, the light source 1502 can include one or more lasers, light emitting diodes, or any combination thereof. For example, the light source 1502 can include one laser 1503 and one light emitting diode 1504, which can provide light at different wavelengths or wavelength ranges selected depending on the fluorophore used. In an alternative embodiment, the computer system 106 may automatically determine which wavelength or range of wavelengths will be used in the sequencing reaction, for example by scanning a label and / or barcode attached to a sample vial and determining the particular fluorophore to be used in the sequencing reaction based on the label and / or barcode, or by looking up information stored in memory related to a previous sequencing reaction, and then control the light source 1502 to select and output the desired wavelength or range of wavelengths.
[0041] Emission light from the flow cell 1600 may be captured by the objective lens 1501, reflected by a beam splitter, and passed through beam conditioning optics 1507 to a detector 1508 (e.g., a CMOS sensor). The detector 1508 may be configured to detect fluorescent emissions in a particular optical channel (e.g., a particular wavelength range) depending on the fluorophores used in the sequencing reaction. The beam splitter 1506 may direct the emission light in a direction orthogonal to the path of the excitation light. The position of the objective lens 1501 may be moved in the z dimension to change the focusing of the excitation light on the flow cell 1600. The imaging system 1500 may be moved back and forth in the y dimension to capture images of several regions of the flow cell 1600.
[0042] The computer system 106 of the exemplary sequencing system 100 illustrated in Figure 1A can be configured to control the optical system 102 and the fluidic system 104. While many configurations of the computer system 106 are possible, one embodiment is illustrated in Figure 2. As shown in Figure 2, the computer system 106 can include a processor 202 in electrical communication with a memory 204, a storage device 206, and a communication interface 208.
[0043] The processor 202 can be configured to execute instructions that cause the fluidic system 104 to provide reagents to the flow cell 114 during a sequencing reaction. The processor 202 can execute instructions to control the light source 120 of the optical system 102 to generate light near a predetermined wavelength. The processor 202 can execute instructions to control the detector 126 of the optical system 102 and receive data from the detector 126. The processor 202 can execute instructions to process the data received from the detector 126, e.g., a fluorescent image, and determine a nucleotide sequence of a polynucleotide based on the data received from the detector 126. The memory 204 can be configured to store instructions for configuring the processor 202 to perform the functions of the computer system 106 when the sequencing system 100 is powered on. The storage device 206 can store instructions for configuring the processor 202 to perform the functions of the computer system 106 when the sequencing system 100 is powered off. The communication interface 208 can be configured to facilitate communication between the computer system 106, the optical system 102, and the fluidic system 104.
[0044] The computer system 106 can include a user interface 210 configured to communicate with a display device (not shown) for displaying sequencing results of the sequencing system 100. The user interface 210 can be configured to receive input from a user of the sequencing system 100. The optical system interface 212 and the fluidic system interface 214 of the computer system 106 can be configured to control the optical system 102 and the fluidic system 104 through communication links 108a and 108b illustrated in FIG. 1A. For example, the optical system interface 212 can communicate with the computer interface 110 of the optical system 102 via communication link 108a.
[0045] The computer system 106 may include a nucleic acid base determiner 216 configured to determine a nucleotide sequence of a polynucleotide using data received from the detector 126. The nucleic acid base determiner 216 may include one or more of a template generator 218, a position register 220, an intensity extractor 222, an intensity corrector 224, a base caller 226, and a quality score determiner 228. The template generator 218 may be configured to generate a template of the position of the polynucleotide cluster in the flow cell 114 using the fluorescent image captured by the detector 126. The position register 220 may be configured to register the position of the polynucleotide cluster in the flow cell 114 in the fluorescent image captured by the detector 126 based on the position template generated by the template generator 218. The intensity extractor 222 may be configured to extract the intensity of the fluorescent emission from the fluorescent image to generate an extracted intensity. For example, a peak intensity value found at a diffraction-limited spot of a DNA cluster may be extracted from the image and used to represent the signal of the DNA cluster. In another example, the total intensity contained within the diffraction-limited spot of the DNA cluster can be extracted from the image and used to represent the signal of the DNA cluster. Alternatively, intensity estimation can be done through the use of equalization and channel estimation.
[0046] The intensity corrector 224 can be configured to reduce or eliminate noise or aberrations inherent to the sequencing reaction or optical system. For example, the intensity can be affected by laser intensity variations, DNA cluster shape / size variations, non-uniform illumination, optical distortions or aberrations, and / or phasing / prephasing occurring within the DNA cluster. In some embodiments, the intensity corrector 224 can phase correct or pre-phase correct the extracted intensities. The base caller 226 can be configured to determine the nucleobase of the polynucleotide from the corrected intensities. The base of the polynucleotide determined by the base caller 226 can be associated with a quality score determined by the quality score determiner 228. Quality scoring refers to the process of assigning a quality score to each base call. To assess the quality of the base calls from the sequencing reads, an exemplary process can include calculating a set of predictor variable values for the base calls and using the predictor variable values to look up the quality score in a quality table. The quality score can be presented in any suitable format that allows a user to determine the probability of error for any given base call. In some embodiments, the quality score is presented as a numerical value. For example, a quality score may be quoted as QXX, where XX is the score and it indicates that that particular call is 10 -XX / 10 It means that the Q30 has an error probability of 1 in 1000, i.e., 0.1%, and the Q40 has an error rate of 1 in 10,000, i.e., 0.01%, by way of example. The error rate can be calculated using a control nucleic acid. In addition, some metric displays can include error rates on a cycle-by-cycle basis. In some embodiments, the quality table is generated using a calibration data set, where the calibration set represents run and sequence variability. Further details of the calculations that can be performed by the nucleic acid base determiner, calculation of error rates, and quality scores can be found in U.S. Patent No. 8,392,126, U.S. Patent Application Publication No. 2020 / 0080142, and U.S. Patent Application Publication No. 2012 / 0020537, each of which is incorporated herein by reference in its entirety.
[0047] Sequencing using a single optical channel The disclosed technology may use a sequencing-by-synthesis process. During each sequencing cycle, four types of nucleotide analogs may be added and incorporated onto the growing primer-polynucleotide. The four types of nucleotide analogs may have different modifications. For example, three types of nucleotide analogs may be labeled with fluorescent dyes and the fourth type of nucleotide analog may be unlabeled. In some embodiments, the coupling of the dyes to the nucleotides may not result in significant changes in their absorption or emission spectra. After incorporation of the nucleotide analogs, unincorporated nucleotide analogs may be washed away. The polynucleotide may then be stimulated by excitation light and a fluorescence image may be acquired to determine the identity of the nucleotide analogs based on their fluorescence emission.
[0048] Figures 3 and 4 illustrate the working principle of the disclosed sequencing technology. Figure 3 shows an exemplary dye labeling scheme according to some embodiments of the disclosed sequencing technology. In "labeling scheme 1", different types of nucleotide analogs are labeled with the same fluorescent label, but with different percentages of labeled nucleotides for each analog type. The labeled percentages can be about 0, 10, 20, 30, 40, 50, 60, 70, 80, 90, and 100%, or any value therebetween. For example, 100% of dATP is unlabeled, 0% of dATP is labeled with a dye, 67% of dGTP is unlabeled, 33% of dGTP is labeled with the same dye, 34% of dTTP is unlabeled, 66% of dTTP is labeled with a dye, 0% of dCTP is unlabeled, and 100% of dCTP is labeled with a dye. In "labeling scheme 2", different types of nucleotide analogues are labeled with different fluorescent labels having different absorption and / or emission spectra. For example, dATP is unlabeled, dGTP is labeled with a first dye, dTTP is labeled with a second dye, and dCTP is labeled with a third dye. In yet another labeling scheme, different types of nucleotide analogues are labeled with the same fluorescent label, but with different numbers of copies of each fluorescent molecule in each nucleotide. The number of copies can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 10-15, or 15-20. For example, dATP is unlabeled, dGTP is labeled with one copy of a dye, dTTP is labeled with two copies of the same dye, and dCTP is labeled with four copies of the same dye. In general, the above labeling schemes can be combined such that a first type of nucleotide is labeled with a first percentage with a first copy number of a first dye, a second type of nucleotide is labeled with a second percentage with a second copy number of a second dye, a third type of nucleotide is labeled with a third percentage with a third copy number of a third dye, and a fourth type of nucleotide is labeled with a fourth percentage with a fourth copy number of a fourth dye, such that the different types of nucleotides can produce distinct levels of fluorescent emission.
[0049] In some embodiments, the coupling of the dye to the nucleotides may not result in significant changes in their absorption or emission spectra. In some embodiments, a single light source, such as a "blue" laser, can excite the fluorescent labels at a predetermined wavelength, such as about 450 nm. However, the embodiments are not limited to light sources that generate light at this particular wavelength, and other wavelengths corresponding to red, green, violet, or other available wavelengths of light are contemplated. In various embodiments, the output light frequency of the single light source may or may not be tunable. Detection of the dNTPs may include capturing their fluorescent emission by obtaining a fluorescent image ("Image 1" in FIG. 3) with a detector tuned to a predetermined optical channel. In some embodiments, the fluorescent image may be stored for later offline processing. In some embodiments, the fluorescent image may be processed to determine the sequence of the growing primer-polynucleotide in each cluster in real time. The three types of nucleotide analogs may be identified by their distinct detected emission levels (e.g., intensity). The fourth type of nucleotide analog may be identified as having a low signal intensity received by the detector (e.g., substantially close to background levels). For example, dATP may be identified as having a brightness level close to 0, dCTP may be identified as having a full brightness level, dGTP may be identified as having a quarter of full brightness, and dTTP may be identified as having half of full brightness. For example, dATP may be identified as having a brightness level close to 0, dCTP may be identified as having a full brightness level, dGTP may be identified as having a third of full brightness, and dTTP may be identified as having two-thirds of full brightness.
[0050] FIG. 4A is a line graph showing expected fluorescence detection results that may result from an analysis performed according to one embodiment of the disclosed sequencing technique. FIG. 4B is a scatter plot showing similar expected fluorescence detection results detecting the intensities of four nucleotides from a similar image. In these prophetic exemplary embodiments, dATP produces a brightness level close to 0, dCTP produces a full brightness level, dGTP produces approximately one-third of full brightness, and dTTP produces approximately two-thirds of full brightness. However, the detected signal levels of DNA clusters may be affected by various noise sources such as laser intensity fluctuations over time, DNA cluster shape / size fluctuations, non-uniform illumination of the excitation light on the sample tile, optical distortions or aberrations in the light path, shot noise, and / or phasing / prephasing occurring in the DNA clusters. Thus, a collection of DNA cluster signal intensities detected on a sample tile shows a distribution instead of four distinct intensity levels.
[0051] Furthermore, the distribution may change over sequencing cycles, e.g., the variance of the distribution may increase over sequencing cycles due to phasing / prephasing. The four peaks of the distribution may be identified as expected signal levels that allow base calling as the four types of nucleotides. The distribution may be fitted to a statistical model, e.g., the distribution may be decomposed into four Gaussian distributions, and the signal-to-noise ratio of each Gaussian distribution may be estimated as an expectation over a standard deviation. The signal-to-noise ratio of multiple DNA clusters in a sample tile may be defined as a function (e.g., weighted average, geometric mean, etc.) of the signal-to-noise ratios of the four Gaussian distributions. Over sequencing cycles, the expectation and / or standard deviation of the Gaussian distributions may change, and the change in the signal-to-noise ratio may be determined accordingly.
[0052] FIG. 5 shows an exemplary emission spectrum of a collection of nucleotide analogs according to the dye "labeling scheme 2" in FIG. 3, where dATP is unlabeled, dGTP is labeled with "dye 1", dTTP is labeled with "dye 2", and dCTP is labeled with "dye 3". The dyes have different fluorescence emission spectra. Thus, by appropriately selecting a bandpass filter with a transmission window as shown in FIG. 5, the different emission spectrum curves encompass different regions within the transmission window, and thus a detector using a bandpass filter will detect dye 1 with a first brightness level, dye 2 with a second brightness level, and dye 3 with a third brightness level, with different brightness levels. For example, fully-functionalized dGTP (ffG) is labeled with a first dye that has an emission spectrum curve with a peak at about 500 nm when excited by a 450 nm light source. Fully-functionalized dTTP (ffT) is labeled with a second dye having an emission spectrum curve with a peak at about 540 nm when excited by a 450 nm light source. Fully-functionalized dCTP (ffC) is labeled with a third dye having an emission spectrum curve with a peak at about 580 nm when excited by a 450 nm light source, and dATP is unlabeled. In embodiments, the transmission window of the bandpass filter can be 530 nm to 570 nm, such that ffC is brighter than ffT, which is brighter than ffG. In some embodiments, one of the three fluorescent dyes can be a normal Stokes shift dye. As used herein, a normal Stokes shift dye refers to a dye with a Stokes shift of 55 to 95 nm, or a dye with a Stokes shift of about 55, 60, 70, 80, 90, 95 nm, or any value therebetween. One of the three fluorescent dyes can be a short Stokes shift dye. As used herein, a short Stokes shift dye refers to a dye having a Stokes shift of about 5, 10, 20, 30, 40, 50 nm, or any value therebetween.One of the three fluorescent dyes can be a long Stokes shift dye. As used herein, a long Stokes shift dye refers to a dye having a Stokes shift of about 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200 nm, or any value therebetween.
[0053] In some embodiments, the disclosed system for identifying nucleobases in a polynucleotide bound to a substrate may include a first detector configured to detect an intensity of light within a first detection wavelength range, a first light source configured to output light at a first excitation wavelength, and a processor configured to control the first light source to generate light at the first wavelength to stimulate an emission from the polynucleotide bound to the substrate, and identify the nucleobases in the polynucleotide based on the intensity of the emission received by the first detector. In some embodiments, the first nucleobase is identified based on receiving a full intensity emission by the first detector, and the second nucleobase is identified by receiving an emission less than the full intensity emission. In some embodiments, at least four types of nucleobases may be identified from an image captured by only the first detector. In some embodiments, the four types of nucleobases are selected from the group consisting of adenine (A), aminoadenine (Z), cytosine (C), guanine (G), thymine (T), and uracil (U). In some embodiments, the substrate comprises a plurality of chemically functionalized regions, a plurality of cavities, a plurality of optical resonators, a plurality of optical waveguides, or any combination thereof.
[0054] FIG. 6 is a line graph showing the signal-to-noise ratio requirements supporting the disclosed sequencing technique ("PAM4") compared to the prior art two-channel sequencing technique ("PAM2"). The disclosed PAM4 or single optical channel sequencing system requires only half as many images, but requires a higher signal-to-noise ratio to achieve the same level of base calling error rate. FIG. 6 illustrates that PAM4 signaling may require a higher signal-to-noise ratio (SNR) to maintain the same base error rate as PAM2 signaling. In the early part of the sequencing process, the sequencing system may have an excess SNR that can be used to meet the target error rate requirements using the disclosed PAM4 encoding, thereby enabling a single cycle base calling process as described. In the present sequencing system, the maximum intensity may be limited by the fluorophore intensity and the fluorescent molecule binding rate to each nucleotide. Thus, the peak PAM2 signal intensity may be the same as the peak PAM4 signal.
[0055] However, in the PAM4 system, two intermediate levels are introduced, thus encoding more information in each cycle: the distance between signaling levels is smaller in PAM4, resulting in more SNR useful for distinguishing between the intensities of each fluorescently labeled base.
[0056] Considering the potential degradation of signal-to-noise ratio over a sequencing cycle, and considering the higher signal-to-noise ratio requirements to support single optical channel sequencing techniques, in some embodiments, the disclosed system may be configured to switch from a single optical channel sequencing mode ("PAM4") to a two optical channel sequencing mode ("PAM2"), as shown in the flowchart of FIG. 7.
[0057] As shown in Figure 7, the sequencing by synthesis process 700 begins at a start state 705 and then moves to state 710 where a one-color SBS process using one excitation light and one detector as described herein is initiated in the sequencing system. The process 700 then moves to state 715 where a light source, such as a laser or LED, is activated to stimulate fluorescent emission from the labeled nucleotides in the SBS system. The process 700 then moves to state 720 where the fluorescent emission from each cluster in the SBS system is captured by a detector. For example, a CMOS camera may be used to capture an image of the colonies in the flow cell and capture the fluorescent intensity of each cluster.
[0058] Process 700 then moves to state 725 where the nucleotides being added to the growing nucleotide sequence strand in each cluster are identified based on the fluorescence intensity captured in each cluster. As described herein, each of the four different nucleotides may be labeled with the same fluorescent molecule and identified based on the intensity of the fluorescent image captured after excitation by a light source. Process 700 then moves to state 730 where the SNR of the system for identifying the nucleotide is calculated as described above to determine the SNR level for the current sequencing cycle of the system. Process 700 then moves to decision state 735 where a determination is made whether the SNR exceeds a certain threshold or the system has already reached a predetermined number of sequencing cycles. If process 700 makes a determination that the SNR exceeds a predetermined threshold or is less than a predetermined number of sequencing cycles, the process moves to state 740 to continue the one-color SBS sequencing process. However, if at decision state 735 a determination is made that the SNR is below a certain threshold or that a predetermined number of sequencing cycles have elapsed, process 700 may move to state 745 to perform a secondary sequencing process that is less efficient but may operate satisfactorily at a lower overall SNR. For example, a second fluorescent detector may be used for the subsequent SBS cycles. Alternatively, a second light source and / or a second detector may be used for the subsequent SBS cycles. Process 700 may then continue to perform the alternative SBS process until the SNR returns to a predetermined level, or may default to continue performing the alternative sequencing process until the run is complete.
[0059] As described above, in one embodiment, a second detector is used to detect light within a second detection wavelength range. A processor in the system may be configured to identify the nucleobase in the polynucleotide based on the intensity of the luminescence received by the first detector and the second detector. In another example, the system may further include a second light source configured to output light at a second excitation wavelength, and the processor is further configured to control the second light source to generate light at the second excitation wavelength to stimulate luminescence from the polynucleotide, and identify the nucleobase in the polynucleotide based on the intensity of the luminescence received by the first detector. In yet another example, the system may further include a second detector configured to detect light within a second detection wavelength range, and a second light source configured to output light at a second excitation wavelength, and the processor is further configured to control the second light source to generate light at the second excitation wavelength to stimulate luminescence from the polynucleotide, and identify the nucleobase in the polynucleotide based on the intensity of the luminescence received by the first detector and the second detector.
[0060] In some embodiments, the first light source and / or the second light source comprises a laser or a light emitting diode. In some embodiments, the first detector and / or the second detector comprises a complementary metal oxide semiconductor image sensor, a charge coupled device image sensor, a photomultiplier tube, a photodiode, or any combination thereof. In some embodiments, the disclosed system further comprises one or more optical filter materials, one or more diffraction gratings, one or more light dispersing elements, or any combination thereof.
[0061] In certain embodiments, the first detection wavelength range and the second detection wavelength range do not overlap. In some embodiments, the first excitation wavelength is shorter than all of the wavelengths in the first detection wavelength range, and / or the second excitation wavelength is shorter than all of the wavelengths in the second detection wavelength range. In some embodiments, the first excitation wavelength is within 405 nm to 460 nm, and / or the second excitation wavelength is within 405 nm to 460 nm. In some embodiments, the first excitation wavelength is longer than all of the wavelengths in the first detection wavelength range, and / or the second excitation wavelength is longer than all of the wavelengths in the second detection wavelength range. In some embodiments, the first excitation wavelength is within 700 nm to 1400 nm, and / or the second excitation wavelength is within 700 nm to 1400 nm.
[0062] To switch sequencing modes, in some embodiments, the processor is further configured to determine the quality of identifying the nucleobase, and in response to the determined quality, activate the second detector, the second light source, or both. In other embodiments, the processor is further configured to determine the error rate of identifying the nucleobase, and in response to the determined error rate, activate the second detector, the second light source, or both to switch from a first mode of identifying the nucleobase based on the intensity of the emission received by the first detector to a second mode of identifying the nucleobase based on the intensity of the emission received by the first detector and the second detector. In an alternative embodiment, the processor herein is further configured to determine the signal-to-noise ratio of the emission from the plurality of polynucleotides bound to the substrate, and in response to the determined signal-to-noise ratio, activate the second detector, the second light source, or both to switch from a first mode of identifying the nucleobase based on the intensity of the emission received by the first detector to a second mode of identifying the nucleobase based on the intensity of the emission received by the first detector to a second mode of identifying the nucleobase based on the intensity of the emission received by the first detector and the second detector. In further embodiments, the processor is further configured to activate the second detector, the second light source, or both after a predetermined number of cycles of identifying a nucleobase in the polynucleotide. In still further embodiments, the processor is further configured to activate the second detector, the second light source, or both after a predetermined number of cycles of identifying a nucleobase in the polynucleotide and switch from a first mode of identifying the nucleobase based on the intensity of the luminescence received by the first detector to a second mode of identifying the nucleobase based on the intensities of the luminescence received by the first detector and the second detector.
[0063] In some embodiments, the disclosed system further comprises a fluidic device configured to deliver a set of nucleotide analogs to the polynucleotide, the set of nucleotide analogs comprising a first nucleotide analog bound to the fluorophore with a first probability, a second nucleotide analog bound to the fluorophore with a second probability, and a third nucleotide analog bound to the fluorophore with a third probability. In some embodiments, the set of nucleotide analogs further comprises a fourth nucleotide analog bound to the fluorophore with a fourth probability, or a fourth nucleotide analog not bound to the fluorophore. To facilitate switching between sequencing modes, in some embodiments, the processor is further configured to control the fluidic device to deliver alternative sets of nucleotide analogs to the polynucleotide based on a determined quality of identifying the nucleobase, based on a determined error rate of identifying the nucleobase, based on a determined signal-to-noise ratio, or after a predetermined number of cycles of identifying the nucleobase. The alternative sets of nucleotide analogs are suitable for a two-channel sequencing process.
[0064] In some embodiments, the disclosed system further comprises a fluidic device configured to deliver a set of nucleotide analogs to the polynucleotide, the set of nucleotide analogs comprising a first nucleotide analog bound to a first number of fluorophores, a second nucleotide analog bound to a second number of fluorophores, and a third nucleotide analog bound to a third number of fluorophores. In some embodiments, the set of nucleotide analogs further comprises a fourth nucleotide analog bound to a fourth number of fluorophores, or a fourth nucleotide analog not bound to a fluorophore. To facilitate switching between sequencing modes, in some embodiments, the processor is further configured to control the fluidic device to deliver alternative sets of nucleotide analogs to the polynucleotide based on a determined quality of identifying the nucleobase, based on a determined error rate of identifying the nucleobase, based on a determined signal-to-noise ratio, or after a predetermined number of cycles of identifying the nucleobase. The alternative sets of nucleotide analogs are suitable for a two-channel sequencing process.
[0065] In some embodiments, the disclosed system further comprises a fluidic device configured to deliver a set of nucleotide analogs to the polynucleotide, the set of nucleotide analogs comprising a first nucleotide analog bound to a first fluorescent label, a second nucleotide analog bound to a second fluorescent label, and a third nucleotide analog bound to a third fluorescent label. In some embodiments, the set of nucleotide analogs further comprises a fourth nucleotide analog bound to a fourth fluorescent label or a fourth nucleotide analog not bound to a fluorescent label. In some embodiments, the fluorescent labels have different fluorescent brightness and / or emission spectra. To facilitate switching between sequencing modes, in some embodiments, the processor is further configured to control the fluidic device to deliver alternative sets of nucleotide analogs to the polynucleotide based on a determined quality of identifying the nucleobase, based on a determined error rate of identifying the nucleobase, based on a determined signal-to-noise ratio, or after a predetermined number of cycles of identifying the nucleobase. The alternative sets of nucleotide analogs are suitable for a two-channel sequencing process.
[0066] In some embodiments, the disclosed system further comprises a fluidic device configured to deliver a set of nucleotide analogs to the polynucleotide, the set of nucleotide analogs comprising a first nucleotide analog bound to a first number of first fluorescent labels, a second nucleotide analog bound to a second number of second fluorescent labels, and a third nucleotide analog bound to a third number of third fluorescent labels. In some embodiments, the set of nucleotide analogs further comprises a fourth nucleotide analog bound to a fourth number of fourth fluorescent labels, or a fourth nucleotide analog not bound to a fluorescent label. In some embodiments, the fluorescent labels have different fluorescent brightness and / or emission spectra. To facilitate switching between sequencing modes, in some embodiments, the processor is further configured to control the fluidic device to deliver alternative sets of nucleotide analogs to the polynucleotide based on a determined quality of identifying the nucleobase, based on a determined error rate of identifying the nucleobase, based on a determined signal-to-noise ratio, or after a predefined number of cycles of identifying the nucleobase. The alternative sets of nucleotide analogs are suitable for a two-channel sequencing process.
[0067] In some embodiments, the nucleotide analogs used in the disclosed sequencing system can be fully functionalized nucleotides. The linker located between the nucleotide base and the fluorescent molecule can include one or more cleavage groups. Prior to a subsequent sequencing cycle, the fluorescent label can be removed from the nucleotide analog by cleavage of the linker. For example, the linker attaching the fluorescent label to the nucleotide analog can include an azide and / or alkoxy group, e.g., on the same carbon, such that the linker can be cleaved after each incorporation cycle with a phosphine reagent, thereby releasing the fluorescent label. The nucleotide triphosphates can be reversibly blocked at the 3' position so that sequencing is controlled, and only a single nucleotide analog can be added to each extended primer-polynucleotide in each cycle. For example, the 3' ribose position of the nucleotide analog can include both an alkoxy and an azide functional group that can be removed by cleavage with a phosphine reagent, thereby creating a nucleotide that can be further extended. Prior to a subsequent sequencing cycle, the reversible 3' block can be removed and another nucleotide analog can be added to each extended primer-polynucleotide.
[0068] In some embodiments, the fluorescent label is selected from the group consisting of polymethine derivatives, coumarin derivatives, benzopyran derivatives, chromenoquinoline derivatives, and compounds containing bis-boron heterocycles, such as BOPPY and BOPYPY. In some embodiments, the fluorescent label is attached to the nucleotide via a cleavable linker. In some further embodiments, the labeled nucleotide may have a fluorescent label attached to the C5 position of a pyrimidine base or the C7 position of a 7-deazapurine base, optionally through a linker moiety. For example, the nucleobase may be 7-deazaadenine, and the dye is attached to the 7-deazaadenine at the C7 position, optionally through a cleavable linker. The nucleobase may be 7-deazaguanine, and the dye is attached to the 7-deazaguanine at the C7 position, optionally through a cleavable linker. The nucleobase may be cytosine, and the dye is attached to the cytosine at the C5 position, optionally through a cleavable linker. As another example, the nucleobase may be thymine or uracil, and the dye is attached to the thymine or uracil at the C5 position, optionally through a cleavable linker. In some further embodiments, the cleavable linker may contain a chemical moiety similar or the same as the reversible terminator 3' hydroxy blocking group, such that the 3' hydroxy blocking group and the cleavable linker may be removed under the same reaction conditions or in a single chemical reaction. Non-limiting examples of cleavable linkers include the LN3 linker, the sPA linker, and the AOL linker, each of which is exemplified below. [ka] [ka] [ka]
[0069] In some embodiments, the nucleotides are selected from the group consisting of an analog of dGTP, an analog of dTTP, an analog of dUTP, an analog of dCTP, and an analog of dATP. In some embodiments, the first nucleotide is a first reversibly blocked nucleotide triphosphate (rbNTP), the second nucleotide is a second rbNTP, the third nucleotide is a third rbNTP, and the fourth nucleotide is a fourth rbNTP, and each of the first nucleotide, the second nucleotide, the third nucleotide, and the fourth nucleotide are different types of nucleotides from each other. In some embodiments, the four rbNTPs are selected from the group consisting of rbATP, rbTTP, rbUTP, rbCTP, and rbGTP. In some embodiments, each of the four rbNTPs includes a modified base and a reversible terminator 3' blocking group. Non-limiting examples of 3' blocking groups include azidomethyl ( * -CH2N3), substituted azidomethyl (e.g. * -CH(CHF2)N3 or * -CH(CH2F)N3) and * -CH2-O-CH2-CH=CH2, where the asterisk * indicates a point attachment to the 3' oxygen of the ribose or deoxyribose ring of a nucleotide.
[0070] Further details regarding dyes and fully functionalized nucleotides can be found in U.S. Patent Application Publication Nos. 2018 / 0094140 and 2020 / 0277670, International Patent Application Publication No. 2017 / 051201, and U.S. Provisional Patent Application Nos. 63 / 057758 and 63 / 127061, the disclosures of which are incorporated by reference in their entireties.
[0071] sample In some embodiments, the sample comprises or consists of purified or isolated polynucleotides from tissue samples, biological fluid samples, cell samples, and the like. Suitable biological fluid samples include, but are not limited to, blood, plasma, serum, sweat, tears, phlegm, urine, sputum, ear fluid, lymph, saliva, cerebrospinal fluid, ravage, bone marrow suspension, vaginal flow, trans-cervical lavage, brain fluid, peritoneal fluid, milk, secretions of the respiratory tract, intestinal tract, and urogenital tract, amniotic fluid, milk, and leukophoresis samples. In some embodiments, the sample is a sample that can be easily obtained by a non-invasive procedure, such as, for example, blood, plasma, serum, sweat, tears, phlegm, urine, sputum, ear fluid, saliva, or feces. In certain embodiments, the sample is a peripheral blood sample, or the plasma and / or serum fraction of a peripheral blood sample. In other embodiments, the biological sample is a swab or smear, a biopsy specimen, or a cell culture. In another embodiment, the sample is a mixture of two or more biological samples, for example, the biological sample can include two or more of a biological fluid sample, a tissue sample, and a cell culture sample. As used herein, the terms "blood," "plasma," and "serum" expressly include fractions thereof or processed portions thereof. Similarly, if a sample is taken from a biopsy, swab, smear, etc., the "sample" expressly includes processed fractions or portions derived from the biopsy, swab, smear, etc.
[0072] In certain embodiments, samples may be obtained from sources including, but not limited to, samples from different individuals, samples from different developmental stages of the same individual or different individuals, samples from different diseased individuals (e.g., individuals having cancer or suspected of having a genetic disease), normal individuals, samples obtained at different stages of a disease in an individual, samples obtained from individuals receiving different treatments for a disease, samples from individuals subjected to different environmental factors, samples from individuals predisposed to a disease condition, samples from individuals exposed to an infectious agent, etc.
[0073] In one exemplary but non-limiting embodiment, the sample is a maternal sample obtained from a pregnant woman, e.g., a pregnant woman. The maternal sample can be a tissue sample, a biological fluid sample, or a cell sample. In another exemplary but non-limiting embodiment, the maternal sample is a mixture of two or more biological samples, e.g., the biological sample can include two or more of a biological fluid sample, a tissue sample, and a cell culture sample.
[0074] In certain embodiments, samples can also be obtained from in vitro cultured tissues, cells, or other polynucleotide-containing sources. Cultured samples can be taken from sources including, but not limited to, cultures (e.g., tissues or cells) maintained in different media and conditions (e.g., pH, pressure, or temperature), cultures (e.g., tissues or cells) maintained for different time periods, cultures (e.g., tissues or cells) treated with different factors or reagents (e.g., drug candidates, or modulators), or cultures of different types of tissues and / or cells.
[0075] In some embodiments, the use of the disclosed sequencing techniques does not include the preparation of a sequencing library. In other embodiments, the sequencing techniques contemplated herein include the preparation of a sequencing library. In one exemplary approach, the preparation of a sequencing library includes the production of a random collection of adaptor-modified DNA fragments (e.g., polynucleotides) that are ready to be sequenced.
[0076] A polynucleotide sequencing library can be prepared from DNA or RNA, including equivalents, analogs of either DNA or cDNA, such as DNA or cDNA that is a complementary or copy DNA produced from an RNA template, for example, by the action of reverse transcriptase. The polynucleotides may be derived from a double-stranded form (e.g., dsDNA, such as genomic DNA fragments, cDNA, PCR amplification products, etc.), or in certain embodiments, the polynucleotides may be derived from a single-stranded form (e.g., ssDNA, RNA, etc.) and converted to a dsDNA form. By way of illustration, in certain embodiments, single-stranded mRNA molecules can be copied into double-stranded cDNA suitable for use in preparing a sequencing library. The exact sequence of the primary polynucleotide molecule is generally not critical to the method of library preparation and may be known or unknown. In one embodiment, the polynucleotide molecules are DNA molecules. More specifically, in certain embodiments, the polynucleotide molecule represents the entire genetic complement of an organism or substantially the entire genetic complement of an organism and is a genomic DNA molecule (e.g., cellular DNA, cell free DNA (cfDNA), etc.), but typically includes intronic and exonic sequences (coding sequences), as well as non-coding regulatory sequences such as promoter and enhancer sequences. In certain embodiments, the primary polynucleotide molecule comprises a human genomic DNA molecule, e.g., a cfDNA molecule present in the peripheral blood of a pregnant subject.
[0077] Methods for isolating nucleic acids from biological sources may vary depending on the nature of the source. One skilled in the art can easily separate nucleic acids from sources as required for the methods described herein. In some cases, it may be advantageous to fragment large nucleic acid molecules (e.g., cellular genomic DNA) in a nucleic acid sample to obtain polynucleotides of a desired size range. Fragmentation may be random or specific, for example, as achieved using restriction endonuclease digestion. Methods for random fragmentation may include, for example, limited DNase digestion, alkaline treatment, and physical shearing. Fragmentation may also be achieved by any of a number of methods known to those skilled in the art. For example, fragmentation may be achieved by mechanical means, including but not limited to nebulization, sonication, and hydroshearing.
[0078] In some embodiments, sample nucleic acid is obtained from cfDNA that has not undergone fragmentation.For example, cfDNA typically exists as fragments of less than about 300 base pairs, and therefore fragmentation is typically not required to use cfDNA sample to generate sequencing library.
[0079] Typically, whether polynucleotides are forcibly fragmented (e.g., fragmented in vitro) or naturally present as fragments, they are converted to blunt-ended DNA with a 5'-phosphate and a 3'-hydroxyl. Standard protocols, such as protocols for sequencing using the Illumina platform, instruct users to purify the end-repaired product prior to dA-tailing for end-repaired sample DNA, and to purify the dA-tailed product prior to the adapter-ligating step of library preparation.
[0080] In various embodiments, verification of sample integrity and sample tracking can be achieved by sequencing a mixture of sample genomic nucleic acid, e.g., cfDNA, and associated marker nucleic acid that has been introduced into the sample, e.g., prior to processing.
[0081] Sequencing technology The disclosed sequencing systems and methods can be compatible with any sequencing technology based on optical detection, such as next-generation sequencing (NGS), fluorescent in situ sequencing (FISSEQ), and massively parallel signature sequencing (MPSS). In one embodiment, the disclosed systems and methods can be compatible with NGS technologies that allow multiple samples to be sequenced individually as genomic molecules (i.e., singleplex sequencing) or as pooled samples containing indexed genomic molecules in a single sequencing run (e.g., multiplex sequencing). These methods can generate up to millions of DNA sequence reads.
[0082] The disclosed technology may implement sequencing reactions such as those incorporating sequencing-by-synthesis methods described in U.S. Patent Application Publication Nos. 2007 / 0166705, 2006 / 0188901, 2006 / 0240439, 2006 / 0281109, 2005 / 0100900, U.S. Patent No. 7,057,026, International Application Nos. PCT 2005 / 065814, 2006 / 064199, and 2007 / 010251, the disclosures of which are incorporated herein by reference in their entireties. In some embodiments, the sequencer may implement sequencing-by-synthesis methods similar to those used in the HiSeq, MiSeq, or HiScanSQ systems from Illumina (San Diego, Calif.).
[0083] Alternatively, sequencing by ligation techniques may be used in techniques that disclose sequencing by ligation techniques, such as those described in U.S. Patent Nos. 6,969,488, 6,172,218, and 6,306,597, the disclosures of which are incorporated herein by reference in their entireties. Ligation sequencing library techniques use DNA ligase to incorporate oligonucleotides and identify the incorporation of such oligonucleotides.
[0084] The disclosed technology may be implemented in several sequencing technologies, such as the sequencing-by-hybridization platform from Affymetrix Inc. (Sunnyvale, CA), as well as sequencing-by-synthesis platforms from 454 Life Sciences (Bradford, CT) and Helicos Biosciences (Cambridge, MA), the sequencing-by-ligation platform from Applied Biosystems (Foster City, CA), or the SMRT technology from Pacific Biosciences.
[0085] In one exemplary, but non-limiting embodiment, the methods described herein include obtaining sequence information about nucleic acids in a sample using Illumina sequencing-by-synthesis and reversible terminator-based sequencing chemistry (e.g., as described in Bentley et al., Nature 6:53-59
[2009] ). Illumina sequencing technology can include attachment of fragmented genomic DNA to a planar, optically transparent surface to which oligonucleotide anchors are attached. For example, the template DNA is end-repaired to generate 5' phosphorylated blunt ends, and a single A base is added to the 3' ends of the blunt phosphorylated DNA fragments using the polymerase activity of the Klenow fragment. This addition prepares the DNA fragments for ligation to oligonucleotide adaptors, which have single T base overhangs at their 3' ends to enhance ligation efficiency. The adaptor oligonucleotides are complementary to the flow cell anchor oligos. Under limiting dilution conditions, adaptor-modified single-stranded template DNA is added to the flow cell and immobilized by hybridization to the anchor oligo. The attached DNA fragments are extended and the bridges are amplified to create an ultra-high density sequencing flow cell with hundreds of millions of clusters, each containing about 1,000 copies of the same template. In one embodiment, the randomly fragmented genomic DNA is amplified using PCR before undergoing cluster amplification. Alternatively, amplification-free (e.g., PCR-free) genomic library preparation is used, in which the randomly fragmented genomic DNA is enriched using only cluster amplification (Kozarewa et al., Nature Methods 6:291-295
[2009] ). Sequencing-by-synthesis reactions may use reversible terminators with removable fluorescent dyes. Short sequence reads of about tens to hundreds of base pairs are aligned against a reference genome to identify unique mapping of the short sequence reads to the reference genome. After the first read is completed, the template can be regenerated in situ to allow a second read from the opposite end of the fragment.Thus, either single-end or paired-end sequencing of DNA fragments can be used. Detailed information regarding paired-end sequencing can be found in U.S. Patent No. 7,601,499 and U.S. Patent Application Publication No. 2012 / 0,053,063, which are incorporated by reference.
[0086] In some embodiments, the Illumina sequencing-by-synthesis platform includes fragment clustering. Clustering is a process in which each fragment molecule is isothermally amplified. In some embodiments, the fragments have two different adapters attached to the two ends of the fragment, which allow the fragments to hybridize to two different oligos on the surface of a flow cell lane. The fragments further include or are connected to two index sequences at the two ends of the fragment, which provide labels to identify different samples in multiplex sequencing.
[0087] In some embodiments, a flow cell for clustering in the Illumina platform is a glass slide with lanes. Each lane is a glass channel coated with a lawn of two types of oligos. Hybridization is enabled by the first of the two types of oligos on the surface. This oligo is complementary to the first adapter at one end of the fragment. A polymerase creates a complementary strand of the hybridized fragment. The double-stranded molecule is denatured and the original template strand is washed away. The remaining strand is clonally amplified by bridge application in parallel with many other remaining strands.
[0088] In bridge amplification, the strands fold back and a second adapter region on the second end of the strand hybridizes to a second type of oligo on the flow cell surface. Polymerase generates a complementary strand, forming a double-stranded bridge molecule. This double-stranded molecule is denatured, resulting in two single-stranded molecules tethered to the flow cell via two different oligos. This process is then repeated across millions of clusters, which occur simultaneously, resulting in clonal amplification of all fragments. After bridge amplification, the reverse strand is cleaved and washed away, leaving only the forward strand. The 3' end is blocked to prevent undesired priming.
[0089] After clustering, sequencing begins by extending the first sequencing primer to generate the first read. In each cycle, fluorescently labeled nucleotides compete to add to the growing strand. Only one is incorporated based on the sequence of the template. After each nucleotide addition, the cluster is excited by a light source and a characteristic fluorescent signal is emitted. The number of cycles determines the length of the read. The emission wavelength and signal intensity determine the base call. For a given cluster, all identical strands are read simultaneously. Hundreds of millions of clusters, or tens of thousands to millions of clusters, are sequenced in a massively parallel fashion. Upon completion of the first read, the read products are washed away.
[0090] In a process involving two index primers, an index 1 primer is introduced and hybridizes to the index 1 region on the template. The index region provides fragment identification useful for demultiplexing samples in a multiplex sequencing process. The index 1 read is generated similarly to the first read. After the index 1 read is completed, the read product is washed away and the 3' end of the strand is deprotected. The template strand then folds over the second oligo on the flow cell and binds to the second oligo. The index 2 sequence is read in the same manner as index 1. The index 2 read product is then washed away upon completion of the step.
[0091] With two index reads, read 2 first extends the second flow cell oligo with a polymer to form a double-stranded bridge. This double-stranded DNA is denatured and the 3' end is blocked. The original forward strand is cleaved and washed away, leaving the reverse strand. Read 2 begins with the introduction of the read 2 sequencing primer. As with read 1, the sequencing steps are repeated until the desired length is achieved. The read 2 product is washed away. This entire process generates millions of reads, representing all fragments. Sequences from the pooled sample library are separated based on the unique index introduced during sample preparation. For each sample, reads of similar extension base calls are locally clustered. Forward and reverse reads are paired to create contiguous sequences. These contiguous sequences are aligned to the reference genome for variant identification.
[0092] Computer Systems In some embodiments, the disclosed systems and methods may involve approaches to shift or distribute certain sequence data analysis functions and sequence data storage to a cloud computing environment or cloud-based network. User interaction with sequencing data, genomic data, or other types of biological data may be mediated through a central hub that stores the data and controls access to various interactions with the data. In some embodiments, the cloud computing environment may also provide sharing of protocols, analysis methods, libraries, sequence data, and distributed processing for sequencing, analysis, and reporting. In some embodiments, the cloud computing environment facilitates correction or annotation of sequence data by users. In some embodiments, the systems and methods may be implemented in a computer browser, on-demand, or online.
[0093] In some embodiments, software written to perform the methods described herein is stored on some form of computer readable medium, such as memory, a CD-ROM, a DVD-ROM, a memory stick, a flash drive, a hard drive, an SSD hard drive, a server, a mainframe storage system, or the like.
[0094] In some embodiments, the method may be written in any of a variety of suitable programming languages, for example, compiled languages such as C, C#, C, Fortran, and Java. Other programming languages may include scripting languages such as Perl, MatLab, SAS, SPSS, Python, Ruby, Pascal, Delphi, R, and PHP. In some embodiments, the method is written in C, C#, C++, Fortran, Java, Perl, R, Java, or Python. In some embodiments, the method may be a separate application having data entry and data display modules. Alternatively, the method may be a computer software product, and distributed objects may include classes that include applications that include the computational methods described herein.
[0095] In some embodiments, the method may be incorporated into existing data analysis software such as that found on sequencing equipment. Software including the computer-implemented methods described herein may be either installed directly on a computer system or indirectly held on a computer-readable medium and loaded onto the computer system as needed. Additionally, the method may be located on a computer that is remote to where the data is being generated, such as software found on a server or the like that is maintained at a separate location to where the data is being generated, such as that provided by a third-party service provider.
[0096] The assay instrument, desktop computer, laptop computer, or server may house a processor in operative communication with accessible memory that includes instructions for implementation of the systems and methods. In some embodiments, the desktop computer or laptop computer is in operative communication with one or more computer-readable storage media or devices and / or output devices. The assay instrument, desktop computer, and laptop computer may operate under many different computer-based operating languages, such as those utilized by Apple-based computer systems or PC-based computer systems. The assay instrument, desktop, and / or laptop computer, and / or server system may further provide a computer interface for creating or modifying experimental definitions and / or conditions, viewing data results, and monitoring experimental progress. In some embodiments, the output device may be a graphic user interface such as a computer monitor or computer screen, a printer, a portable device such as a portable digital assistant (i.e., personal digital assistant, PDA, Blackberry®, iPhone®), a tablet computer (e.g., iPAD®), a hard drive, a server, a memory stick, a flash drive, etc.
[0097] The computer readable storage device or medium may be any device, such as a server, mainframe, supercomputer, magnetic tape system, etc. In some embodiments, the storage device may be located in close proximity to the assay instrument, e.g., adjacent or proximate to the assay instrument. For example, the storage device may be located in relation to the assay instrument, in the same room, in the same building, in an adjacent building, on the same floor in a building, on a different floor in a building, etc. In some embodiments, the storage device may be located off-site or distal to the assay instrument. For example, the storage device may be located in a different area of a city, in a different city, in a different state, in a different country, etc., relative to the assay instrument. In embodiments where the storage device is located distal to the assay instrument, communication between the assay instrument and one or more of a desktop, laptop, or server is typically via an internet connection, either wirelessly or by network cable via an access point. In some embodiments, the storage device may be maintained and managed by an individual or entity directly associated with the assay instrument, while in other embodiments, the storage device may be maintained and managed by a third party, typically in a distal location relative to an individual or entity associated with the assay instrument. In the embodiments described herein, the output device can be any device for visualizing data.
[0098] The assay instruments, desktops, laptops, and / or server systems may be used to store and / or retrieve computer-implemented software programs incorporating computer code for executing and implementing the computational methods described herein, data for use in implementing the computational methods, and the like. One or more of the assay instruments, desktops, laptops, and / or servers may comprise one or more computer-readable storage media for storing and / or retrieving software programs incorporating computer code for executing and implementing the computational methods described herein, data for use in implementing the computational methods, and the like. The computer-readable storage media may include, but are not limited to, one or more of a hard drive, SSD hard drive, CD-ROM drive, DVD-ROM drive, floppy disk, tape, flash memory stick, or card. Additionally, a network, including the Internet, may be a computer-readable storage medium. In some embodiments, a computer-readable storage medium refers to a computational resource storage accessible by a computer network via the Internet or a corporate network provided by a service provider, rather than, for example, from a local desktop or laptop computer at a distal location to the assay instrument.
[0099] In some embodiments, computer-implemented software programs incorporating computer code for executing and implementing the computational methods described herein, computer readable storage media for storing and / or retrieving data used in implementing the computational methods, etc. are operated and maintained by a service provider in operative communication with the assay instruments, desktops, laptops, and / or server systems via an internet connection or a network connection.
[0100] In some embodiments, a hardware platform for providing a computing environment includes a processor (i.e., CPU) where processor time and memory layout, such as random access memory (RAM), are system considerations. For example, smaller computer systems offer cheaper, faster processors and larger memory and storage capabilities. In some embodiments, a graphics processing unit (GPU) can be used. In some embodiments, a hardware platform for performing the computational methods described herein comprises one or more computer systems having one or more processors. In some embodiments, smaller computers are clustered together to create a supercomputer network.
[0101] In some embodiments, the computational methods described herein are implemented on a collection of inter- or intra-connected computer systems (i.e., grid technologies) that may cooperatively run a variety of operating systems. For example, the CONDOR framework (University of Wisconsin-Madison) and systems available from United Devices are illustrative of the cooperation of multiple independent computer systems for the purpose of handling large amounts of data. These systems may provide a Perl interface for submitting, monitoring, and managing large sequence analysis jobs on a cluster in a serial or parallel configuration.
[0102] definition Unless otherwise defined, technical and scientific terms used in this disclosure have the same meaning as commonly understood by those skilled in the art to which this disclosure belongs.See, for example, Singleton et al., Dictionary of Microbiology and Molecular Biology 2nd ed., J. Wiley & Sons (New York, NY 1994); Sambrook et al., Molecular Cloning, A Laboratory Manual, Cold Spring Harbor Press (Cold Spring Harbor, NY 1989).For purposes of this disclosure, the following terms are defined below.
[0103] As used herein, the terms "well," "cavity," and "chamber" are used interchangeably and refer to discrete features defined in a device that can contain a fluid (e.g., liquid, gel, gas). An example of an array of the present devices can have one or more wells. Furthermore, it should be understood that a cross-section of a well taken parallel to a surface of the substrate that at least partially defines the well can be curved, square, polygonal, hyperbolic, conical, angular, etc.
[0104] As used herein, the term "cluster" or "clamp" refers to a group of molecules, e.g., a group of DNA, or a group of signals. In some embodiments, the signals of a cluster are derived from different features. In some implementations, a signal clump represents a physical area covered by one amplified oligonucleotide. Each signal clump could ideally be observed as several signals. Thus, overlapping signals could be detected from the same clump of signals. In some embodiments, a cluster or clump of signals can include one or more signals or spots that correspond to a particular feature. When used in connection with a microarray device or other molecular analysis device, a cluster can include one or more signals that together occupy a physical area occupied by an amplified oligonucleotide (or other polynucleotide or polypeptide with the same or similar sequence). For example, if the feature is an amplified oligonucleotide, the cluster can be a physical area covered by one amplified oligonucleotide. In other embodiments, a cluster or clump of signals does not have to correspond strictly to a feature. For example, a spurious noise signal can be included in a signal cluster, but not necessarily within the feature area. For example, a cluster of signals from a four-cycle sequencing reaction could include at least four signals.
[0105] As used herein, the term "spot radius" or "cluster radius" refers to a defined radius that encompasses a diffraction-limited spot or cluster of signals. Thus, by defining the cluster radius larger or smaller, a larger number of signals can fall within the radius for subsequent ordering and selection. The cluster radius can be defined by any distance measure, such as pixels, meters, millimeters, or any other useful measure of distance.
[0106] As used herein, a "signal" refers to a detectable event, such as a light emission, e.g., a light emission in an image. Thus, in some embodiments, a signal can represent any detectable light emission (i.e., a "spot") captured in an image. Thus, as used herein, a "signal" can refer to an actual light emission from a feature of the specimen, or it can refer to a spurious emission that does not correlate with an actual feature. Thus, a signal can result from noise and can later be discarded when it is not representative of an actual feature of the specimen.
[0107] As used herein, the "intensity" of emitted light refers to the intensity of light transmitted per unit area, where the area is measured on a plane perpendicular to the direction of propagation of the light beam, and the intensity is the amount of energy transmitted per unit time. In some embodiments, the signal "strength", "amplitude", "magnitude", or "level" may be used synonymously with signal intensity. In some embodiments, the image obtained by the detector is approximated or proportional to an intensity map integrated over some amount of time. In some embodiments, the signal of a diffraction-limited spot of DNA clusters is extracted from the image as the total intensity contained in the spot up to a factor of the integration time. For example, the signal of a DNA cluster may be defined as the intensity contained within the spot radius of the DNA cluster up to a factor of the integration time. In other embodiments, the peak intensity value found within the spot radius may be used to represent the signal of the DNA cluster up to a factor of the integration time.
[0108] As used herein, the process of aligning a template of signal locations onto a given image is referred to as “registration,” and the process of determining the intensity values or intensity values for each signal within the template in a given image is referred to as “intensity extraction.” For registration, the methods and systems provided herein may take advantage of the random nature of signal clamp locations by using image correlation to align the template to the image.
[0109] As used herein, a "nucleotide" comprises a nitrogen-containing heterocyclic base, a sugar, and one or more phosphate groups. A nucleotide is a monomeric unit of a nucleic acid sequence. Examples of nucleotides include, for example, ribonucleotides or deoxyribonucleotides. In a ribonucleotide (RNA), the sugar is ribose, and in a deoxyribonucleotide (DNA), the sugar is deoxyribose, i.e., a sugar lacking the hydroxyl group present at the 2' position of the ribose. The nitrogen-containing heterocyclic base can be a purine base or a pyrimidine base. Purine bases include adenine (A) and guanine (G), as well as modified derivatives or analogs thereof. Pyrimidine bases include cytosine (C), thymine (T), and uracil (U), as well as modified derivatives or analogs thereof. The C-1 atom of the deoxyribose is attached to the N-1 of the pyrimidine or the N-9 of the purine. The phosphate group can be in mono-, di-, or triphosphate form. These nucleotides may be naturally occurring nucleotides, however it should be further understood that non-naturally occurring nucleotides, modified nucleotides or analogs of the aforementioned nucleotides may also be used.
[0110] As used herein, a "nucleobase" is a heterocyclic base, such as adenine, guanine, cytosine, thymine, uracil, inosine, xanthine, hypoxanthine, or a heterocyclic derivative, analog, or tautomer thereof. Nucleobases can be naturally occurring or synthetic. Non-limiting examples of nucleobases include adenine, guanine, thymine, cytosine, uracil, xanthine, hypoxanthine, 8-azapurine, purine substituted with methyl or bromine at the 8-position, 9-oxo-N6-methyladenine, 2-aminoadenine, 7-deazaxanthine, 7-deazaguanine, 7-deaza-adenine, N4-ethanocytosine, 2,6-diaminopurine, N6-ethano-2,6-diaminopurine, 5-methylcytosine, 5-(C3-C6)-alkynylcytosine, 5-fluorouracil, 5-bromouracil, thiamine, uracil ... auracil, pseudoisocytosine, 2-hydroxy-5-methyl-4-triazolopyridine, isocytosine, isoguanine, inosine, 7,8-dimethylalloxazine, 6-dihydrothymine, 5,6-dihydrouracil, 4-methyl-indole, ethenoadenine, and the non-naturally occurring nucleobases described in U.S. Pat. Nos. 5,432,272 and 6,150,510, and International Application Nos. PCT 92 / 002258, 93 / 10820, 94 / 22892, and 94 / 24144, and in Fasman (Practical Handbook of Biochemistry and Molecular Biology, pp. 385-394, 1989, CRC Press, Boca Raton, LO), all of which are incorporated herein by reference in their entireties.
[0111] The term "nucleic acid" or "polynucleotide" refers to deoxyribonucleotide or ribonucleotide polymers in single- or double-stranded form and includes, unless specifically limited, known analogues of natural nucleotides that hybridize to nucleic acids in a manner similar to naturally occurring nucleotides, such as peptide nucleic acid (PNA) and phosphorothioate DNA. Unless otherwise specified, a particular nucleic acid sequence includes its complementary sequence. Nucleotides include, but are not limited to, ATP, dATP, CTP, dCTP, GTP, dGTP, UTP, TTP, dUTP, 5-methyl-CTP, 5-methyl-dCTP, ITP, dITP, 2-amino-adenosine-TP, 2-amino-deoxyadenosine-TP, 2-thiothymidine triphosphate, pyrrolo-pyrimidine triphosphate, and 2-thiocytidine, as well as alpha thiotriphosphate for all of the above, and 2'-O-methyl-ribonucleotide triphosphate for all of the above bases. Modified bases include, but are not limited to, 5-Br-UTP, 5-Br-dUTP, 5-F-UTP, 5-F-dUTP, 5-propynyl dCTP, and 5-propynyl-dUTP.
[0112] The polymerase used is generally an enzyme for joining 3'-OH 5'-triphosphate nucleotides, oligomers and their analogs. Polymerases include DNA-dependent DNA polymerase, DNA-dependent RNA polymerase, RNA-dependent DNA polymerase, RNA-dependent RNA polymerase, T7 DNA polymerase, T3 DNA polymerase, T4 DNA polymerase, T7 RNA polymerase, T3 RNA polymerase, SP6 RNA polymerase, DNA polymerase I, Klenow fragment, Thermophilus aquaticus DNA polymerase, Tth DNA polymerase, VentR® DNA polymerase (New England Biolabs), Deep VentR® DNA polymerase (New England Biolabs), Bst DNA polymerase large fragment, Stoeffel fragment, 90N DNA polymerase, 90N DNA polymerase, Pfu DNA polymerase, TfI DNA polymerase, Tth DNA polymerase, RepliPHI Phi29 polymerase, TIi Examples of suitable polymerases include, but are not limited to, DNA polymerases, eukaryotic DNA polymerase beta, telomerase, Therminator™ polymerase (New England Biolabs), KOD HiFi™ DNA polymerase (Novagen), KOD1 DNA polymerase, Q-beta replicase, terminal transferase, AMV reverse transcriptase, M-MLV reverse transcriptase, Phi6 reverse transcriptase, HIV-1 reverse transcriptase, novel polymerases discovered by bioprospecting, and polymerases cited in US Patent Application Publication No. 2007 / 0048748, US Patent No. 6,329,178, US Patent No. 6,602,695, and US Patent No. 6,395,524 (incorporated by reference). These polymerases include wild-type, mutant isoforms, and engineered variants. "Encode" or "parse" are verbs referring to transferring from one format to another, and refer to transferring the genetic information of the target template sequence into the reporter configuration.
[0113] Nucleosides and nucleotides may be labeled at sites on the sugar or nucleobase. Dyes may be attached at any position on the nucleotide base, for example, via a linker. In certain embodiments, Watson-Crick base pairing may still be performed on the resulting analog. Specific nucleobase labeling sites include the C5 position of pyrimidine bases, or the C7 position of 7-deazapurine bases. A linker group may be used to covalently attach the dye to the nucleoside or nucleotide. As used herein, the terms "covalently attached" or "covalently bonded" refer to the formation of a chemical bond characterized by the sharing of electron pairs between atoms. For example, a covalently attached polymer coating refers to a polymer coating that forms a chemical bond with the functionalized surface of a substrate, as compared to attachment to the surface by other means, for example, adhesion or electrostatic interactions. It will be understood that polymers that are covalently attached to a surface may be attached by means in addition to covalent bonds.
[0114] A variety of different types of linkers with different lengths and chemical properties can be used. The term "linker" encompasses any moiety that is useful for connecting one or more molecules or compounds to each other, to other components of a reaction mixture, and / or to a reaction site. For example, a linker can attach a reporter molecule or "label" (e.g., a fluorescent dye) to a reaction component. In certain embodiments, the linker is a member selected from substituted or unsubstituted alkyl (e.g., 2-5 carbon chains), substituted or unsubstituted heteroalkyl, substituted or unsubstituted aryl, substituted or unsubstituted heteroaryl, substituted or unsubstituted cycloalkyl, and substituted or unsubstituted heterocycloalkyl. In one example, the linker moiety is selected from straight and branched carbon chains, optionally containing at least one heteroatom (e.g., at least one functional group such as ether, thioether, amide, sulfonamide, carbonate, carbamate, urea, and thiourea), and optionally containing at least one aromatic, heteroaromatic, or non-aromatic ring structure (e.g., cycloalkyl, phenyl). In certain embodiments, molecules with trifunctional linking capabilities are used, including, but not limited to, cynuric chloride, mealamine, diaminopropanoic acid, aspartic acid, cysteine, glutamic acid, pyroglutamic acid, S-acetylmercaptosuccinic anhydride, carbobenzoxylidine, histidine, lysine, serine, homoserine, tyrosine, piperidinyl-1,1-aminocarboxylic acid, diaminobenzoic acid, etc. In certain specific embodiments, hydrophilic PEG (polyethylene glycol) linkers are used.
[0115] In certain embodiments, the linker is derived from a molecule that includes at least two reactive functional groups (e.g., one at each end), which can react with complementary reactive functional groups on various reactive components or can be used to immobilize one or more reactive components at a reaction site. "Reactive functional group," as used herein, includes olefins, acetylenes, alcohols, phenols, ethers, oxides, halides, aldehydes, ketones, carboxylic acids, esters, amides, cyanates, isocyanates, thiocyanates, isothiocyanates, amines, hydrazines, hydrazones, hydrazides, diazos, diazonium, nitro, nitriles, mercaptans, sulfides, disulfides, sulfoxides, sulfones, sulfonic acids, sulfinic acids, acetals, ketones ... It refers to groups including, but not limited to, tars, anhydrides, sulfates, sulfenic acid isonitriles, amidines, imides, imidates, nitrones, hydroxylamines, oximes, hydroxamic acids, thiohydroxamic acids, allenes, orthoesters, sulfites, enamines, ynamines, ureas, pseudoureas, semicarbazides, carbodiimides, carbamates, imines, azides, azo compounds, azoxy compounds, and nitroso compounds. Reactive functional groups also include those used to prepare bioconjugates, such as N-hydroxysuccinimide esters, maleimides, and the like.
[0116] The cleavable linker may be, by way of non-limiting example, an electrophilically cleavable linker, a nucleophilically cleavable linker, a photocleavable linker, a linker cleavable under reducing conditions (e.g., a disulfide or azide-containing linker), a linker cleavable under oxidative conditions, a linker cleavable by use of a safety lock linker, a linker that is cleavable by a removal mechanism. The use of a cleavable linker to attach the dye compound to the substrate moiety allows for the label to be removed after detection, if necessary, to avoid any interfering signals in downstream steps.
[0117] In some embodiments, one or more dye molecules or labeling molecules may be attached to the nucleotide base by non-covalent interactions or by a combination of covalent and non-covalent interactions via multiple intervening molecules. In one example, the nucleotide or nucleotide analog newly incorporated by the polymerase synthesizing from the target polynucleotide is initially unlabeled. One or more fluorescent labels may then be introduced to the nucleotide or nucleotide analog by binding to a labeled affinity reagent containing one or more fluorescent dyes. The use of unlabeled nucleotides and affinity reagents in sequencing by synthesis is disclosed in US Patent Application Publication No. 2013 / 0079232, incorporated herein by reference. For example, one, two, three, or each of the four different types of nucleotides (e.g., dATP, dCTP, dGTP, and dTTP or dUTP) in the reaction mix may be initially unlabeled. Each of the four types of nucleotides (e.g., dNTPs) may have a 3' hydroxy blocking group to ensure that only a single base can be added by the polymerase to the 3' end of the copy polynucleotide being synthesized from the target polynucleotide. After incorporation of the unlabeled nucleotide, an affinity reagent that specifically binds to the incorporated dNTP can then be introduced to provide a labeled extension product containing the incorporated dNTP. The affinity reagent can be designed to specifically bind to the incorporated dNTP via, for example, an antibody-antigen interaction or a ligand-receptor interaction. The dNTP can be modified to include a specific antigen that pairs with a specific antibody included in the corresponding affinity reagent. Thus, one, two, three, or each of the four different types of nucleotides can be specifically labeled via their corresponding affinity reagent. In some embodiments, the affinity reagent can include a small molecule or protein tag that can bind to a hapten portion of the nucleotide (e.g., streptavidin-biotin, anti-DIG and DIG, anti-DNP and DNP), an antibody (including but not limited to a binding fragment of an antibody, a single chain antibody, a bispecific antibody, etc.), an aptamer, a knottin, an affimer, or any other known agent that binds to the incorporated nucleotide with suitable specificity and affinity.In some embodiments, the hapten moiety of the unlabeled nucleotide may be attached to the nucleobase via a cleavable linker, which may be cleaved under the same reaction conditions to remove the 3' blocking group. In some embodiments, one affinity reagent may be labeled with multiple copies of the same fluorescent dye, for example, 1, 2, 3, 4, 5, 6, 8, 10, 12, 15 copies of the same dye. In some embodiments, each affinity reagent may be labeled with a different number of copies of the same fluorescent dye. In some embodiments, the first affinity reagent may be labeled with a first number of first fluorescent dyes, the second affinity reagent may be labeled with a second number of second fluorescent dyes, the third affinity reagent may be labeled with a third number of third fluorescent dyes, and the fourth affinity reagent may be labeled with a fourth number of fourth fluorescent dyes. In some embodiments, each affinity reagent may be labeled with a different combination of one or more types of dyes, each type of dye having a specific number of copies. In some embodiments, different affinity reagents may be labeled with different dyes that can be excited by the same light source, but each dye has a distinguishable fluorescence intensity or a distinguishable emission spectrum, hi some embodiments, different affinity reagents may be labeled with the same dye at different molar ratios, resulting in a measurable difference in their fluorescence intensities.
[0118] The nucleotide analogs may be attached to or associated with one or more optically detectable labels to provide a detectable signal. In some embodiments, the optically detectable label may be a fluorescent compound, such as a small molecule fluorescent label. Fluorescent molecules (fluorophores) suitable as fluorescent labels include 1,5 IAEDANS; 1,8-ANS; 4-methylumbelliferone; 5-carboxy-2,7-dichlorofluorescein; 5-carboxyfluorescein (5-FAM); fluorescein amidite (fluorescein amidite ... amidite, FAM; 5-carboxyfluorescein; tetrachloro-6-carboxyfluorescein (TET); hexachloro-6-carboxyfluorescein (HEX); 2,7-dimethoxy-4,5-dichloro-6-carboxyfluorescein (JOE); VIC®; NED®; tetramethylrhodamine (TMR); 5-carboxytetramethylrhodamine (5-TAMRA); 5-HAT (hydroxytryptamine); 5-hydroxytryptamine (HAT); 5-ROX (carboxy-X-rhodamine); 6-carboxyrhodamine 6G; 6-JOE; Light Cycler® Red 610; Light Cycler® Red 640; Light Cycler® Red 670; Light Cycler® Red 705; 7-amino-4-methylcoumarin; 7-aminoactinomycin D (7-AAD); 7-hydroxy-4-methylcoumarin; 9-amino-6-chloro-2-methoxyacridine; 6-methoxy-N-(4-aminoalkyl)quinolinium bromide hydrochloride (ABQ); acid fuchsin; ACMA (9-amino-6-chloro-2-methoxyacridine); acridine orange; acridine red; acridine yellow; acriflavine;Acriflavine Feulgen SITSA;AFP-Autofluorescent Protein- (Quantum Biotechnologies);Texas Red;Texas Red-X Conjugate;Thiadicarbocyanine (DiSC3);Thiazine Red R;Thiazole Orange;Thioflavin 5;Thioflavin S;Thioflavin TCN;Thiolight;Thiosol Orange;Tinopol CBS (Calcofluor White);TMR;TO-PRO-1;TO-PRO-3;TO-PRO-5;TOTO-1;TOTO-3;TriColor (PE-Cy5);TRITC (TetramethylRodamine-lsoThioCyanate);True Blue;TruRed;Ultralight;Uranine B;Uvitex SFC;WW 781; X-rhodamine; X-rhodamine-5-(and-6)-isothiocyanate (5(6)-XRITC); Xylene Orange; Y66F; Y66H; Y66W; YO-PRO-1; YO-PRO-3; YOYO-1; YOYO-3, interchelating dyes such as Sybr Green, thiazole orange; Alexa Fluor® dye series (Molecular members of the Cy Dye fluorophore series (GE Healthcare) including broad spectrum such as Cy3, Cy3B, Cy3.5, Cy5, Cy5.5, Cy7; members of the Oyster® dye fluorophores (Denovo Biolabels) including Oyster-500, -550, -556, 645, 650, 656;For example, members of the DY-Labels series (Dyomics) having absorption maxima in the range of 418 nm (DY-415) to 844 nm (DY-831), such as DY-415, -495, -505, -547, -548, -549, -550, -554, -555, -556, -560, -590, -610, -615, -630, -631, -632, -633, -634, -635, -636, -647, -648, -649, -650, -651, -652, -675, -676, -677, -680, -681, -682, -700, -701, -730, -731, -732, -734, -750, -751, -752, -776, -780, -781, -782, -831, -480XL, -481XL, -485XL, -510XL, -520XL, -521XL; fluorescently labeled ATTO series (ATTO-TEC members of the CAL Fluor® series (Biosearch Technologies), such as, but not limited to, CAL Fluor® Gold 540, CAL Fluor® Orange 560, Quasar® 570, CAL Fluor® Red 590, CAL Fluor® Red 610, CAL Fluor® Red 635, Quasar® 570, and Quasar® 670. In some embodiments, a first optically detectable label interacts with a second optically detectable moiety to alter the detectable signal, for example, via fluorescence resonance energy transfer ("FRET"; also known as Förster resonance energy transfer);
[0119] Fluorescent labels utilized by the systems and methods disclosed herein can have different peak absorption wavelengths, for example, ranging from 400 nm to 800 nm. In some embodiments, the peak absorption wavelength of the fluorescent label can be, or can be approximately, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720, 730, 740, 750, 760, 770, 780, 790, 800 nm, or a number or range between any two of these values. In some embodiments, the peak absorption wavelength of the fluorescent label can be at least or at most 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720, 730, 740, 750, 760, 770, 780, 790, or 800 nm.
[0120] Fluorescent labels can have different peak emission wavelengths, for example ranging from 400 nm to 800 nm. In some embodiments, the peak emission wavelength of the fluorescent label can be, or can be approximately, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720, 730, 740, 750, 760, 770, 780, 790, 800 nm, or a number or range between any two of these values. In some embodiments, the peak emission wavelength of the fluorescent label can be at least or at most 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720, 730, 740, 750, 760, 770, 780, 790, or 800 nm.
[0121] Fluorescent labels can have different Stokes shifts, for example ranging from 10 nm to 200 nm. In some embodiments, the Stokes shift can be, or can be about, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200 nm, or a number or range between any two of these values. In some embodiments, the Stokes shift can be at least, or at most, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 nm.
[0122] In some embodiments, the distance between the peak emission wavelengths of any two fluorescent labels may vary, for example, in the range of 10 nm to 200 nm. In some embodiments, the distance between the peak emission wavelengths of any two fluorescent labels may be, or may be approximately, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200 nm, or a number or range between any two of these values. In some embodiments, the distance between the peak emission wavelengths of any two fluorescent labels may be at least, or at most, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 nm.
[0123] A "light source" may be any device capable of emitting energy along the electromagnetic spectrum. The light source may be a visible light (VIS), ultraviolet light (UV) and / or infrared light (IR) source. "Visible light" (VIS) generally refers to a band of electromagnetic radiation having wavelengths from about 400 nm to about 750 nm. "Ultraviolet (UV) light" generally refers to electromagnetic radiation having wavelengths shorter than those of visible light, or in the range of about 10 nm to about 400 nm. "Infrared light" or infrared radiation (IR) generally refers to electromagnetic radiation having wavelengths greater than the VIS range, or from about 750 nm to about 50,000 nm. The light source may also provide full spectrum light. The light source may output light from a selected wavelength or range of wavelengths. In some embodiments of the invention, the light source may be configured to provide light above or below a predetermined wavelength, or may provide light within a predetermined range. The light source may be used in combination with a filter to selectively transmit or block selected wavelengths of light from the light source. The light source may be connected to the intensity source by one or more electrical connectors. An array of light sources may be connected in series or parallel to the intensity source. The intensity source may be a battery, or a vehicle electrical system or a building electrical system. The light source may be connected to the intensity source via control electronics (control circuitry). The control electronics may comprise one or more switches. The one or more switches may be automated, controlled by a sensor, timer, or other input, or controlled by a user, or a combination thereof. For example, a user may operate a switch to turn on a UV light source. The light source may be applied on a constant basis until it is turned off, or may be pulsed (repeated on / off cycles) until it is turned off. In some embodiments, the light source may be switched from a continuous on state to a pulsed state or vice versa. In some embodiments, the light source may be configured to get brighter or dimmer over time.
[0124] For operation, the light source may be connected to an intensity source capable of providing sufficient intensity to illuminate the sample. The control electronics may be used to switch the intensity on or off based on input from a user or some other input, and may also be used to modulate the intensity to a suitable level (e.g., to control the brightness of the output light). The control electronics may be configured to turn the light source on and off as needed. The control electronics may include switches for manual, automatic, or semi-automatic operation of the light source. The one or more switches may be, for example, transistors, relays, or electromechanical switches. In some embodiments, the control circuitry may further comprise an AC-DC and / or DC-DC converter for converting a voltage from a voltage source to a suitable voltage for the light source. The control circuitry may comprise a DC-DC regulator for regulating the voltage. The control circuitry may further comprise a timer and / or other circuit elements for energizing the optical filter for a fixed period following receipt of the input. The switch may be activated manually, automatically in response to a predetermined condition, or with the aid of a timer. For example, the control electronics may process information such as user input, stored instructions, etc.
[0125] One or more of a plurality of light sources may be provided. In some embodiments, each of the plurality of light sources may be the same. Alternatively, one or more of the light sources may vary. The light characteristics of the light emitted by the light sources may be the same or different. The plurality of light sources may or may not be independently controllable. One or more characteristics of the light sources may or may not be controlled, including but not limited to whether the light source is on or off, the brightness of the light source, the wavelength of the light, the intensity of the light, the angle of illumination, the position of the light source, or any combination thereof.
[0126] In some embodiments, the light output from the light source may be from about 350 to about 750 nm, or any amount or range therebetween, such as from about 350 nm to about 360, 370, 380, 390, 400, 410, 420, 430, or about 450 nm, or any amount or range therebetween. In other embodiments, the light from the light source may be from about 550 to about 700 nm, or any amount or range therebetween, such as from about 550 to about 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, or about 700 nm, or any amount or range therebetween. In some embodiments, the wavelength of the light generated by the light source may vary, for example, in the range of 400 nm to 800 nm. In some embodiments, the wavelength of the light generated by the light source can be, or can be approximately, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720, 730, 740, 750, 760, 770, 780, 790, 800 nm, or a number or range between any two of these values. In some embodiments, the wavelength of light generated by the light source may be at least or at most 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720, 730, 740, 750, 760, 770, 780, 790, or 800 nm. The light source may be capable of emitting electromagnetic radiation in any spectrum. In some embodiments, the light source may have a wavelength between 10 nm and 100 μm. In some embodiments, the wavelength of the light can be between 100 nm and 5000 nm, between 300 nm and 1000 nm, or between 400 nm and 800 nm.In some embodiments, the wavelength of the light may be less than and / or equal to 10 nm, 100 nm, 200 nm, 300 nm, 400 nm, 500 nm, 600 nm, 700 nm, 800 nm, 900 nm, 1000 nm, 1100 nm, 1200 nm, 1300 nm, 1500 nm, 1750 nm, 2000 nm, 2500 nm, 3000 nm, 4000 nm, or 5000 nm.
[0127] In one example, the light source may be a light-emitting diode (LED) (e.g., a gallium arsenide (GaAs) LED, an aluminum gallium arsenide (AlGaAs) LED, a gallium arsenide phosphide (GaAsP) LED, an aluminum gallium indium phosphide (AlGaInP) LED, a gallium (III) phosphide (GaP) LED, an indium gallium nitride (InGaN) / gallium (III) nitride (GaN) LED, or an aluminum gallium phosphide (AlGaP) LED). In another example, the light source can be a laser, e.g., a vertical cavity surface emitting laser (VCSEL), or other suitable light emitter, such as an indium-gallium-aluminum-phosphide (InGaAIP) laser, a gallium-arsenide phosphide / gallium phosphide (GaAsP / GaP) laser, or a gallium-aluminum-arsenide / gallium-aluminum-arsenide (GaAlAs / GaAs) laser.Other examples of light sources may include, but are not limited to, electronically excited light sources (e.g., cathodoluminescence, electronically stimulated luminescence (ESL bulbs), cathode ray tubes (CRT monitors), Nixie tubes), incandescent light sources (e.g., carbon button lamps, conventional incandescent light bulbs, halogen lamps, glow bars, Nernst lamps), electroluminescent (EL) light sources (e.g., light emitting diodes, organic light emitting diodes, polymer light emitting diodes, solid state lighting, LED lamps, electroluminescent sheets, electroluminescent wires), gas discharge light sources (e.g., fluorescent lamps, induction lighting, hollow cathode lamps, neon and argon lamps, plasma lamps, xenon flash lamps), or high intensity discharge light sources (e.g., carbon arc lamps, ceramic discharge metal halide lamps, hydrogen iodide arc lamps, mercury vapor lamps, metal halide lamps, sodium vapor lamps, xenon arc lamps). Alternatively, the light source may be a bioluminescent, chemiluminescent, phosphorescent, or fluorescent light source.
[0128] Optical filters can be tailored for transparency or haze, translucency, transparency or opacity, light transmittance (LT), switching speed, durability, light stability, contrast ratio, and light transmittance state (e.g., dark or light state). "Light transmittance" (LT) refers to the amount of light transmitted or passing through an optical filter or a device or apparatus comprising the same. LT may be expressed with reference to light transmittance and / or change in a particular type or wavelength of light (e.g., from about 10% visible light transmittance (LT) to about 90% LT, etc.). Alternatively, LT may be expressed as absorbance, and may optionally include reference to one or more wavelengths absorbed. According to some embodiments, an optical filter can be selected or configured to have an LT in one state of less than 80%, or less than 70%, or less than 60%, or less than 50%, or less than 40%, or less than 30%, or less than 20%, or less than 10%, or any amount or range therebetween. According to some embodiments, the optical filter can be selected or configured to have an LT in another state of greater than 80%, or greater than 70%, or greater than 60%, or greater than 50%, or greater than 40%, or greater than 30%, or greater than 20%, or greater than 10%, or any amount or range therebetween.
[0129] The filters can be bandpass filters and can have peak transmissions at various wavelengths ranging from 400 nm to 800 nm, in some embodiments the peak transmissions can be at or about 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720, 730, 740, 750, 760, 770, 780, 790, 800 nm, or a number or range between any two of these values. In some embodiments, the peak transmittance can be at least or at most 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720, 730, 740, 750, 760, 770, 780, 790, or 800 nm. The width of the transmission window of the filter can vary, for example, from 1 nm to 50 nm. In some embodiments, the width of the filter may be, or may be approximately, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50 nm, or a number or range between any two of these values. In some embodiments, the width of the filter may be at least or at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, or 50 nm. A short-pass filter may be considered as a special band-pass filter with a lower limit of the transmission window close to 0 nm. A long-pass filter may be considered as a special band-pass filter with an upper limit of the transmission window close to infinity. A band-stop filter may be defined as the complement of a band-pass filter.
[0130] As used herein, an "optical channel" is a predefined profile of optical frequencies (or equivalently wavelengths). For example, a first optical channel may have wavelengths from 500 nm to 600 nm. To capture an image in the first optical channel, a detector responsive only to light from 500 nm to 600 nm may be used, or a bandpass filter with a transmission window of 500 nm to 600 nm may be used to filter the light incident on a detector responsive to light from 300 nm to 800 nm. A second optical channel may have wavelengths from 300 nm to 450 nm and 850 nm to 900 nm. To capture an image in the second optical channel, a detector responsive to light from 300 nm to 450 nm and another detector responsive to light from 850 nm to 900 nm may be used, and the detection signals of the two detectors may then be combined. Alternatively, to capture an image in the second optical channel, a bandstop filter that rejects light from 300 nm to 900 nm may be used in front of a detector responsive to light from 451 nm to 849 nm.
[0131] Additional Notes The embodiments described herein are exemplary. Modifications, rearrangements, alternative processes, etc. may be made to these embodiments and still fall within the teachings described herein. One or more of the steps, processes, or methods described herein may be performed by one or more suitably programmed processing and / or digital devices.
[0132] The various exemplary imaging or data processing techniques described in connection with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or a combination of both. To illustrate this interchangeability of hardware and software, various exemplary components, blocks, modules, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the particular application and design constraints imposed on the overall system. The described functionality may be implemented in various ways for each particular application, and such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure.
[0133] Various exemplary detection systems described in connection with the embodiments disclosed herein may be implemented or performed by a machine such as a processor configured with specific instructions, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. The processor may be a microprocessor, but alternatively the processor may be a controller, microcontroller, or state machine, combinations thereof, and the like. A processor may also be implemented as a combination of computing devices, such as a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in association with a DSP core, or any other such configuration. For example, the systems described herein may be implemented using discrete memory chips, a portion of memory within a microprocessor, flash, EPROM, or other types of memory.
[0134] Elements of the methods, processes, or algorithms described in connection with the embodiments disclosed herein may be embodied directly in hardware, in software modules executed by a processor, or in a combination of the two. The software modules may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disk, removable disk, CD-ROM, or any other form of computer-readable storage medium known in the art. An exemplary storage medium may be coupled to the processor such that the processor can read information from, and write information to, the storage medium. Alternatively, the storage medium may be integral to the processor. The processor and the storage medium may reside in an ASIC. The software modules may include computer-executable instructions that cause a hardware processor to execute the computer-executable instructions.
[0135] Unless otherwise indicated, conditional language used in this specification, such as "can," "might," "may," "eg," and the like, unless otherwise indicated and as otherwise understood within the context of use, is intended to generally convey that certain embodiments are included, and other embodiments do not include the particular features, elements, and / or steps. Thus, such conditional language does not generally imply that features, elements, and / or conditions are in any manner required for one or more embodiments, or that one or more embodiments necessarily include logic for determining or prompting, with or without author input, whether those features, elements, and / or conditions are included or performed in any particular embodiment. Terms such as "comprising," "including," "having," "involving," and the like, are synonymous and used in an inclusive, open-ended manner and do not exclude additional elements, features, acts, operations, and the like. Also, the term "or," when used, for example, to connect a list of elements, is used in its inclusive sense (and not its exclusive sense), so as to mean one, some, or all of the elements in the list.
[0136] Disjunctive language, such as the phrase "at least one of X, Y, or Z," is otherwise understood in context as it is generally used to state that an item, term, etc. can be either X, Y, or Z, or any combination thereof (e.g., X, Y, and / or Z), unless specifically stated otherwise. Thus, such disjunctive language is generally not intended to, and should not, imply that a particular embodiment requires that at least one of X, at least one of Y, or at least one of Z, respectively, be present.
[0137] Terms such as "about" or "approximately" are synonymous and are used to indicate that the value modified by the term has an understood range associated with it, which may be ±20%, ±15%, ±10%, ±5%, or ±1%. The term "substantially" is used to indicate that a result (e.g., a measurement) is close to a target value, where close may mean, for example, that the result is within 80% of the value, within 90% of the value, within 95% of the value, or within 99% of the value.
[0138] Unless otherwise noted, articles such as "a" or "an" should generally be construed to include one or more of the listed items. Thus, phrases such as "a device configured to" or "a device to" are intended to include one or more of the listed devices. Such one or more listed devices may also be collectively configured to perform the described detailed descriptions. For example, "a processor configured to perform detailed descriptions A, B, and C" may include a first processor to perform operations in conjunction with a second processor configured to perform detailed descriptions A and B and C.
[0139] While the above detailed description has illustrated, described, and pointed out novel features applied to the exemplary embodiments, it will be understood that various omissions, substitutions, and changes in the form and details of the illustrated devices or algorithms can be made without departing from the spirit of the present disclosure. As will be recognized, certain embodiments described herein may be embodied in forms that do not provide all of the features and advantages described herein, since some features may be used or practiced separately from others. All changes that come within the meaning and range of equivalency of the claims are intended to be embraced within their scope.
[0140] It is to be understood that all combinations of the foregoing concepts (provided that such concepts are not mutually inconsistent) are intended to be part of the inventive subject matter disclosed herein. In particular, all combinations of claimed subject matter appearing at the end of this disclosure are contemplated to be part of the inventive subject matter disclosed herein.
Claims
1. 1. A system for identifying nucleobases in a polynucleotide bound to a substrate, comprising: a first detector configured to detect an intensity of light within a first detection wavelength range; a first light source configured to output light at a first excitation wavelength; 1. A processor, comprising: controlling the first light source to emit light at the first excitation wavelength to stimulate emission of light from the polynucleotides bound to the substrate; a processor configured to identify a nucleobase in the polynucleotide based on the intensity of the luminescence received by the first detector.
2. The system described in claim 1, further comprising a second detector configured to detect light within a second detection wavelength range.
3. 2. The system of claim 1, wherein a first nucleobase is identified based on receiving a full intensity emission by the first detector, and a second nucleobase is identified by receiving an emission that is less than the full intensity emission.
4. 10. The system of claim 1, wherein at least four types of nucleobases can be identified from an image captured by the first detector.
5. The system described in claim 2, wherein the processor is further configured to identify nucleic acid bases in the polynucleotide based on the intensity of the luminescence received by the first detector and the second detector.
6. a second light source configured to output light at a second excitation wavelength, controlling the second light source to generate light at the second excitation wavelength to stimulate emission of light from the polynucleotide; and identifying a nucleobase in the polynucleotide based on the intensity of the emitted light received by the first detector; or identifying the nucleobases in the polynucleotide based on the intensities of the emitted light received by the first detector and the second detector. The system of claim 2 , further configured to:
7. the processor: determining a quality score identifying the nucleobase, an error rate identifying the nucleobase, or a signal-to-noise ratio of the light emitted from the plurality of polynucleotides bound to the substrate; 6. The system of claim 5, further configured to: determine whether to activate the second detector, the second light source, or both in response to the determined quality score, error rate, or signal-to-noise ratio.
8. and a fluidic device configured to deliver a set of nucleotide analogues to the polynucleotide, the set of nucleotide analogues comprising: a first population of nucleotide analogues, wherein a first predetermined percentage of the nucleotide analogues in the first population are conjugated to a fluorophore; a second population of nucleotide analogues, wherein a second predetermined percentage of the nucleotide analogues in the second population are conjugated to the fluorophore; and a third population of nucleotide analogues, wherein a third predetermined percentage of the nucleotide analogues in the third population are conjugated to the fluorophore.
9. 9. The system of claim 8, wherein the set of nucleotide analogues further comprises a fourth population of nucleotide analogues, wherein a fourth predetermined percentage of the nucleotide analogues in the fourth population are conjugated to the fluorophore or the fourth population of nucleotide analogues is not conjugated to the fluorophore.
10. and a fluidic device configured to deliver a set of nucleotide analogues to the polynucleotide, the set of nucleotide analogues comprising: a first nucleotide analogue linked to a first number of fluorophores; a second nucleotide analogue attached to a second number of said fluorophores or second fluorescent labels; and a third nucleotide analog attached to a third number of said fluorophores or a third fluorescent label.
11. 11. The system of claim 10, wherein the set of nucleotide analogues further comprises a fourth nucleotide analogue linked to a fourth number of fluorophores or a fourth nucleotide analogue that is not linked to a fluorophore.
12. and a fluidic device configured to deliver a set of nucleotide analogues to the polynucleotide, the set of nucleotide analogues comprising: a first nucleotide attached to a first fluorescent label; a second nucleotide attached to a second fluorescent label; and a third nucleotide bound to a third fluorescent label.
13. 13. The system of claim 12, wherein the set of nucleotide analogs further comprises a fourth nucleotide analog attached to a fourth fluorescent label or a fourth nucleotide analog not attached to a fluorescent label.
14. The system of claim 12 , wherein the fluorescent labels have different fluorescent brightness and / or emission spectra.
15. 13. The system of any one of claims 8, 10, or 12, wherein the processor is further configured to control the fluidic device to deliver an alternative set of nucleotide analogs to the polynucleotide based on a determined quality score identifying the nucleobase, based on a determined error rate identifying the nucleobase, based on a determined signal-to-noise ratio, or after a predetermined number of cycles of identifying the nucleobase.