Primary Analysis in Next Generation Sequencing
Patent Information
- Application Number
- JP2024534498
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-06-30
- Filing Date
- 2022-12-09
- Publication Date
- 2025-12-16
AI Technical Summary
Existing methods for identifying cluster centers in DNA sequencing images face challenges such as merged clusters due to accuracy and alignment issues, leading to improper sequence identification and excessive processing time.
The method involves capturing flow cell images at sub-pixel resolution, determining intensity and purity of candidate cluster centers, and using interpolation and additional cluster location identification to improve accuracy without increasing imager resolution, utilizing specialized processors and FPGAs for real-time processing.
This approach enhances cluster center detection accuracy, reduces processing time, and decreases system cost by identifying more clusters with lower error rates, enabling real-time primary analysis in DNA sequencing.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is a continuation-in-part of U.S. Patent Application No. 17 / 854,042, filed on June 30, 2022, which is a continuation-in-part of U.S. Patent Application No. 17 / 547,602, filed on December 10, 2021, which is a continuation-in-part of U.S. Patent Application No. 17 / 219,556, filed on March 31, 2021, which claims priority to U.S. Provisional Patent Application No. 63 / 072,649, filed on August 31, 2020. This application is a continuation-in-part of U.S. Patent Application No. 17 / 725,065, filed on April 20, 2022, which claims priority to U.S. Provisional Patent Application No. 63 / 316,784, filed on March 4, 2022. This application is a continuation-in-part of U.S. Patent Application No. 17 / 725,042, filed April 20, 2022, which claims priority to U.S. Provisional Patent Application No. 63 / 316,790, filed March 4, 2022. This application claims priority to U.S. Provisional Patent Application No. 63 / 349,421, filed June 6, 2022. The entireties of the above patent applications are incorporated herein by reference in their entireties.
[0002] The present disclosure relates generally to image data analysis, and more particularly to identifying cluster locations for performing base calling within digital images of flow cells during DNA sequencing. [Background technology]
[0003] Next generation sequencing by synthesis using a flow cell can be used to identify the sequence of DNA. When single-stranded DNA fragments from a sequencing library are flooded across a flow cell, the fragments are typically randomly attached to the surface of the flow cell due to complementary oligomers bound to the surface of the flow cell or beads present thereon. The DNA fragments are then subjected to an amplification process, so that copies of a given fragment form clusters or polonies of denatured cloned nucleotide strands. In some embodiments, a single bead can comprise the cluster, and the bead can be attached to the flow cell at a random position.
[0004] To identify the sequence of the strands, the strand pairs are reassembled one nucleotide base at a time. During each base assembly cycle, a mixture of single nucleotides, each attached to a fluorescent label (or tag) and a blocker, is flooded across the flow cell. The nucleotides are attached to complementary positions on the strands. A blocker is included so that only one base is attached to any given strand during a single cycle. The flow cell is exposed to an excitation light, exciting the label to fluoresce. As the cloned strands are clustered together, the fluorescent signal of any one fragment is amplified by the signal from its cloned counterpart, so that the fluorescence of the clusters can be recorded by the imager. After the flow cell is imaged, the blocker is cleaved and washed from the flowed-through nucleotides, more nucleotides are flooded over the flow cell, and the cycle is repeated. With each flow cycle, one or more images are recorded.
[0005] A base calling algorithm is applied to the recorded images to "read" the successive signals from each cluster and convert the light signals into the identification of the nucleotide base sequence attached to each fragment. Accurate base calling requires accurate identification of the cluster centers to ensure that successive signals are assigned to the correct fragment. Summary of the Invention
[0006] Provided herein are system, apparatus, article, method and / or computer program product aspects, and / or combinations and sub-combinations thereof, that computationally improve the resolution of an imager beyond its physical resolution limit and / or provide more accurate source location within an image.
[0007] As a particular application of such, aspects of a method and system for identifying a set of base calling locations within a flow cell are described. These include capturing flow cell images after each flow cycle and identifying candidate cluster centers in at least one of the flow cell images. An intensity is determined for each candidate cluster center. A purity is determined for each candidate cluster center based on the intensity. In some aspects, the intensity and / or purity is determined at a sub-pixel level. Each candidate cluster center that has a purity higher than the purity of the surrounding candidate cluster centers within a distance threshold is added to the set of base calling locations. The set of base calling locations may be referred to herein as a template.
[0008] In some embodiments, identifying the candidate cluster centers comprises labeling each pixel of the flow cell image as a candidate cluster center.
[0009] In some aspects, identifying candidate cluster centers includes detecting a set of potential cluster center locations using a spot search algorithm and then identifying additional cluster locations around each potential cluster center location.
[0010] Other aspects include corresponding computer systems, apparatus, and computer program products stored on computer storage device(s) configured to perform the actions or operations of the methods, either alone or in combination. For a computer system configured or configured to perform an operation or action, the computer system has installed thereon software, firmware, hardware, or a combination thereof that, during operation, causes the computer system to perform the operation or action. For a computer program product configured or configured to perform an operation or action, the computer program product includes instructions that, when executed by a hardware processor, cause the hardware processor to perform the operation or action.
[0011] Further aspects, features, and advantages of the present disclosure, as well as the structure and operation of the various embodiments of the present disclosure, are described in detail below with reference to the accompanying drawings.
[0012] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate aspects of the present disclosure and, together with the description, further serve to explain the principles of the present disclosure and to enable a person skilled in the art(s) to make and use the aspects. [Brief description of the drawings]
[0013] [Figure 1] FIG. 1 shows a block diagram of a system for identifying cluster positions on a flow cell, according to some embodiments. [Diagram 2] 1 illustrates an exemplary flow cell image with candidate cluster centroids, according to some embodiments. [Diagram 3] 1 is a flowchart illustrating a method for identifying positions to perform base calling, according to some embodiments. [Figure 4] FIG. 1 illustrates a block diagram of a computer that may be used to implement various aspects of the disclosure, according to some aspects. [Figure 5A]1 is a schematic diagram of an exemplary linear single-stranded library molecule (500) including a surface pinning primer binding site (520), a left index sequence (560) with an optional k-mer sequence (561), a forward sequencing primer binding site (540), an insert region (510) with a sequence of interest, a reverse sequencing primer binding site (550), a right index sequence (570) with an optional k-mer sequence (571), and a surface capture primer binding site (530). According to some embodiments, the library molecule (500) can have an optional right unique identification sequence between the right index sequence (570) and the surface capture primer binding site (530). [Figure 5B]FIG. 5C is a schematic diagram of an exemplary linear single-stranded library molecule (500 in FIG. 5A) that hybridizes to a double-stranded splint molecule (580), thereby circularizing the library molecule to form a library-sprint complex (590) with two nicks. The library molecule (500 in FIG. 5A) can include one or more selected from a first added left universal adaptor sequence, a first left universal adaptor sequence (520), a first left splice adaptor sequence, a left index sequence (560), a second left splice adaptor sequence, a second left universal adaptor sequence (540), a third left splice adaptor sequence, a sequence of interest (510), a third right splice adaptor sequence, a second right universal adaptor sequence (540), a second right splice adaptor sequence, a right index sequence (570), a first right splice adaptor sequence, a first right unique identification sequence, a first right universal adaptor sequence (530), and a first added right universal adaptor sequence. The double-stranded splint molecule 580 includes a first splint strand hybridized to a second splint strand. The first splint strand comprises a first sequence (320) that hybridizes to a sequence on one end of the linear single-stranded library molecule, and a second sequence (330) that hybridizes to a sequence on the other end of the linear single-stranded library molecule. The inner region (310) of the first splint strand hybridizes to the second splint strand. For simplicity, the library-sprint complex (580) does not show any splice adapter sequences or any added universal adapter sequences. Those skilled in the art will recognize that the linear library molecule (500) can include any one of the splice adapters, or any combination of two or more of the splice adapters, with or without one or both of the added universal adapter sequences. Those skilled in the art will recognize that the library-sprint complex (590), according to some embodiments, can include any one of the splice adapters, or any combination of two or more of the splice adapters, with or without one or both of the added universal adapter sequences present in the library molecule (500). [Figure 6] 1 is a flowchart illustrating a method for identifying positions to perform base calling, according to some embodiments. [Figure 7] Schematic diagrams of various exemplary configurations of multivalent molecules. Left (Class I): Schematic diagram of a multivalent molecule having a "starburst" or "helter skelter" configuration. Center (Class II): Schematic diagram of a multivalent molecule having a dendrimer configuration. Right (Class III) Schematic diagram of a number of multivalent molecules formed by reacting streptavidin with 4-arm or 8-arm PEG-NHS bearing biotin and dNTPs. According to some embodiments, the nucleotide unit is represented as "N", biotin is represented as "B", and streptavidin is represented as "SA". [Figure 8] FIG. 1 is a schematic diagram of an exemplary multivalent molecule comprising a generic core attached to multiple nucleotide arms, according to some embodiments. [Figure 9] FIG. 1 is a schematic diagram of an exemplary multivalent molecule comprising a dendrimer core attached to multiple nucleotide arms, according to some embodiments. [Figure 10] FIG. 1 shows a schematic diagram of an exemplary multivalent molecule according to some embodiments, comprising a core attached to multiple nucleotide arms, the nucleotide arms comprising biotin, a spacer, a linker, and a nucleotide unit. [Figure 11] FIG. 1 is a schematic diagram of an exemplary nucleotide arm comprising a core attachment moiety, a spacer, a linker, and a nucleotide unit, according to some embodiments. [Figure 12] 1 shows the chemical structure of an exemplary spacer (top) and various exemplary linkers (bottom), including an 11-atom linker, a 16-atom linker, a 23-atom linker, and an N3 linker, according to some embodiments. [Figure 13] 1 shows the chemical structures of various exemplary linkers, including linkers 1-9, according to some embodiments. [Figure 14] 1 shows the chemical structures of various exemplary linkers linked / attached to nucleotide units, according to some embodiments. [Figure 15]1 shows the chemical structures of various exemplary linkers linked / attached to nucleotide units, according to some embodiments. [Figure 16] 1 shows the chemical structures of various exemplary linkers linked / attached to nucleotide units, according to some embodiments. [Figure 17] 1 shows the chemical structures of various exemplary linkers linked / attached to nucleotide units, according to some embodiments. [Figure 18] 1 shows the chemical structure of an exemplary biotinylated nucleotide arm, in which, according to some embodiments, the nucleotide unit is connected to the linker via a propargylamine attachment at the 5-position of the pyrimidine base or the 7-position of the purine base. [Figure 19] A schematic diagram of one embodiment of a low binding solid support of the present disclosure is provided, which according to some embodiments comprises a glass substrate and alternating layers of a hydrophilic coating covalently or non-covalently adhered to the glass, and further comprises chemically reactive functional groups that serve as attachment sites for oligonucleotide primers.
[0014] In the drawings, like reference numbers generally indicate the same or similar elements. Additionally, the leftmost digit(s) of a reference number generally identifies the drawing in which the reference number first appears. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0015] Aspects of systems, devices, products, methods and / or computer program products, and / or combinations and subcombinations thereof are provided herein that computationally improve the resolution of an imager beyond its physical resolution limit and / or provide more accurate source location within an image. The image processing techniques described herein are particularly useful for base calling in next generation sequencing, and base calling is used as the primary example herein to illustrate the application of these techniques. However, such image analysis techniques may also be particularly useful in other applications in which spot finding and / or CCD imaging is used. For example, identifying the actual center (e.g., source location) of a perceived light signal has utility in many other fields such as location detection and tracking, astronomical photography, heat mapping, etc. In addition, techniques such as those described herein may be useful in any other application that benefits from computationally increasing the resolution once the physical resolution limit of the imager is reached.
[0016] In DNA sequencing, identifying the centers of clusters or polonies (often formed on beads) is sometimes called primary analysis. Primary analysis involves the formation of a template for the flow cell. The template contains the estimated positions of all detected clusters in a common coordinate system. The template is generated by identifying the cluster positions in all images in the first few flows of the sequencing process. The images can be aligned across all images to provide a common coordinate system. Cluster positions from different images can be merged based on their proximity in the coordinate system. Once the template is generated, all further images are aligned to it and sequencing is performed based on the cluster positions in the template.
[0017] Various algorithms exist for identifying cluster centers in an image. These existing algorithms have several shortcomings. As discussed above, cluster centers may appear to be merged if they are close to each other. The proximity may be due to precision issues or registration issues. Thus, different clusters may be treated as a single cluster, leading to improper identification of sequences or even missing sequences.
[0018] In addition, the algorithm may need to find clusters across several images to identify cluster locations for the template, which may require excessive processing time.
[0019] 1 shows a block diagram of a system 100 for identifying cluster locations on a flow cell, according to one embodiment. The system 100 has a sequencing system 110 that may include a flow cell 112, a sequencer 114, an imager 116, data storage 122, and a user interface 124. The sequencing system 110 may be connected to a cloud 130. The sequencing system 110 may include one or more of a dedicated processor 118, a field programmable gate array(s) (FPGA(s)) 120, and a computer system 126.
[0020] In some embodiments, the flow cell 112 is configured to capture DNA fragments and form a DNA sequence for base calling on the flow cell. The sequencer 114 may be configured to flow the nucleotide mixture over the flow cell 112, cleave blockers from the nucleotides during the flowing step, and perform other steps to form a DNA sequence on the flow cell 112. The nucleotides may have fluorescent elements that emit light or energy at wavelengths indicative of the type of nucleotide. Each type of fluorescent element may correspond to a particular nucleotide base (e.g., A, G, C, T). The fluorescent elements may emit light in visible wavelengths.
[0021] For example, each nucleotide base may be assigned a color. Adenine may be, for example, red, cytosine may be blue, guanine may be green, and thymine may be yellow. The color or wavelength of the fluorescent element of each nucleotide may be selected such that the nucleotides are distinguishable from one another based on the wavelength of light emitted by the fluorescent element.
[0022] The imager 116 can be configured to capture an image of the flow cell 112 after each flowing step. In one embodiment, the imager 116 is a camera configured to capture digital images, such as a CMOS or CCD camera. The camera can be configured to capture images at the wavelength of the fluorescent elements attached to the nucleotides.
[0023] The resolution of the imager 116 controls the level of detail in the flow cell image, including the pixel size. In existing systems, this resolution is very important because it controls the accuracy with which the spot search algorithm identifies cluster centers. One way to increase the accuracy of the spot search is to improve the resolution of the imager 116 or to improve the processing performed on the image captured by the imager 116. The methods described herein may detect cluster centers in pixels other than those detected by the spot search algorithm. These methods allow for improved accuracy of cluster center detection without increasing the resolution of the imager 116. The imager resolution may be smaller than existing systems with comparable performance, which may reduce the cost of the sequencing system 110.
[0024] In one aspect, images of the flow cell may be captured in groups, with each image in the group taken at wavelengths or spectra that correspond to or include only one of the fluorescent elements, hi another aspect, the images may be captured as a single image that captures all wavelengths of the fluorescent elements.
[0025] The sequencing system 100 may be configured to identify cluster locations on the flow cell 112 based on the flow cell image. The processing to identify the clusters may be performed by a dedicated processor 118, an FPGA(s) 120, a computing system 126, or a combination thereof. Identifying or determining the cluster locations may involve performing traditional cluster detection in combination with the cluster detection methods more specifically described herein.
[0026] A general purpose processor provides an interface for running a variety of programs within an operating system, such as Windows™ or Linux™, which typically provide great flexibility to the user.
[0027] In some aspects, the dedicated processors 118 may be configured to perform the steps of the cluster detection methods described herein. These may not be general-purpose processors, but rather custom processors with specific hardware or instructions to perform those steps. The dedicated processors execute specific software directly without an operating system. The lack of an operating system reduces overhead at the expense of flexibility in what the processor may execute. The dedicated processors may utilize custom programming languages that may be designed to operate more efficiently than software running on a general-purpose processor. This may increase the speed at which steps are executed and allow for real-time processing.
[0028] In some aspects, the FPGA(s) 120 may be configured to perform the steps of the cluster detection method described herein. The FPGA is programmed as hardware to perform only specific tasks. Special programming languages may be used to translate software steps into hardware components. Once the FPGA is programmed, the hardware directly processes the provided digital data without executing software. Instead, the FPGA uses logic gates and registers to process the digital data. Because there is no overhead required for an operating system, FPGAs generally process data faster than general-purpose processors. As with special-purpose processors, this comes at the expense of flexibility.
[0029] Also, without software overhead, an FPGA can run faster than a dedicated processor, although this depends on the exact operations being performed and the particular FPGA and dedicated processor.
[0030] A group of FPGA(s) 120 may be configured to perform steps in parallel. For example, several FPGA(s) 120 may be configured to perform a processing step on an image, a set of images, or cluster locations within one or more images. Each FPGA(s) 120 may perform its own portion of the processing step simultaneously, reducing the time required to process the data. This allows the processing step to be completed in real time. Further discussion of the use of FPGAs is provided below.
[0031] By performing the processing steps in real time, the system may use less memory since the data may be processed as it is received, which is an improvement over conventional systems that may need to store the data before it can be processed, which may require more memory or access to computer systems located in the cloud 130.
[0032] In some embodiments, data storage 122 is used to store information used in identifying the cluster locations. This information may include the image itself or information derived from the image captured by imager 116. The DNA sequence determined from base calling may be stored in data storage 122. Parameters identifying the cluster locations may also be stored in data storage 122.
[0033] The user interface 124 may be used by a user to operate the sequencing system or to access data stored in the data storage 122 or the computer system 126 .
[0034] The computer system 126 may control the general operation of the sequencing system and may be coupled to the user interface 124. It may also perform the steps of cluster location identification and base calling. In some embodiments, the computer system 126 is computer system 400, as described in more detail in FIG. 4. The computer system 126 may store information regarding the operation of the sequencing system 110, such as configuration information, instructions for operating the sequencing system 110, or user information. The computer system 126 may be configured to pass information between the sequencing system 110 and the cloud 130.
[0035] As discussed above, the sequencing system 110 may have a dedicated processor 118, an FPGA(s) 120, or a computer system 126. The sequencing system may use one, two, or all of these elements to accomplish the necessary processing described above. In some embodiments, when these elements are present together, processing tasks are divided between them. For example, the FPGA(s) 120 may be used to perform the cluster center detection method described herein, while the computer system 126 may perform other processing functions for the sequencing system 110. Those skilled in the art will appreciate that various combinations of these elements enable various system embodiments that balance efficiency and speed of processing with the cost of the processing elements.
[0036] Cloud 130 may be a network, remote storage, or some other remote computing system separate from sequencing system 110. Connection to cloud 130 may allow access to data stored outside of sequencing system 110 or may allow updates to software within sequencing system 110.
[0037] FIG. 3 is a flow chart illustrating a method 300 for identifying actual cluster center locations for performing base calling. Cluster centers in a flow cell image are locations in the image that correspond to the location of cloned clusters on the physical flow cell. The wavelength of the optical signal detected at the cluster center correlates to the nucleotide base added to the fragment on the flow cell at that location. For the DNA sequence to be correctly determined, successively detected optical signals must be consistently attributed to the correct DNA fragment. Thus, accurately identifying the location of the cluster center improves the base calling accuracy of that fragment. In some embodiments, once the actual cluster center is identified, such location can be mapped onto a template for use in subsequent base calling cycles using the same flow cell. Method 300 can be performed by a dedicated processor 118, FPGA(s) 120, or computer system 126.
[0038] In step 310, a flow cell image is captured. As discussed above, the flow cell image may be captured by the imager 116. Step 310 may involve capturing images one at a time to be processed by the following steps, or may involve capturing a set of images for simultaneous processing. In examples where a set of images is captured, each image in the set of images may correspond to a different detected wavelength. For example, considering the above color designations associated with nucleotides, the set of images may include four images, each corresponding to a signal captured at one of the red, blue, green, and yellow wavelengths. In examples where a single image is captured, the image may include all detected wavelengths of interest. Each image or set of images may be captured for a single flow step on the flow cell. In some embodiments, the flow cell images are captured with reference to a coordinate system.
[0039] 2 shows a schematic diagram of a flow cell image 200 in which signals from clusters are present. The flow cell image 200 is composed of pixels 210, such as pixels 210A, 210B, and 210C. During step 310, the imager records the light signal received from the flow cell after excitation of, for example, fluorescent elements bound to fragments on the flow cell, such fragments being located within a cloned cluster of fragments.
[0040] In step 320, the locations of potential cluster centers within the flow cell image are identified. For example, in some embodiments, the optical signal imaged in step 310 may be input to a spot finder algorithm such that the spot finder algorithm outputs a set of potential cluster centers. In some embodiments, potential cluster centers may be identified using only a single flow cell image (e.g., one image including all wavelengths of interest). In some other embodiments, potential cluster centers may be identified from a set of images from a single running cycle on the flow cell (e.g., one image at each wavelength of interest). Using only a flow cell image or set of flow cell images from a single running cycle advantageously reduces the amount of processing time because the spot finder algorithm does not need to wait for additional images from future running cycles to be acquired.
[0041] In yet other embodiments, the spot search algorithm may be applied to images from two or more flow cycles, and potential cluster centers may be searched for using some combination of those images. For example, potential cluster centers may be identified by the presence of spots occupying the same location in images from two or more flow cycles.
[0042] Potential cluster center locations identified by the spot finding algorithm are indicated by "X's" in FIG. 2, such as potential cluster center locations 220A, 220B, and 220C. Due to the random nature of fragment attachment to the flow cell, some of the cloned clusters may be close to each other, while others may be further apart or may be alone. As a result, some "X's" in FIG. 2 are located closer to each other than others. In addition, some pixels may be identified as containing potential cluster centers, while other pixels are not. For example, pixel 210A may be identified as containing potential cluster center location 220A, while pixel 210B is not initially identified as containing potential cluster center location 220.
[0043] In some embodiments, the spot finding algorithm may identify potential cluster center locations 220 at sub-pixel resolution by interpolating across pixels 210. For example, potential cluster center location 220A is not located at the center of pixel 210A, but is located on the lower right side of pixel 210A. Other potential cluster center locations may be located in different regions of their respective pixels 210. For example, potential cluster center location 220B is located at the upper right side of the pixel, and potential cluster center location 220C is located at the upper left side of the pixel. The interpolation may be performed by an interpolation function.
[0044] In some embodiments, the interpolation function is a Gaussian interpolation function known to those skilled in the art. Sub-pixel resolution may allow potential cluster locations to be determined, for example, at a tenth pixel resolution, although other resolutions are also contemplated. In embodiments, for example, the resolution may be a quarter pixel resolution, a half pixel resolution, etc. The interpolation function may be configured to determine this resolution.
[0045] An interpolation function may be used to fit the light intensity at one or more pixels 210. This interpolation allows for the identification of sub-pixel locations. The interpolation function may be applied across a set of pixels 210 that include potential cluster center locations 220. In one aspect, the interpolation function may be fitted to a pixel 210 that has a potential cluster center location 220 therein, as well as surrounding pixels 210 that touch the edges of that pixel 210.
[0046] In some embodiments, the interpolation function may be determined at several points in the image. The resolution determines the number of points located at each pixel 210. For example, if the resolution is one tenth of a pixel, then nine points are calculated across the pixel 210 and one point at each edge along a line perpendicular to the pixel edge, dividing the pixel 210 into ten portions. In some embodiments, the interpolation function is calculated at each point and the difference between the interpolation function and the pixel intensity at each point is determined. The center of the interpolation function is shifted to minimize the difference between the interpolation function and the intensity of each pixel 210. This sub-pixel interpolation allows the system to achieve higher resolution with a lower resolution imager, reducing the cost and / or complexity of the system.
[0047] In some embodiments, the interpolation may be performed on a 5 x 5 grid. The grid may be centered on pixel centers 210 within the pixels having potential cluster center locations.
[0048] Some embodiments of step 320 use a spot-finding algorithm to identify potential cluster center locations, while some other embodiments of step 320 first identify each pixel in the captured flow cell image as a potential cluster center location. For example, in FIG. 2, all pixels 210 may be identified as potential cluster center locations 220. This approach eliminates the need for a spot-finding algorithm, which may simplify the type of processing required to perform method 300. This approach may be advantageous when massively parallel processing is available, since each potential cluster center location may be processed in parallel. This may reduce processing time, at the potential cost of additional hardware, such as increased dedicated processor 118 or FPGA(s) 120. In some embodiments, an interpolation function may then be used as described above to identify intensities at sub-pixel resolution throughout the flow cell image.
[0049] As discussed above, a cluster center identifies a location in an image, such as pixel 210, that corresponds to the location of a cloned cluster on a physical flow cell. Potential cluster center locations 220 are locations in an image where one or more wavelengths of light are detected by an imager. In some cases, while a cluster's physical location corresponds to one set of pixels 210, it is possible that the light signal from that cluster overflows into additional pixels 210 adjacent to the one set, for example, due to saturation (also referred to as "blooming") of the corresponding sensor in the camera. In addition, when clusters are placed closely together, the light signals from the clusters may overlap even if the clusters themselves do not. Identifying the true cluster center allows the detected signals to be attributed to the correct DNA fragments, thus improving the accuracy of the base calling algorithm.
[0050] Thus, in step 325, additional cluster signal locations 225 are identified around each potential cluster center location 220. These are indicated by "+" in Figure 2, e.g., 225A, 225B, and 225C. These additional cluster signal locations 225 correspond to other locations within the flow cell that may constitute cluster centers, instead of, or in addition to, the locations already identified as potential cluster center locations 220.
[0051] In some embodiments, additional cluster locations 225 are placed around potential cluster center locations 220. In some embodiments, these additional cluster locations 225 are not initially identified by the spot search algorithm, but are placed in a pattern around each potential cluster center location 220 identified by the spot search algorithm. The additional cluster locations 225 do not represent actual detected cluster centers, but rather potential locations for checking cluster centers that might not otherwise be detected. Such cluster centers may not be detected due to mixing between signals from nearby cluster centers, errors in the spot search algorithm, or other effects.
[0052] As an example, the additional cluster location 225A may be placed at pixel 210A based on the location of the potential cluster center location 220A. In this context, this may mean that the additional cluster location 225A is placed a pixel width away from the potential cluster center location 220A. Other additional cluster locations 225 are also placed around the potential cluster center location 220A. It should be understood that the additional cluster location 225 is not placed where no pixel 210 exists, such as when the potential cluster center location 220 is near an edge of the flow cell image 200.
[0053] In some embodiments, the additional cluster signal locations 225 are arranged in a grid centered on the potential cluster center location 220. The additional cluster locations 225 may be separated from each other and from the potential cluster center location 220 by a pixel width. In some embodiments, the additional cluster locations 225 and the potential cluster center location 220 form a square grid. The grid may have an area of 5 pixels by 5 pixels, 9 by 9, 15 by 15, or other dimensions.
[0054] In some embodiments, potential cluster center locations 220 may be close enough to one another to overlap corresponding grids of additional cluster locations 225. This may result in the same pixel 210 including additional cluster locations 225 from more than one potential cluster center location 220, and more than one additional cluster location 225 being attributed to the same pixel 210. For example, pixel 210C includes additional cluster location 225B (identified based on potential cluster center location 220B) as well as additional cluster location 225C (identified based on potential cluster center location 220C).
[0055] In some aspects, if the additional cluster locations 225 within the same pixel 210 are sufficiently close to one another, one of the additional cluster locations 225 is discarded and the other is used to represent both. Two additional cluster locations 225 may be considered sufficiently close to one another for such processing if they are within, for example, but not limited to, two tenths of a pixel, one tenth of a pixel, or some other sub-pixel distance.
[0056] One skilled in the art will understand that if, in step 320, all pixel locations have been identified as potential cluster center locations 220, then step 325 may be skipped because no other pixels remain in the flow cell image to consider in addition to the identified potential cluster center locations.
[0057] In some embodiments, the potential cluster center locations 220 (and surrounding additional cluster locations 225, if identified) together constitute a set of all candidate cluster centers, which can be processed to identify actual cluster centers for each captured flow cell image.
[0058] Thus, once the potential cluster center locations 220 and their surrounding grid of additional cluster locations 225 (i.e., candidate cluster centers) have been identified, they may be used as a starting point for determining the actual location of the cluster center, which may or may not be the same as the originally identified potential cluster center location 220.
[0059] 3, a purity value for each candidate cluster center on each captured flow cell image is determined in step 330. The purity value may be determined based on the wavelength of the fluorescent element attached to the nucleotide and the intensity of the pixel in the captured flow cell image.
[0060] At each candidate cluster center, the intensity of the pixel is the combination of energy or light emitted across the spectral bandwidth of the imager. In some embodiments, the amount of energy or light corresponding to the fluorescence spectral bandwidth of each nucleotide base can be searched. The purity of each signal corresponding to a particular nucleotide base can be searched as the ratio of the amount of energy for one nucleotide base signal to the total amount of energy for each other nucleotide base signal (e.g., the purity of a "red" signal can be determined based on the relative intensity of the red wavelength detected for that pixel or subpixel compared to each of the blue, green, and yellow wavelengths detected). The overall purity of a pixel can be the maximum ratio, the minimum ratio, the average ratio, or the median ratio. The calculated purity can then be assigned to that pixel or subpixel.
[0061] As described above, a set of flow cell images can be captured for a single flow cycle. Each image in the set is captured at a different wavelength, each wavelength corresponding to one of the fluorescent elements attached to the nucleotides. The purity of a given cluster center across the set of images can be the highest, lowest, median, or average purity from the set of purities for the set of images.
[0062] In some embodiments, the purity of a candidate cluster center may be determined as the ratio of the second highest intensity or energy from a wavelength of a pixel minus the highest intensity or energy from that wavelength of the pixel. A threshold may be set for what constitutes high purity or low purity. For example, the highest purity may be 1, and low purity pixels may have a purity value close to zero. A threshold may be set in between.
[0063] In some embodiments, the ratio of the two intensities may be modified by adding an offset to both the second highest intensity and the highest intensity. The offset may provide improved accuracy of the quality score. For example, in some cases, the two ratio intensities may differ by a small amount of the absolute maximum intensity (which is also a high percentage of the highest intensity). As a specific non-limiting example, the highest intensity may be 10 and the lowest intensity may be the maximum possible intensity of 1000. The ratio in this case would be 0.1 and the purity would be 0.9. If there was no more, this could be read as a high quality score. This is in contrast to an intensity of, for example, 500 for the highest intensity and 490 for the second highest intensity. This example has almost the same absolute difference, but the ratio is close to 1 and the purity is close to zero. In the first case, the purity is misleading, since the low overall intensity suggests that there is no polony. In the second case, the purity is more accurate, indicating that the pixel is displaying intensities or energies from two different cluster centers located nearby.
[0064] An offset can be a value added to the strength of the ratio to solve such problems. For example, if the offset is 10 percent of the maximum amplitude, then in the example above, the offset would be 100 and the first ratio would be 101 over 110, which is much closer to 1, resulting in a purity near zero that accurately reflects the small delta between the two wavelength intensities. For the second ratio, the ratio would be 600 over 590, which is still closer to 1, also resulting in a purity near zero.
[0065] As another example incorporating an offset, if the highest intensity is 800 and the lowest intensity is 1, the purity without the offset is close to 1 since the ratio is near zero. If the offset is again 100, the ratio becomes 101 for 900. This slightly reduces the purity from 1 to about 0.89. While this may reduce the purity, the calculated purity is still high. The offset value may be set to reduce this effect. For example, the offset in another case may be 10. Using the previous example where the highest intensity is 10 and the lowest intensity is 1, the purity would be 1-10 / 11, or about 0.09, which accurately reflects the small difference between the intensities. In the example where the highest intensity is 600 and the lowest intensity is 590, the purity would be 1-600 / 610, or about 0.016, which again reflects the small difference between the intensities. In the example where the highest intensity is 800 and the lowest intensity is 1, the purity is 1-11 / 810, or about 0.99, which is a much smaller drop in purity and reflects a larger difference between the intensities.
[0066] In step 340, actual cluster centers are identified based on the purity values calculated for the candidate cluster centers for the flow cell. The actual cluster centers may be a subset of the candidate cluster locations identified in steps 320 and 325.
[0067] In some embodiments, the actual cluster center is identified by comparing the purity of each candidate cluster center across the entire flow cell image with nearby candidate cluster centers in the same image. In some embodiments, considering two candidate cluster centers being compared, the candidate cluster center with the higher purity is kept. In some embodiments, the candidate cluster center is only compared with other candidate cluster centers within a certain distance. For example, this distance threshold can be based on pixel size and the size of a cluster of a given nucleotide.
[0068] For example, if the average size of a cluster is four pixels wide / height, the distance threshold could be two pixels wide / height, since any candidate cluster centers within two pixel widths / heights of each other are more likely to belong to the same cluster or to have a higher intensity (indicating that the candidate cluster center is actually on the edge of two separate clusters).
[0069] In some embodiments where purity is calculated over multiple flow cell cycles, determining that a candidate cluster center consistently has a higher purity than surrounding candidate cluster centers across multiple flow cell images may further increase the likelihood that the location is an actual cluster center. Lower purity may indicate that the signal detected at the candidate cluster center is not an actual cluster center, but rather noise, a mixture of other signals, or some other phenomenon.
[0070] In step 350, the actual cluster centers are used to perform base calling on the flow cell image. For example, a wavelength detected from the actual cluster center may be determined, which correlates to a particular nucleotide base (e.g., A, G, C, T). That nucleotide base is then recorded as being added to the sequence corresponding to the actual cluster center.
[0071] By successive repetitions of the flow cycle and fluorescence wavelength discrimination at actual cluster centers in successive flow cell images, a sequence of DNA fragments corresponding to each actual cluster center on the flow cell can be constructed.
[0072] In some embodiments, a template is formed from the actual cluster centers identified for a single flow cycle. The actual cluster locations in the template can then be used to identify locations to perform base calling in images from subsequent flow cycles.
[0073] Flow cell images captured at different flow cycles may have alignment issues due to shifts in the position of the flow cell or imager between flow cycles. Thus, in some embodiments, step 350 may include an alignment step to properly align consecutive images. This ensures that the actual clusters accurately map the centers of the templates to the same positions on each flow cell image, thus improving the accuracy of base calling.
[0074] In some aspects where templates are used to identify actual cluster centers in subsequent images, only data corresponding to relevant locations in those subsequent images need to be maintained and / or processed. This reduction in the amount of data processed increases the speed and / or efficiency of processing such that accurate results can be obtained more quickly than with legacy systems. Additionally, the reduction in the amount of data stored reduces the amount of storage required for the sequencer, thus reducing the amount and / or cost of resources required.
[0075] In addition, some legacy systems require different images to be compared to each other to identify cluster locations. This comparison may involve applying a spot finder algorithm to images from multiple flow cell cycles and then comparing the spot finder results across images. This may require storing images or spot finder results for each of multiple flow cycles. Method 300 may improve processing and storage efficiency of cluster finders because images do not need to be compared directly. Instead, images may be processed in real time and only purity information and / or final template locations need to be stored.
[0076] The sequencing flow cycles and image creation processes often run faster than the spot finding and base calling programs that analyze the images, and this difference in execution time may require storing the flow cell images after each flow cycle, or delaying the sequencing flow cycle while waiting for some or all of the image analysis processes to complete.
[0077] The use of FPGAs can increase processing speed without sacrificing accuracy. Implementing some or all of the processes described herein on an FPGA can reduce processor overhead and enable parallel processing on the FPGA. For example, each possible cluster location may be processed by a different FPGA, or by a single FPGA configured to process the possible cluster locations simultaneously in parallel. If properly implemented, this can enable real-time processing. Real-time processing has the advantage that images can be processed as they are generated. The FPGA is ready to process the next image by the time the sequencing system prepares the flow cell. The sequencing system does not have to wait for post-processing, and the entire process of primary analysis can be completed in a matter of minutes. In addition, since the entire image is processed as it is received, the only information that needs to be stored is the data to perform base calling. Instead of storing the entire image, only the purity or intensity of a particular pixel needs to be stored. This greatly reduces the need for data storage or remote storage of images in the sequencing system.
[0078] In some embodiments, the entire process, including image registration, intensity extraction, purity calculations, base calling, and other steps, is performed by an FPGA, which may provide the most compact implementation and may provide the speed and throughput required for real-time processing.
[0079] In some embodiments, processing responsibilities are shared, such as between the FPGA and an associated computer system. For example, in some embodiments, the FPGA may handle image registration, intensity extraction, and purity calculations. The FPGA then passes information to the computer system for base calling. This approach balances the load between the FPGA and computer system resources, including scaling down communication between the FPGA and the computer system. It also provides flexibility for software on the computer system that handles base calling with rapid algorithm tune-up capabilities. Such an approach may provide real-time processing.
[0080] Those skilled in the art will recognize that different configurations of FPGAs, special purpose processors, and computer systems can be used to perform the various steps. The selection of a given configuration may be based on the flow cell image size, imager resolution, number of images to process, desired accuracy, and required speed. Implementation and hardware costs of the FPGAs, special purpose processors, and computer systems may also influence the selection of the configuration.
[0081] As a non-limiting example of a performance comparison between existing methods and embodiments of the methods described herein, tests were performed on two exemplary flow cells, one with a low density of clusters and one with a high density of clusters. For comparison, tests were performed using each method to target a particular average error rate of false positives on the identified clusters.
[0082] For the low density flow cell, the average error rate was 0.3%. The existing method identified about 78,000 cluster centers, whereas the method described herein identified about 98,000 cluster centers. For the high density flow cell, the average error rate was 1.1%. The existing method identified about 63,000 cluster centers, whereas the method described herein identified about 170,0000 cluster centers.
[0083] The results suggest that the method described herein effectively identifies more clusters than existing methods. Moreover, even with the same error rate, when the density of clusters on the flow cell increases, the method disclosed herein performs even better, identifying nearly three times as many clusters. In some embodiments, this may allow the flow cell to run at a higher density without the performance loss typically experienced with existing methods.
[0084] 6 shows a flow chart of an exemplary embodiment of a method 600 for identifying base calling positions in a primary analysis of NGS data analysis. Method 600 may include some or all of the operations disclosed herein. The operations may be performed in the order described herein, but are not limited to such.
[0085] The method 600 may be performed by one or more processors (e.g., 404 of FIG. 4) disclosed herein. In some aspects, the processor may include one or more of a processing unit, an integrated circuit, or a combination thereof. For example, the processing unit may include a central processing unit (CPU) and / or a graphics processing unit (GPU). The integrated circuit may include a chip such as a field programmable gate array (FPGA). In some aspects, the processor may include the computing system 400.
[0086] In some aspects, some or all of the operations in method 600 may be performed by FPGA(s). In aspects in which some operations are performed by FPGA(s), data after the operations performed by the FPGA(s) may be communicated by the FPGA(s) to the CPU(s) so that the CPU(s) can use such data to perform subsequent operation(s) in method 500. Similarly, data may also be communicated from the CPU(s) to the FPGA(s) for processing by the FPGA(s). In some aspects, all of the operations in method 500 may be performed by the CPU(s). Alternatively, the operations performed by the CPU(s) may be performed by other processors, such as dedicated processors or GPU(s). In some aspects, all of the operations in method 600 may be performed by FPGA(s).
[0087] In some embodiments, some or all of the operations of method 600 are performed during or before sequencing cycle N in a sequencing run. Base calling templates, e.g., polony maps, may be generated in some or all of cycles 1-N. Polonies or clusters from one or more channels in such cycles may be included in the template in a reference coordinate system, but flow cell images for cycle N and / or subsequent cycles have not yet been captured or are not currently being captured. In some embodiments, cycle N is the current cycle. N may be any non-zero integer, e.g., 3, 4, or 5.
[0088] Methods for multiple library molecules In some embodiments, the method 600 can include an operation 610 of providing a first plurality of library molecules immobilized on a support. Each of the first plurality of library molecules can include a first insert sequence derived from a first sample source and a first sample index sequence. The first sample index sequence can include a first k-mer sequence and a first universal sample index sequence. The first universal sample index can be configured for unique identification of the first sample source of the first insert sequence. In some embodiments, the first plurality of library molecules are derived from the same sample source.
[0089] 5A shows an exemplary embodiment of a library molecule 500 disclosed herein. The library molecule 500 can include a first insert sequence 510, i.e., a first sequence of interest, obtained from a first sample, and a first sample index sequence having at least a first universal sample sequence 570. The first sample index sequence can also include a k-mer sequence 571.
[0090] In some embodiments, each library molecule 500 may contain only a single insert sequence 510. In other embodiments, each library molecule 500 may contain multiple insert sequences, either immediately adjacent to each other or separated by some other portion of the library molecule.
[0091] In some embodiments, different library molecules 500 immobilized on a support may have insert sequences of the same size, e.g., an insert of 120 nucleotide bases. In other embodiments, different library molecules 500 immobilized on a support may have insert sequences of different sizes.
[0092] In some embodiments, the method 600 includes an operation 620 of providing a second plurality of library molecules immobilized on a support. Each of the second plurality of library molecules can include a second insert sequence derived from a second sample source and a second sample index sequence. The second sample index sequence can include a second k-mer sequence and a second universal sample index sequence. The second universal sample index can be used to uniquely identify the second sample source of the second insert sequence. In some embodiments, the second sample source is different from the first sample source.
[0093] In some embodiments, each of the first plurality of library molecules further comprises a third sample index sequence having a third universal sample index sequence. The combination of the first and third universal sample index sequences can be configured to uniquely identify the first sample source of the first insert sequence. In some embodiments, the third sample index lacks a k-mer sequence or a random sequence. In some embodiments, the third sample index includes a k-mer sequence or a random sequence. The first sample index sequence can be on the left side of the first insert sequence, i.e., the 5' end, or on the right side of the insert sequence, i.e., the 3' end. The third sample index sequence can then be on the opposite end of the insert sequence.
[0094] 5A shows an exemplary embodiment of a library molecule 500 including a first insert sequence 510, i.e., a first sequence of interest, obtained from a first sample, and a third sample index sequence having at least a third universal sample sequence 560. The first sample index sequence can also include a k-mer sequence 561.
[0095] In some embodiments, each of the second plurality of library molecules further comprises a fourth sample index having a fourth universal sample index sequence. The combination of the second and fourth universal sample index sequences can uniquely identify the second sample source of the second insert sequence. In some embodiments, the fourth sample index is a random sequence or lacks a random sequence. In some embodiments, the fourth sample index comprises a k-mer sequence or a random sequence. The second sample index sequence can be on the left side of the second insert sequence, i.e., the 5' end, or on the right side of the insert sequence, i.e., the 3' end. The third sample index sequence can then be on the opposite end of the second insert sequence.
[0096] In some embodiments, each of the library molecules further comprises one or more universal adaptor sequences (eg, 520, 540 in Figures 5A-5B).
[0097] In some embodiments, method 600 can include repeating operations 610 and / or 620 for a third, fourth, fifth, sixth, or even a plurality of library molecules, each of the plurality of library molecules being derived from a different sample. For example, operation 610 can be repeated 95 times after performing operations 610 and 620 once. Each repetition can be of a plurality of library molecules from a different sample, so that 96 samples can be prepared for sequencing on the support.
[0098] For different library molecules from the same sample, each library molecule can have a different insert sequence, a different k-mer sequence, and the same universal sample index sequence(s). For different library molecules from different samples, each library molecule can have a different insert sequence, a different k-mer sequence, and a different universal sample index sequence(s).
[0099] The first, second, or other sample source of the insert sequence may be genomic DNA, double-stranded cDNA, and / or cell-free circulating DNA.
[0100] In some embodiments, the method 600 can include pooling the first and second plurality of library molecules together, distributing the pooled library molecules on a support, and performing an amplification reaction to generate a plurality of nucleic acid template molecules immobilized on the support. The plurality of nucleic acid template molecules can be clonally amplified as disclosed herein. In embodiments involving three or more samples, e.g., 96 samples, all of the plurality of library molecules can be pooled together and distributed on a support for the amplification reaction.
[0101] In some embodiments, the nucleic acid template molecules are clonally amplified from a corresponding library molecule. Each nucleic acid template molecule can have multiple clonal copies of the corresponding library molecule. In some embodiments, each of the multiple nucleic acid template molecules is clonally amplified from a corresponding splint complex (e.g., 590 in FIG. 5B) that includes a corresponding library molecule and a corresponding splint molecule / adapter (e.g., 580 in FIG. 5B). The splint molecule / adapter can be single-stranded or double-stranded (e.g., FIG. 5B). Details of the formation of splint complexes and amplification to form template molecules are disclosed in U.S. Patent Application Nos. 17 / 725,042 and 17 / 725,065, the contents of both of which are incorporated herein by reference in their entireties.
[0102] In some embodiments, each nucleic acid template molecule may have a single copy of a corresponding library molecule, and multiple template molecules can be clustered together to form a "cluster."
[0103] Each template molecule can have one or more copies of the corresponding library molecule, so that for different template molecules from the same sample, each template molecule can have a different insert sequence, a different k-mer sequence, and the same universal sample index sequence(s). For different template molecules from different samples, each template molecule can have a different insert sequence, a different k-mer sequence, and a different universal sample index sequence(s).
[0104] In some embodiments, the plurality of nucleic acid template molecules are immobilized at random positions on the support. In some embodiments, the plurality of nucleic acid template molecules are immobilized at predetermined positions on the support. Each of the plurality of nucleic acid template molecules immobilized on the support can correspond to a polony or a cloned cluster. The positions of the plurality of immobilized template molecules on the support can correspond to base calling positions. The immobilized nucleic acid template molecules can be sequenced when a sequencing reaction cycle is performed.
[0105] The density of nucleic acid template molecules on the support is 1 mm 2 About 10 per 2 ~10 12 The density of the nucleic acid template molecules on the support can be 1 mm 2 About 10 per 4 ~10 8 The density of the nucleic acid template molecules on the support can be 1 mm 2 About 10 per 4 ~10 12 The density of the nucleic acid template molecules on the support can be 1 mm 2 About 10 per 4 ~10 5 The density of the nucleic acid template molecules on the support can be 1 mm 2 10 per 2 ~10 12 The density of the nucleic acid template molecules on the support can be 1 mm 2 10 per 4 ~10 8The density of the nucleic acid template molecules on the support can be 1 mm 2 10 per 4 ~10 12 The density of the nucleic acid template molecules on the support can be 1 mm 2 10 per 4 ~10 5 It could be.
[0106] The support may include one or more substrates. The support may include a glass substrate or a plastic substrate. The support may include a transparent top substrate that is closest to the objective lens of the optical system. The support may include one or more microfluidic channels, and the template molecules are immobilized on the surface of the microfluidic channels. In some embodiments, the support is included in a flow cell device.
[0107] In some embodiments, the support or a portion thereof (e.g., the surface of a microfluidic channel) is passivated with at least one hydrophilic polymer coating having a water contact angle of 45 degrees or less. The at least one hydrophilic polymer coating may comprise a molecule selected from the group consisting of polyethylene glycol (PEG), poly(vinyl alcohol) (PVA), poly(vinylpyridine), poly(vinylpyrrolidone) (PVP), poly(acrylic acid) (PAA), polyacrylamide, poly(N-isopropylacrylamide) (PNIPAM), poly(methyl methacrylate) (PMA), poly(2-hydroxylethyl methacrylate) (PHEMA), poly(oligo(ethylene glycol) methyl ether methacrylate) (POEGMA), polyglutamic acid (PGA), polylysine, polyglucoside, streptavidin, and dextran. In some embodiments, the at least one hydrophilic polymer coating comprises a branched hydrophilic polymer molecule having at least four branches. In some embodiments, at least one hydrophilic polymer coating comprises polymer molecules having a molecular weight of at least 1000 Daltons.
[0108] In some embodiments, the k-mer sequence (e.g., 561 or 571 in FIG. 5A) can contain 1, 2, 3, 4, 5, or more nucleotide bases. In some embodiments, the k-mer sequence contains a random sequence of at least two or three nucleotide bases: A, G, C, and T / U. In some embodiments, the k-mer sequence contains only random sequences of A, G, C, and T / U nucleotide bases.
[0109] The method 600 may include performing an operation 630 of performing k sequencing reaction cycles of the first and second k-mer sequences. The operation 630 may be performed by a sequencing system disclosed herein, thereby generating a first plurality of flow cell images. In each sequencing cycle, different k-mer sequences from different template molecules of different samples may be sequenced in parallel. For example, in the first sequencing cycle, the first nucleotide bases in 96,000 k-mer sequences from 96 samples may be sequenced in parallel, thereby generating a flow cell image per channel with 1000 base calling positions.
[0110] In some embodiments, operation 630 is performed prior to performing one or more sequencing reaction cycles of the insert sequence, e.g., the first and second insert sequences.
[0111] In some embodiments, operation 630 is based on the sequencing order (i.e., read order) of the sequencing run. The sequencing order can include sequencing the k-mer sequence, then sequencing the first universal sample index sequence, then sequencing the first insert sequence. The sequencing order can include sequencing the k-mer sequence, then sequencing the first and second universal sample index sequences, then sequencing the first and second insert sequences. The sequencing order can include sequencing the first and second universal sample index sequences, then sequencing the k-mer sequence, then sequencing the first and second insert sequences. The sequencing order can include sequencing the k-mer sequence, then sequencing the first and second insert sequences, then sequencing the first and second universal sample index sequences. The sequencing order can include sequencing a number of bases (e.g., the first 3-8 bases) of the first and second insert sequences, sequencing the k-mer sequence, and then sequencing the first and second universal sample index sequences. In the sequencing order, if the third and fourth sample index sequences lack random sequences, sequencing the third and fourth sample index sequences can occur after sequencing the first and second sample index sequences.
[0112] In some embodiments, the first plurality of flow cell images are from 2, 3, 4, 5, or 6 different color channels in each sequencing cycle. In some embodiments, the first plurality of flow cell images are from k cycles or k+1 sequencing cycles, where one or more of the k or k+1 cycles include balanced diversity of A, G, C, and T / U nucleotide bases among the multiple nucleic acid template molecules immobilized on the support.
[0113] In some embodiments, each sequencing reaction cycle is polymerase-mediated. The operation 630 of performing k sequencing reaction cycles of a k-mer sequence can include contacting nucleotide acids or polonies of a plurality of nucleic acid template molecules with a plurality of nucleotide reagents that include a mixture of different types of nucleotide bases A, G, C, and T / U. In some embodiments, each individual nucleotide reagent includes a different detectable color label corresponding to each different type of nucleotide base.
[0114] In some other embodiments, operation 630 includes contacting a nucleotide acid or polony of a plurality of nucleic acid template molecules with a plurality of sequencing primers, a plurality of polymerases, and a mixture of different types of avidites, each of which includes a core having a plurality of nucleotide arms attached thereto, each arm of each of the individual avidites including the same type of nucleotide base.
[0115] In some embodiments, the operation 630 of performing k sequencing reaction cycles of the k-mer sequence includes imaging optical color signals emitted from nucleotide reagents bound to the template molecule by an optical system herein. In some embodiments, the operation 630 of performing k sequencing reaction cycles of the k-mer sequence includes acquiring a first plurality of flow cell images including optical color signals emitted from nucleotide reagents bound to the template molecule by an optical system herein for the k cycles.
[0116] A first plurality of flow cell images can be generated over k cycles corresponding to the conduction of a sequencing reaction on a k-mer sequence. The first plurality of flow cell images can include light signals emitted from nucleotide reagents bound to A, G, C, and T / U nucleotide bases of balanced diversity among a plurality of nucleic acid template molecules immobilized on a support over the k cycles.
[0117] The balanced diversity of A, G, C, and T / U nucleotide bases among a plurality of nucleic acid template molecules can include a percentage of the number of each type of nucleotide base (1) relative to the total number of bases (2) in one or more cycles. The percentage can be greater than 10%, 15%, or 20%. For example, the unbalanced diversity of nucleotide bases can include a number of nucleotide bases A, G, C, T that is 26%, 15%, 27%, and 32%, respectively, of the total number of all nucleotide bases in the template molecules of the cycle.
[0118] The method 600 may include an operation 640 of determining pixel intensities for pixels of the first plurality of flow cell images and respective color purities of each of the pixel intensities. The operation 640 may be performed by a processor as disclosed herein. Each of the pixel intensities may include an intensity of the pixel and / or one or more sub-pixel intensities of a sub-pixel corresponding to the pixel. The respective color purities of each of the pixel intensities may include one or more color purities corresponding to one of the color purities of the pixel and / or one or more sub-pixel intensities of the corresponding pixel. The respective color purities of each of the pixel intensities may include respective color purities of one or more color channels.
[0119] Determining the pixel intensities can include determining each channel intensity in a set of channel intensities, each channel corresponding to a different fluorescent wavelength, and such determination can be based on a comparison of the set of channel intensities at corresponding pixel or one or more sub-pixel locations.
[0120] Determining the color purity of each of the pixel intensities can include determining a ratio of (1) a signal corresponding to a particular type of nucleotide base to (2) a total amount of signals for other types of nucleotide bases. Determining the color purity of each of the pixel intensities can include, at least in part, operations 330 disclosed herein.
[0121] The method 600 may include an operation 650 of determining a base calling template including base calling positions based on the pixel intensities and the color purity of each of the pixel intensities determined in operation 640 .
[0122] As disclosed herein, determining the base calling template based on the pixel intensities and the color purity of each of the pixel intensities may include determining for each of the pixels or sub-pixels whether the respective color purity is greater than the color purity of the other pixels or sub-pixels within the distance threshold. In response to determining that the respective color purity is greater than the color purity of the other pixels or sub-pixels within the threshold distance, adding the corresponding pixel or sub-pixel location to the base calling template. In response to determining that the respective color purity is not greater than the color purity of the other pixels or sub-pixels within the distance threshold, making no changes to the base calling template and proceeding to the next pixel and repeating the determining operation until all pixels or at least a subset of the pixels have passed through a similar determining operation.
[0123] The operation of determining the base calling template may include, at least in part, the operation of 340.
[0124] After the base calling template is determined, it can be configured to align flow cell images in a sequencing run, for example, a second plurality of flow cell images of the flow cell device in one or more cycles following the k cycles.
[0125] In some embodiments, the method 600 can further include performing one or more sequencing reaction cycles of the first insert sequence, thereby generating a second plurality of flow cell images. In embodiments having at least two different samples, the method can further include performing one or more sequencing reaction cycles of the first and second insert sequences by the sequencing system, thereby generating a second plurality of flow cell images.
[0126] The second plurality of flow cell images may be generated in one or more cycles following the k cycles corresponding to performing a sequencing reaction on the k-mer sequence. The second plurality of flow cell images may include light signals emitted from nucleotide reagents bound to A, G, C, and T / U nucleotide bases of unbalanced diversity among the plurality of nucleic acid template molecules immobilized on the support in such one or more cycles. The one or more cycles may correspond to a sequencing reaction of a portion of the template molecule that does not include the k-mer sequence. For example, the one or more cycles may correspond to at least a portion of an optional second sample index sequence or an insertion sequence.
[0127] The unbalanced diversity of A, G, C, and T / U nucleotide bases among the plurality of nucleic acid template molecules may include a percentage of the number of one or more types of nucleotide bases (1) to the total number of bases (2) that is less than 20%, 15%, 10%, or 5% in one or more cycles. For example, the unbalanced diversity of nucleotide bases includes a number of nucleotide bases A that is about 5% of the total number of all nucleotide bases in the template molecules of the cycle. As another example, the unbalanced diversity of nucleotide bases includes a number of nucleotide bases C that is about 8% and a number of nucleotide bases T that is about 1% of the total number of all nucleotide bases among the template molecules of the cycle.
[0128] During a single sequencing run, base calling positions, i.e., polonies or clusters, may be shifted, rotated, or otherwise spatially translated within flow cell images obtained from different cycles and / or cross channels. As a result, templates are required to ensure that base calling positions in a sequencing run are spatially aligned and base calls are accurately assigned to the corresponding polonies or clusters, template molecules, and samples. It may be advantageous to generate templates early in a sequencing run to allow alignment of base calling positions all sequencing cycles after generating the template. Early generation of templates may also advantageously reduce delays in primary analysis of flow cell images in subsequent cycles, allowing real-time analysis while the sequencing run is still in progress. For example, if templates are generated from the first 3-5 cycles, the primary analysis may be performed in real-time after the flow cell images are acquired in cycle 6 and in parallel with the sequencing and imaging operations in cycle 7. Similar analysis of each cycle after cycle 6 may be performed while the subsequent sequencing cycles are still in progress. Thus, the primary analysis can be quite soon after, if not when, the sequencing run is completed.
[0129] In some embodiments, the method 600 may further include registering or aligning a second plurality of flow cell images from one or more subsequent flow cycles to the base calling template. In some embodiments, registering the second plurality of flow cell images may include generating coordinates of polonies in the second plurality of flow cell images in a common coordinate system. The base calling template is also in the coordinate system. In some embodiments, registering the second plurality of flow cell images may utilize multiple transformations corresponding to subtiles of the flow cell image to estimate an image transformation of the entire flow cell image. Registering the flow cell images may include generating a transformation of the subtile that provides an estimate of the image transformation of the flow cell image. Information in neighboring subtiles may be used when determining each individual transformation of the subtile.
[0130] In some embodiments, the coordinates of the polony may be stored in a one-dimensional vector or list, where each entry in the vector or list may contain a unique identification of the polony, its coordinates, and other relevant information such as pixel intensities in one or more channels of the cycle.
[0131] In some embodiments, the method 600 can further include performing base calling of the second plurality of flow cell images at base calling positions in the base calling template using signals from the aligned second plurality of flow cell images. Base calling in the primary analysis herein can be based on aligned image intensities in a common coordinate system. In some embodiments, base calling is performed only at base calling positions in the template. In other words, the template serves as a polony map that identifies all the correct polony positions on the flow cell and ignores other signal positions that are unlikely to be polony. As disclosed herein, the base calling positions can be at sub-pixel positions.
[0132] Performing base calling of the second plurality of flow cell images may include, at least in part, operation 350 of FIG. 3 .
[0133] In some embodiments, instead of including operations 610 and 620, method 600 may include receiving a first plurality of library molecules immobilized on a support, each of the first plurality of library molecules including a first insert sequence and a first sample index sequence derived from a first sample source, the first sample index sequence including a first k-mer sequence and a first universal sample index sequence, and the first universal sample index identifies the first sample source of the first insert sequence. In some embodiments, method 600 may include receiving a second plurality of library molecules immobilized on a support, each of the second plurality of library molecules including a second insert sequence and a second sample index sequence derived from a second sample source, the second sample index sequence including a second k-mer sequence and a second universal sample index sequence, and the second universal sample index identifies the second sample source of the second insert sequence. In such embodiments, other operations of method 600 are still similar to those disclosed above.
[0134] In some embodiments, instead of including operation 630, method 600 can include performing k+1 sequencing reaction cycles of the first and second k-mer sequences and base positions downstream of the k-mer sequences (e.g., of the first universal sample index sequence and the second universal sample index sequence), thereby generating a first plurality of flow cell images. In such embodiments, the first cycle immediately following sequencing the k-mer sequence and the k-mer sequence together is configured to generate balanced or highly diverse nucleotide bases in at least two or three of the k+1 cycles. Base calling templates can be generated based on the k+1 cycles together. Additional cycles may, but need not, include random bases. Additional cycles can be from the universal sample index sequence. Alternatively, based on a particular sequencing order, additional cycles can be from the insert sequence or any other portion of the template molecule. In such embodiments, other operations of method 600 are still similar to those disclosed above.
[0135] In some embodiments, instead of including operations 610 and 620, method 600 may include receiving a first plurality of library molecules immobilized on a support, each of the first plurality of library molecules comprising a first insert sequence and a first sample index sequence derived from a first sample source, the first sample index sequence comprising a first k-mer sequence and a first universal sample index sequence, and the first universal sample index identifying the first sample source of the first insert sequence. In such embodiments, method 600 may include receiving a second plurality of library molecules immobilized on a support, each of the second plurality of library molecules comprising a second insert sequence and a second sample index sequence derived from a second sample source, the second sample index sequence comprising a second k-mer sequence and a second universal sample index sequence, and the second universal sample index identifying the second sample source of the second insert sequence. In such an embodiment, instead of including operation 630, method 600 can include performing k+1 sequencing reaction cycles of the first and second k-mer sequences and the base positions of the first universal sample index sequence and the second universal sample index sequence, thereby generating a first plurality of flow cell images. In such an embodiment, the first cycle immediately after sequencing the k-mer sequence and the k-mer sequence together is configured to generate balanced or highly diverse nucleotide bases in at least two or three of the k+1 cycles, and the base calling templates can be generated together based on the k+1 cycles. The additional cycle may, but need not, include random bases. The additional cycle can be from the universal sample index sequence. Alternatively, based on a particular sequencing order, the additional cycle can be from the insert sequence or any other part of the template molecule. In such an embodiment, the other operations of method 600 are still similar to those disclosed above.
[0136] In some embodiments, operation 650 may include determining, by the processor and prior to performing one or more sequencing reaction cycles of the insert sequence, a base calling template including base calling positions based on the pixel intensities and the color purity of each of the pixel intensities, the base calling template being configured to register the second plurality of flow cell images of the support in one or more cycles following the k cycles. In such an embodiment, other operations of method 600 remain similar to those disclosed above.
[0137] Methods for uniplex library molecules In the embodiments disclosed above, method 600 is configured for multiplex library molecules, i.e., library molecules from more than one sample. In some embodiments, method 600 and its operations are configured for singlex library molecules, i.e., library molecules from a single sample. In such embodiments, the method does not include operation 620. Operation 610 can still be similar to that disclosed above. The operations can be adjusted accordingly to remove a second plurality of library molecules, a second insert sequence, a second k-mer sequence, or a combination thereof, corresponding to the second sample source.
[0138] In some embodiments, each of the nucleic acid template molecules can further comprise a second sample index sequence (e.g., 540 or 560 in Figures 5A-5B) that has a second universal sample index sequence. In such embodiments, the combination of the first and second universal sample index sequences uniquely identifies the sample source of the insert sequence (e.g., 510 in Figure 5). In some embodiments, the second sample index lacks random sequences (e.g., 561 or 571 in Figure 5).
[0139] In some embodiments, each of the first plurality of library molecules further comprises a third sample index sequence having a third universal sample index sequence. The combination of the first and third universal sample index sequences can uniquely identify the first sample source of the first insert sequence. In some embodiments, the third sample index lacks a k-mer sequence or a random sequence. In some embodiments, the third sample index comprises a k-mer sequence or a random sequence. The first sample index sequence can be on the left side of the first insert sequence, i.e., the 5' end, or on the right side of the insert sequence, i.e., the 3' end. The third sample index sequence can then be on the opposite end of the insert sequence.
[0140] In some embodiments, each of the library molecules further comprises one or more universal adaptor sequences (eg, 520, 540 in Figures 5A-5B).
[0141] For different library molecules from the same sample, each library molecule can have a different insert sequence. For different library molecules from the same sample, each library molecule can have the same universal index sequence(s). For different library molecules from the same sample, each library molecule can have a different k-mer sequence.
[0142] The first sample source of the insert sequence can be genomic DNA, double-stranded cDNA, or cell-free circulating DNA.
[0143] In some embodiments, a nucleic acid template molecule is clonally amplified from a corresponding library molecule. A nucleic acid template molecule can have one or more clonal copies of a corresponding library molecule.
[0144] In some embodiments, each of the plurality of nucleic acid template molecules is clonally amplified from a corresponding splint complex (e.g., 590 in Figure 5B) that includes a corresponding library molecule and a corresponding splint molecule / adaptor.
[0145] In some embodiments, the plurality of nucleic acid template molecules are immobilized at random positions on the support. In some embodiments, the plurality of nucleic acid template molecules are immobilized at predetermined positions on the support. Each of the plurality of nucleic acid template molecules immobilized on the support can correspond to a polony or a cloned cluster. The positions of the plurality of immobilized template molecules on the support can correspond to base calling positions.
[0146] The density of nucleic acid template molecules on the support is 1 mm 2 About 10 per 2 ~10 12 The density of the nucleic acid template molecules on the support can be 1 mm 2 About 10 per 4 ~10 8 The density of the nucleic acid template molecules on the support can be 1 mm 2 About 10 per 4 ~10 12 The density of the nucleic acid template molecules on the support can be 1 mm 2 About 10 per 4 ~10 5 The density of the nucleic acid template molecules on the support can be 1 mm 2 10 per 2 ~10 12 The density of the nucleic acid template molecules on the support can be 1 mm 2 10 per 4 ~10 8 The density of the nucleic acid template molecules on the support can be 1 mm 2 10 per 4 ~10 12 The density of the nucleic acid template molecules on the support can be 1 mm 2 10 per 4 ~10 5 It could be.
[0147] The nucleic acid template molecule can be sequenced as sequencing reaction cycles are performed.
[0148] The support may include one or more substrates. The support may include a glass substrate or a plastic substrate. The support may include a transparent top substrate that is closest to the objective lens of the optical system. The support may include one or more microfluidic channels, and the template molecules are immobilized on the surface of the microfluidic channels. In some embodiments, the support is included in a flow cell device.
[0149] In some embodiments, the support or a portion thereof (e.g., the surface of a microfluidic channel) is passivated with at least one hydrophilic polymer coating having a water contact angle of 45 degrees or less. The at least one hydrophilic polymer coating may comprise a molecule selected from the group consisting of polyethylene glycol (PEG), poly(vinyl alcohol) (PVA), poly(vinylpyridine), poly(vinylpyrrolidone) (PVP), poly(acrylic acid) (PAA), polyacrylamide, poly(N-isopropylacrylamide) (PNIPAM), poly(methyl methacrylate) (PMA), poly(2-hydroxylethyl methacrylate) (PHEMA), poly(oligo(ethylene glycol) methyl ether methacrylate) (POEGMA), polyglutamic acid (PGA), polylysine, polyglucoside, streptavidin, and dextran. In some embodiments, the at least one hydrophilic polymer coating comprises a branched hydrophilic polymer molecule having at least four branches. In some embodiments, at least one hydrophilic polymer coating comprises polymer molecules having a molecular weight of at least 1000 Daltons.
[0150] In some embodiments, the k-mer sequence may comprise 1, 2, 3, 4, 5, or more nucleotide bases. In some embodiments, the k-mer sequence comprises a random sequence of at least two or three nucleotide bases A, G, C, and T / U. In some embodiments, the k-mer sequence is a random sequence of A, G, C, and T / U nucleotide bases.
[0151] The method 600 can include an operation 630 of performing k sequencing reaction cycles of the first and second k-mer sequences. The operation 630 can be performed by a sequencing system disclosed herein, thereby generating a first plurality of flow cell images.
[0152] In some embodiments, operation 630 is prior to performing one or more sequencing reaction cycles of the first and second insert sequences.
[0153] In some embodiments, operation 630 is based on the sequencing order of the sequencing run. The sequencing order can include sequencing the k-mer sequence, then sequencing the first universal sample index sequence, then sequencing the first insert sequence. The sequencing order can include sequencing the first universal sample index sequence, then sequencing the k-mer sequence, then sequencing the first insert sequence. The sequencing order can include sequencing the k-mer sequence, then sequencing the first insert sequence, then sequencing the first universal sample index sequence. The sequencing order can include sequencing a few bases (e.g., the first 3-8 bases) of the first insert sequence, sequencing the k-mer sequence, then sequencing the first universal sample index sequence. In the sequencing order, if the third sample index sequence lacks a random sequence, sequencing the third sample index sequence may be performed after sequencing the first sample index sequence.
[0154] In some embodiments, the first plurality of flow cell images are from 2, 3, 4, 5, or 6 different color channels in each sequencing cycle. In some embodiments, the first plurality of flow cell images are from k cycles or k+1 sequencing cycles, each cycle comprising a balanced diversity of A, G, C, and T / U nucleotide bases among a plurality of nucleic acid template molecules immobilized on a support.
[0155] In some embodiments, each sequencing reaction cycle is polymerase-mediated. The operation 630 of performing k sequencing reaction cycles of a k-mer sequence can include contacting nucleotide acids or polonies of a plurality of nucleic acid template molecules with a plurality of nucleotide reagents that include a mixture of different types of nucleotide bases A, G, C, and T / U. In some embodiments, each individual nucleotide reagent includes a different detectable color label corresponding to each different type of nucleotide base.
[0156] In some other embodiments, operation 630 of performing k sequencing reaction cycles of a k-mer sequence includes contacting a nucleotide acid or polony of a plurality of nucleic acid template molecules with a plurality of sequencing primers, a plurality of polymerases, and a mixture of different types of avidites, each of which includes a core having a plurality of nucleotide arms attached thereto, each arm of each of the individual avidites including the same type of nucleotide base.
[0157] In some embodiments, the operation 630 of performing k sequencing reaction cycles of the k-mer sequence includes imaging optical color signals emitted from nucleotide reagents bound to the template molecule by an optical system herein. In some embodiments, the operation 630 of performing k sequencing reaction cycles of the k-mer sequence includes acquiring, in each of the k cycles, by an optical system herein, a first plurality of flow cell images including optical color signals emitted from nucleotide reagents bound to the template molecule.
[0158] A first plurality of flow cell images are generated in k cycles corresponding to performing a sequencing reaction on the k-mer sequence. The first plurality of flow cell images can include light signals emitted from nucleotide reagents bound to A, G, C, and T / U nucleotide bases of balanced diversity among the plurality of nucleic acid template molecules immobilized on the support in the k cycles.
[0159] The balanced diversity of A, G, C, and T / U nucleotide bases among a plurality of nucleic acid template molecules can include a percentage of the number of each type of nucleotide base (1) relative to the total number of bases (2) in one or more cycles. The percentage can be greater than 10%, 15%, or 20%. For example, the unbalanced diversity of nucleotide bases can include a number of nucleotide bases A, G, C, T that is 26%, 15%, 27%, and 32%, respectively, of the total number of all nucleotide bases in the template molecules of the cycle.
[0160] The method 600 may include an operation 640 of determining pixel intensities for pixels of the first plurality of flow cell images and respective color purities of each of the pixel intensities. The operation 640 may be performed by a processor as disclosed herein. Each of the pixel intensities may include an intensity of the pixel and / or one or more sub-pixel intensities of a sub-pixel corresponding to the pixel. The respective color purities of each of the pixel intensities may include one or more color purities corresponding to one of the color purities of the pixel and / or one or more sub-pixel intensities of the corresponding pixel. The respective color purities of each of the pixel intensities may include respective color purities of one or more color channels.
[0161] Determining the color purity of each of the pixel intensities can include determining a ratio of (1) a signal corresponding to a particular type of nucleotide base to (2) a total amount of signals for other types of nucleotide bases.
[0162] Determining the pixel intensities can include determining each channel intensity in a set of channel intensities, each channel corresponding to a different fluorescent wavelength, and such determination can be based on a comparison of the set of channel intensities at corresponding pixel or one or more sub-pixel locations.
[0163] The method 600 may include an operation 650 of determining a base calling template including base calling positions based on the pixel intensities and the color purity of each of the pixel intensities determined in operation 640 .
[0164] As disclosed herein, determining the base calling template based on the pixel intensities and the color purity of each of the pixel intensities may include determining for each of the pixels or sub-pixels whether the respective color purity is greater than the color purity of the other pixels or sub-pixels within the distance threshold. In response to determining that the respective color purity is greater than the color purity of the other pixels or sub-pixels within the threshold distance, adding the corresponding pixel or sub-pixel location to the base calling template. In response to determining that the respective color purity is not greater than the color purity of the other pixels or sub-pixels within the distance threshold, making no changes to the base calling template and proceeding to the next pixel and repeating the determining operation until all pixels or at least a subset of the pixels have passed through a similar determining operation.
[0165] After determining the template, a base calling template, or equivalently, the template is configured to align flow cell images in a sequencing run, for example, a second plurality of flow cell images of a flow cell device in one or more cycles following the k cycles.
[0166] During a single sequencing run, base calling positions, i.e., polonies or clusters, may shift, rotate, or otherwise spatially translate from cycle to cycle and / or cross-channel. As a result, templates are needed to ensure that base calling positions in a sequencing run are spatially aligned and base calls are accurately assigned to corresponding polonies, template molecules, and samples. It may be advantageous to generate templates early to allow alignment of base calling positions in subsequent cycles. Early generation of templates may also advantageously reduce delays in primary analysis of flow cell images in subsequent cycles, allowing real-time primary analysis while the sequencing run is still running. For example, if templates are generated using the first 3-5 cycles, the primary analysis may be performed in real-time after the flow cell images in the cycle are acquired in cycle 6, and in cycle 7 in parallel with the sequencing and imaging operations. Similar analysis of each cycle may be performed accordingly. Thus, the primary analysis may be completed shortly after, if not when, the sequencing run is completed.
[0167] In some embodiments, the method 600 can further include performing one or more sequencing reaction cycles of the first insert sequence, thereby generating a second plurality of flow cell images. In embodiments having at least two different samples, the method can further include performing one or more sequencing reaction cycles of the first and second insert sequences by the sequencing system, thereby generating a second plurality of flow cell images.
[0168] The second plurality of flow cell images are generated in one or more cycles following the k cycles corresponding to performing a sequencing reaction on the k-mer sequence. The second plurality of flow cell images may include light signals emitted from nucleotide reagents bound to the A, G, C, and T / U nucleotide bases of unbalanced diversity among the plurality of nucleic acid template molecules immobilized on the support in such one or more cycles. The one or more cycles may correspond to a sequencing reaction of a portion of the template molecule that does not include the k-mer sequence. For example, the one or more cycles may correspond to at least a portion of the optional second sample index sequence or the insertion sequence.
[0169] The unbalanced diversity of A, G, C, and T / U nucleotide bases among the plurality of nucleic acid template molecules can include a percentage of the number of one or more types of nucleotide bases (1) relative to the total number of bases (2). The percentage can be less than 20%, 15%, 10%, or 5% in one or more cycles. For example, the unbalanced diversity of nucleotide bases includes a number of nucleotide bases A that is about 5% of the total number of all nucleotide bases in the template molecules of the cycle. As another example, the unbalanced diversity of nucleotide bases includes a number of nucleotide bases C that is about 8% and a number of nucleotide bases T that is about 1% of the total number of all nucleotide bases among the template molecules of the cycle.
[0170] In some embodiments, the method 600 can further include registering or aligning a second plurality of flow cell images from one or more subsequent flow cycles to the base calling template. In some embodiments, registering the second plurality of flow cell images can include generating coordinates of the polony in the second plurality of flow cell images in a common coordinate system. The base calling template is also in the coordinate system. In some embodiments, the coordinates of the polony can be stored in a one-dimensional vector or list. Each entry in the vector or list can include a unique identification of the polony, its coordinates, and other related information, such as pixel intensities in one or more channels of the cycle.
[0171] In some embodiments, the method 600 can further include performing base calling of the second plurality of flow cell images at base calling positions in the base calling template using signals from the aligned second plurality of flow cell images. Base calling in the primary analysis herein can be based on aligned image intensities in a common coordinate system. In some embodiments, base calling is performed only at base calling positions in the template. In other words, the template serves as a polony map that identifies all the correct polony positions on the flow cell and ignores other signal positions that are unlikely to be polony. As disclosed herein, the base calling positions can be at sub-pixel positions.
[0172] In some embodiments, instead of including operations 610 and 620, method 600 may include an operation of receiving a first plurality of library molecules immobilized on a support, each of the first plurality of library molecules comprising a first insert sequence derived from a first sample source and a first sample index sequence, the first sample index sequence comprising a first k-mer sequence and a first universal sample index sequence, the first universal sample index identifying the first sample source of the first insert sequence. In such embodiments, other operations of method 600 remain similar to those disclosed above with respect to single-plex library molecules.
[0173] In some embodiments, instead of including operation 630, method 600 can include performing k+1 sequencing reaction cycles of the first and second k-mer sequences and of base positions downstream of the k-mer sequences, thereby generating a first plurality of flow cell images. In such embodiments, the first cycle immediately following sequencing the k-mer sequence and the k-mer sequence together is configured to generate balanced or highly diverse nucleotide bases in at least two or three of the k+1 cycles, and base calling templates can be generated together based on the k+1 cycles. One additional cycle may, but need not, include random bases. One additional cycle can be from the universal sample index sequence. Alternatively, based on a particular sequencing order, one additional cycle can be from the insert sequence or any other portion of the template molecule. In such embodiments, other operations of method 600 are still similar to those disclosed above for single-plex library molecules.
[0174] In some embodiments, instead of including operations 610 and 620, the method 600 may include receiving a first plurality of library molecules immobilized on a support, each of the first plurality of library molecules including a first insert sequence and a first sample index sequence derived from a first sample source, the first sample index sequence including a first k-mer sequence and a first universal sample index sequence, the first universal sample index identifying the first sample source of the first insert sequence. In such embodiments, instead of including operation 630, the method 600 may include performing k+1 sequencing reaction cycles of the first and second k-mer sequences and of downstream base positions of the k-mer sequences, thereby generating a first plurality of flow cell images. In such embodiments, the first cycle immediately following sequencing the k-mer sequence and the k-mer sequence together may be configured to generate balanced or highly diverse nucleotide bases in at least two or three of the k+1 cycles, and the base calling templates may be generated together based on the k+1 cycles. The additional cycle may, but need not, include random bases. The additional cycle can be from the universal sample index sequence. Alternatively, based on a particular sequencing order, the additional cycle can be from the insert sequence or any other part of the template molecule. In such an embodiment, the other operations of method 600 are still similar to those disclosed above.
[0175] In some embodiments, operation 650 may include determining, by the processor and prior to performing one or more sequencing reaction cycles of the insert sequence, a base calling template including base calling positions based on the pixel intensities and the color purity of each of the pixel intensities, the base calling template being configured to register the second plurality of flow cell images of the support in one or more cycles following the k cycles. In such embodiments, other operations of method 600 remain similar to those disclosed above for single-plex library molecules.
[0176] Method for sequencing template molecules having unbalanced diversity In general, it is desirable to prepare a nucleic acid library that is distributed on a support (e.g., a coated flow cell), and the library molecules are converted into template molecules immobilized at high density on the support for massively parallel sequencing. For template molecules immobilized at high density, e.g., at random positions on the support, it can be difficult to resolve high density fluorescent images for accurate base calling during a sequencing run. The methods disclosed herein are directed to the detection of relatively high density polonies or clusters, e.g., 1 mm 2 10 per 4 ~10 12 The template molecule or polony allows for base calling of the
[0177] The nucleotide diversity of a population of immobilized template molecules can refer to the relative proportions of nucleotides A, G, C, and T / U present in a sequencing cycle. A high diversity library can generally contain a region of sequence of interest (insertion) in which all four nucleotides are represented in each cycle of a sequencing run in approximately equal proportions. A low diversity library can generally contain a region of sequence of interest (insertion) in which a high proportion of certain nucleotides and a low proportion of other nucleotides are present. To overcome the problem of low diversity libraries, a small amount of a high diversity library prepared from a PhiX bacteriophage is typically mixed with a library of interest (e.g., a PhiX spike-in library) and sequenced together on the same flow cell. Although the PhiX library spike-in library can provide nucleotide diversity, it also occupies space on the flow cell, thereby displacing the target library carrying the sequence of interest and reducing the amount of sequencing data obtained from the target library (e.g., reducing sequencing throughput). Another way to overcome the problem of low diversity libraries may be to prepare target library molecules with at least one sample index sequence that is designed to be color balanced and therefore has high diversity. However, it may be desirable to design single index sample sequences or sets of paired index sample sequences for multiple sample index sets, for example, 16-plex, 24-plex, 96-plex, or greater plexy levels. It may be difficult to design sample index sequences for a large sample index set in which all sample index sequences are color balanced as single or paired sample indexes.
[0178] Another way to overcome the challenge of sequencing low diversity library molecules (e.g., at high density on a support) is to prepare a library with at least one sample index sequence that includes a short k-mer sequence (e.g., NNN) directly linked to a universal sample index sequence, where the k-mer sequence provides nucleotide diversity and color balance. In a population of sample-indexed library molecules, the k-mer sequence of the sample index provides high nucleotide diversity with approximately equal proportions of all four nucleotides (e.g., A, G, C, T, and / or U) represented in each cycle of a sequencing run. The high nucleotide diversity of the k-mer sequence also provides color balance during each cycle of a sequencing run. The advantage of designing the sample index to include a k-mer sequence (e.g., NNN) is that in a low plexi population of library molecules (e.g., 2-plex or 4-plex), the universal sample index sequence does not need to exhibit nucleotide diversity to distinguish between two or four different samples. In addition, the nucleotide diversity of k-mer sequences (e.g., NNN) can eliminate the need to include a PhiX spike-in library or can reduce the amount of PhiX spike-in library that is distributed on the flow cell and sequenced.
[0179] The target library molecule can include a single sample index sequence including a k-mer sequence (e.g., sample index) and a universal sample index sequence. FIG. 5 shows an exemplary linear library molecule (500) including a k-mer sequence and a second optional k-mer sequence (571) (561). The exemplary linear library molecule (500) can also include a universal sample index sequence and an optional second universal sample index sequence (570) (560). In some embodiments, sequencing data from only a single sample index sequence (e.g., 570 or 560 in FIG. 5) is used for polony mapping and alignment of base calling templates, since a k-mer sequence (e.g., NNN) provides sufficient nucleotide diversity and color balance. Sequencing data from a universal sample index sequence can be used to distinguish sequences of interest obtained from different sample sources in a multiplex assay.
[0180] The target library molecule may further include a second sample index sequence (e.g., a dual sample index) that includes a second universal sample index sequence. In some embodiments, the sequencing data from only a single sample index sequence is used for polony mapping and / or alignment of base calling templates, because k-mer sequences provide sufficient nucleotide diversity and color balance, and because polony mapping and / or alignment of base calling templates are preferred in earlier sequencing cycles of a sequencing run to provide base calling templates in subsequent cycles. The sequencing data from the first universal sample index sequence and the second universal sample index sequence may be used as a dual sample index to distinguish sequences of interest obtained from different sample sources in a multiplex assay. In some embodiments, the second sample index sequence (e.g., may or may not include a second k-mer sequence (e.g., NNN).
[0181] Also, the order of sequencing the sequence region and the sample index region(s) can be used to improve the challenge of sequencing low diversity library molecules. For example, the sample index region can be sequenced first before sequencing the sequence region of interest, and the sample index sequence can be associated with the sequence region of interest. For example, the sample index region can be sequenced by first sequencing a k-mer sequence (e.g., NNN), and optionally sequencing at least a portion of the universal sample index), and then sequencing the sequence region of interest. In a population of sample-indexed library molecules, the k-mer sequence (e.g., NNN) provides nucleotide diversity that may not provide the sequence region of interest of the library molecule. The sequence of the sample index provides improved nucleotide diversity and color balance for polony mapping and template alignment.
[0182] In addition, when the sample index region is sequenced first, the length of the sequenced sample index region is relatively short (e.g., less than 30 nucleotides in length) so that the dehybridization of the sequenced sample index region product is more complete. Mild dehybridization conditions can be used to remove most or all of the sequenced sample index region product, which reduces the level of residual signal from any sequencing product that remains hybridized to the template molecule. In contrast, the sequence region of interest is typically much longer than the sample index region (e.g., more than 100 nucleotides in length). If the sequence region of interest is sequenced before the sample index region, the sequenced sequence region of interest product must be subjected to more severe dehybridization conditions that may damage the template molecule in order to remove the product that remains hybridized to the template molecule.
[0183] The present disclosure provides a nucleic acid library molecule (500) comprising at least one sample index sequence, each of which can be used to distinguish sequences of interest obtained from different sample sources in a multiplex assay, wherein at least one sample index sequence comprises a k-mer sequence (e.g., NNN) linked to a universal sample index sequence. In some embodiments, the left sample index comprises a k-mer sequence (e.g., NNN) linked to a left universal sample index sequence, and / or the right sample index comprises a k-mer sequence (e.g., NNN) linked to a right universal sample index sequence. At least one sample index sequence can comprise sequence diversity for improved base calling. At least one sample index sequence can be used to improve base calling accuracy.
[0184] In some embodiments, the k-mer sequence (e.g., NNN) is placed upstream of the universal sample index sequence, so that during a sequencing run, the k-mer sequence portion is sequenced before the universal sample index sequence. The upstream of the universal sample index sequence may be on either the left or right side of the universal sample index sequence, depending on the sequencing or reading order. In other words, the upstream of the universal sample index sequence may be on either the left or right side of the universal sample index sequence, depending on whether the universal sample index sequence is sequenced from the '5' end or from the '3' end. In some embodiments, the k-mer sequence is placed downstream of the universal sample index sequence, so that during a sequencing run, the k-mer portion is sequenced after the universal sample index sequence. The downstream of the universal sample index sequence may be on either the right or left side of the universal sample index sequence, depending on the sequencing or reading order. In other words, the downstream of the universal sample index sequence may be on either the right or left side of the universal sample index sequence, depending on whether the universal sample index sequence is sequenced from the '5' end or from the '3' end. 5, in an exemplary embodiment, when the reading or sequencing of the universal sample index sequence 570 is from the '5' end to the 3' end, the upstream of the universal sample index sequence 570 is to the left of the index sequence. In the same embodiment, when the second reading or sequencing of the second universal sample index sequence 560 is from the '5' end to the 3' end, the upstream of the second universal sample index sequence 560 is to its right hand side.
[0185] In some embodiments, in the k-mer sequence, one or more bases are independently and randomly selected from A, G, C, T, or U. In some embodiments, in the k-mer sequence, one or more bases are independently and randomly selected from A, G, C, T, or U, such that the k-mer sequence lacks consecutive repeat sequences having two or three of the same nucleotide bases, e.g., AA, TT, CC, GG, UU, AAA, TTT, CCC, GGG, or UUU. In some embodiments, in the population of library molecules, the universal sample index sequences include k-mer sequences having high diversity sequences in which all four nucleotides (e.g., A, G, C, T, and / or U) are represented in each cycle of a sequencing run in approximately equal proportions.
[0186] In some embodiments, the k-mer sequence comprises 1 to 20 nucleotides, or 1 to 10 nucleotides, or 2 to 8 nucleotides, or 3 to 6 nucleotides, or 3 to 5 nucleotides, or 3 to 4 nucleotides.
[0187] In some embodiments, k-mer sequences include, but are not limited to, AGC, AGT, GAC, GAT, CAT, CAG, TAG, TAC. One of skill in the art will recognize that many more random sequences can be prepared (e.g., 64 possible combinations) in which each base "N" at a given position in a k-mer sequence is independently selected from A, G, C, T, or U.
[0188] In some embodiments, the universal sample index sequence comprises between 3 and 20 nucleotides, or between 7 and 18 nucleotides, or between 9 and 16 nucleotides.
[0189] In some embodiments, the k-mer sequence comprises one or more random nucleotide bases. In some embodiments, k is an integer greater than 0 and less than 10. In some embodiments, k is 3 and the k-mer sequence is NNN, where each "N" represents a base randomly and independently selected from A, G, C, T, and U. In some embodiments, k is 1, 2, 3, 4, or 5. In some embodiments, k is 3 and the k-mer sequence comprises at least two random nucleotide bases. In some embodiments, k is 4 and the k-mer sequence comprises at least two or three random nucleotide bases.
[0190] In some embodiments, a k-mer sequence (e.g., 571 or 561) is combined with one or more bases from the universal sample index sequence that flanks the k-mer sequence to provide a highly diverse nucleotide base and generate a polony map or base calling template as disclosed herein. A "polony map" as disclosed herein can be equivalent to a "base calling template" that includes base calling positions within a common coordinate system.
[0191] In some embodiments, each right sample index sequence in the population of right sample index sequences includes a universal sample index sequence (e.g., 570 in FIG. 5) and a k-mer sequence (e.g., 571 in FIG. 5). In some embodiments, the k-mer sequence in the population of right sample index sequences has an overall base composition of about 25% or about 15-35% of all four nucleotide bases (e.g., A, G, C, and T / U) to provide nucleotide diversity in each sequencing cycle during sequencing of the k-mer sequence.
[0192] In some embodiments, in the population of right sample index sequences, the percentage of adenine (A) at any given position in the k-mer sequence is about 20-30%, or about 15-35%, or about 10-40%. In some embodiments, in the population of right sample index sequences, the percentage of guanine (G) at any given position in the k-mer sequence is about 20-30%, or about 15-35%, or about 10-40%. In some embodiments, in the population of right sample index sequences, the percentage of cytosine (C) at any given position in the k-mer sequence is about 20-30%, or about 15-35%, or about 10-40%. In some embodiments, in the population of right sample index sequences, the percentage of thymine (T) or uracil (U) at any given position in the k-mer sequence is about 20-30%, or about 15-35%, or about 10-40%.
[0193] In some embodiments, in the population of right sample index sequences, the percentage of adenine (A) and thymine (T), or the percentage of adenine (A) and uracil (U), at any given position in the k-mer sequence is about 10-65%. In some embodiments, in the population of right sample index sequences, the percentage of guanine (G) and cytosine (C) at any given position in the k-mer sequence is about 10-65%.
[0194] In some embodiments, in the population of right sample index sequences, the sequence diversity of the k-mer sequences ensures that, during the sequencing of at least the k-mer sequences, no sequencing cycles that contain less than four different nucleotide bases are presented. In some embodiments, in the population of right sample index sequences, the sequence diversity of the k-mer sequences ensures that, during the sequencing of at least the k-mer sequences, no sequencing cycles that contain less than three different nucleotide bases are presented.
[0195] As disclosed herein, "unbalanced diversity" is used synonymously with "low diversity" and "balanced diversity" is used synonymously with "high diversity."
[0196] Exemplary sample index sequences that include k-mer sequences directly linked to a universal sample index sequence include, but are not limited to, NNNGTAGGAGCC (SEQ ID NO: 97), NNNCCGCTGCTA (SEQ ID NO: 98), NNNAACAACAAG (SEQ ID NO: 99), NNNGGTGGTCTA (SEQ ID NO: 100), NNNTTGGCCAAC (SEQ ID NO: 101), NNNCAGGAGTGC (SEQ ID NO: 105), and NNNATCACACTA (SEQ ID NO: 106). One of skill in the art will recognize that the universal sample index can be of any length and have any sequence that can be used to distinguish sequences of interest obtained from different sample sources in a multiplex assay. In a given sample index, for example, a population of NNNGTAGGAGCC (SEQ ID NO: 97), the population includes a mixture of individual sample index molecules each having the same universal sample index sequence (e.g., GTAGGAGCC) and a different k-mer sequence (e.g., NNN), and up to 64 different k-mer sequences can be present within a given population of sample indexes.
[0197] In some embodiments, the k-mer sequence (e.g., NNN or NNNN) provides a balanced ratio of the nucleobases adenine, cytosine, guanine, thymine, and / or uracil. In some embodiments, in a population of sample-indexed library molecules, the k-mer sequence, along with at least a portion of the universal sample index sequence, provides a balanced ratio of the nucleobases adenine, cytosine, guanine, thymine, and / or uracil represented in each cycle of a sequencing run. In some embodiments, in a population of sample-indexed library molecules, the k-mer sequence, along with one, two, three, or more bases of the universal sample index sequence, provides a balanced ratio of the nucleotide bases adenine, cytosine, guanine, thymine, and / or uracil represented in each cycle of a sequencing run.
[0198] In some embodiments, the sequencing reaction includes the use of a polymerase and nucleotides (e.g., nucleotide analogs) labeled with different fluorophores corresponding to the nucleobases. In some embodiments, sequencing a k-mer sequence (e.g., NNN) using labeled nucleotides provides a balanced ratio of fluorescent colors corresponding to the nucleobases adenine, cytosine, guanine, thymine, and / or uracil in each cycle of a sequencing run. In some embodiments, sequencing at least a portion of a k-mer sequence (e.g., NNN) and a universal sample index sequence using labeled nucleotides provides a balanced ratio of fluorescent colors corresponding to the nucleobases adenine, cytosine, guanine, thymine, and / or uracil. The labeled nucleotides emit a fluorescent signal during the sequencing reaction.
[0199] In some embodiments, the sequencing reaction is performed on a sequencing system (e.g., 110 in FIG. 1) disclosed herein. The sequencing system can have an optical system that captures a fluorescent image from the sequencing reaction on the immobilized template molecule. The sequencing system can be configured to relay the fluorescent imaging data captured by the optical system to a processor (e.g., 400 in FIG. 4) disclosed herein. The processor can be programmed to perform one or more operations disclosed herein to determine the position (e.g., mapping) of the immobilized template molecule on the flow cell. The position (e.g., mapping) of the immobilized template molecule on the flow cell can be equivalent to a base calling position herein. The processor can generate a template of the base calling position of the immobilized template molecule based on the fluorescent imaging data of only the k-mer sequence (e.g., NNN) or based on at least a portion of the k-mer sequence (e.g., NNN) and the universal sample index sequence. Thus, a few sequencing cycles used to sequence the k-mer sequence (e.g., NNN) and optionally a portion of the universal sample index sequence can be used to generate a map of the positions of immobilized template molecules that are equivalent to the templates of base calling positions on the flow cell. The processor can be configured to extract the fluorescent color and intensity of only the k-mer sequence (e.g., NNN), or of at least a portion of the k-mer sequence (e.g., NNN) and the universal sample index sequence. The processor can be configured to use the positions of the given immobilized template molecule and the fluorescent color and intensity associated with the given template molecule (established while sequencing the k-mer sequence) for base calling while sequencing the insertion region. The processor can be configured to detect phasing and pre-phasing while sequencing the k-mer sequence (e.g., NNN) and the universal sample index sequence, and the insertion region.In some embodiments, a balanced ratio of fluorescent colors provided by a k-mer sequence (e.g., NNN) in each sequencing cycle can improve the quality of data processed from the fluorescent images captured by the optical system, which in turn can improve the ability of the processor to determine the location, as well as the color and intensity, of the immobilized template molecules on the flow cell, all of which can improve base calling accuracy and quality scores of the sequenced insert regions.
[0200] In some embodiments, the sequencing reaction involves the use of a polymerase and a multivalent molecule labeled with different fluorophores that correspond to the nucleobases (e.g., adenine, guanine, cytosine, thymine, or uracil) of the nucleotide units attached to the nucleotide arms in a given multivalent molecule. In some embodiments, the core of each multivalent molecule is attached to a fluorophore that corresponds to the nucleotide unit (e.g., adenine, guanine, cytosine, thymine, or uracil) attached to the nucleotide arms in a given multivalent molecule. In some embodiments, at least one of the nucleotide arms of the multivalent molecule comprises a linker and / or a nucleotide base that is attached to a fluorophore, and the fluorophore attached to the given linker or nucleotide base corresponds to the nucleotide base (e.g., adenine, guanine, cytosine, thymine, or uracil) of the nucleotide arm. In some embodiments, sequencing a k-mer sequence (e.g., NNN) using a labeled multivalent molecule provides a balanced ratio of fluorescent colors corresponding to the nucleobases adenine, cytosine, guanine, thymine, and / or uracil in each cycle of a sequencing run. In some embodiments, sequencing a k-mer sequence (e.g., NNN) and at least a portion of a universal sample index sequence using a labeled multivalent molecule provides a balanced ratio of fluorescent colors corresponding to the nucleobases adenine, cytosine, guanine, thymine, and / or uracil. The labeled multivalent molecule emits a fluorescent signal during the sequencing reaction.
[0201] In some embodiments, the sequencing reaction is performed on a sequencing system (e.g., 110 in FIG. 1 ) having an optical system that captures a fluorescent image from the sequencing reaction on the immobilized template molecule. The sequencing device can be configured to relay the fluorescent imaging data captured by the detector to a processor (e.g., 400 in FIG. 4 ) that is programmed to determine (e.g., map) the positions of the immobilized template molecule (polony) that are equivalent to the base calling positions on the flow cell. The processor can generate a map of the positions of the immobilized template molecule based on the fluorescent imaging data of only the k-mer sequence (e.g., NNN) or based on at least a portion of the k-mer sequence (e.g., NNN) and the universal sample index sequence. Thus, a map of the positions of the immobilized template molecule can be generated using a small number of sequencing cycles used to sequence the k-mer sequence (e.g., NNN) and optionally a portion of the universal sample index sequence. The processor can be configured to extract the fluorescent color and intensity of only the k-mer sequence (e.g., NNN) or the k-mer sequence (e.g., NNN) and the universal sample index sequence. The processor can be configured to use the position of a given immobilized template molecule and the fluorescent color and intensity associated with the given template molecule (established while sequencing the k-mer sequence) for base calling while sequencing the insertion region. The processor can be configured to detect phasing and pre-phasing while sequencing the k-mer sequence (e.g., NNN) and universal sample index sequence, as well as the insertion region.In some embodiments, a balanced ratio of fluorescent colors provided by the k-mer sequence (e.g., NNN) in each sequencing cycle can improve the quality of the data processed from the fluorescent images captured by the detector, which in turn can improve the processor's ability to determine the location, as well as the color and intensity, of the immobilized template molecules on the flow cell, all of which can improve the base calling accuracy and quality score of the sequenced insertion region.
[0202] Sequencing Order In some embodiments, the order of performing the multiple cycle sequencing reaction includes (1) sequencing the right sample index, where the right index includes a first k-mer sequence (e.g., 571 in FIG. 5 ) and a right universal sample index sequence (e.g., 570 in FIG. 5 ), (2) sequencing the left sample index, and (3) sequencing the insert region (510). In some embodiments, the left sample index includes a left universal sample index sequence (e.g., 560 in FIG. 5 ). In some embodiments, the left sample index includes a second k-mer sequence (e.g., 561 in FIG. 5 ) and a left universal sample index sequence. In some embodiments, sequencing the right sample index region including the first k-mer sequence (e.g., NNN) and a right universal sample index sequence may provide sufficient nucleotide diversity for the purpose of generating base calling templates, such that sequencing the left sample index may be omitted.
[0203] In some embodiments, a method for performing a sequencing reaction over multiple cycles corresponding to template molecules immobilized to a support, wherein each template molecule (e.g., 500 in FIG. 5 ) comprises: (i) a universal binding sequence for a first surface primer (520); (ii) a left sample index sequence having a left universal sample index sequence (560); (iii) a universal binding sequence for a forward sequencing primer (540); (iv) a sequence of interest (510); (v) a universal binding sequence for a reverse sequencing primer (550); (vi) a right sample index sequence having a k-mer sequence (e.g., NNN) (571) directly linked to the right universal sample index sequence (560); and (vii) a universal binding sequence for a second surface primer (130). The method includes step (a): hybridizing the template molecule with a first plurality of soluble sequencing primers that hybridize to a universal binding sequence for the reverse sequencing primer, sequencing the right sample index sequence including a k-mer sequence (e.g., NNN) and sequencing the right universal sample index sequence, thereby generating a first plurality of sample index extension products hybridized to the immobilized template molecule. The first plurality of sample index extension products are complementary to the right sample index sequence having a k-mer sequence (e.g., NNN) directly linked to the right universal sample index sequence.
[0204] In some embodiments, the method for sequencing further comprises step (b): removing the first plurality of sample index extension products and retaining the immobilized template molecules.
[0205] In some embodiments, the method for sequencing further comprises step (c): hybridizing the retained immobilized template molecules with a second plurality of soluble sequencing primers that hybridize to the universal binding sequence (520) for the first surface primer, and sequencing the left sample index sequence, thereby generating a second plurality of sample index extension products hybridized to the immobilized template molecules. The second plurality of sample index extension products are complementary to the left sample index sequence with the left sample index sequence.
[0206] In some embodiments, the method for sequencing further comprises step (d): removing the second plurality of sample index extension products and retaining the immobilized template molecules.
[0207] In some embodiments, the method for sequencing further comprises step (e): hybridizing the retained immobilized template molecules with a third plurality of soluble sequencing primers that hybridize to the universal binding sequence for the forward sequencing primer (540) and sequence the insertion region (510), thereby generating a plurality of insertion extension products hybridized to the immobilized template molecules, the plurality of insertion extension products being complementary to the sequence of interest (510).
[0208] In some embodiments, the method for sequencing further comprises the step (f1): (i) assigning the sequence of the insertion region (510) to (ii) a right sample index sequence having a k-mer sequence (e.g., NNN) directly linked to the right universal sample index sequence, thereby identifying the insertion region as being obtained from the first source.
[0209] In some embodiments, the method for sequencing further comprises the step (f2): (i) assigning the sequence of the insertion region (510) to (ii) a right sample index sequence having a k-mer sequence (e.g., NNN) directly linked to the right universal sample index sequence and (iii) a left sample index sequence, thereby identifying the insertion region as being obtained from the first source.
[0210] In some embodiments, removal of the multiple sequencing extension products in steps (b) and (d) can be performed using a denaturing reagent comprising an SSC (e.g., saline-sodium citrate) buffer with or without formamide at a temperature that promotes nucleic acid denaturation, such as, for example, 50-90 °C.
[0211] In some embodiments, the sequencing of steps (a), (c), and (e) comprises performing any of the sequencing methods described herein using a sequencing polymerase and detectably labeled nucleotide analogs. In some embodiments, the sequencing of steps (a), (c), and (e) comprises performing any of the two-stage sequencing methods described herein using a sequencing polymerase, detectably labeled multivalent molecules, and nucleotide analogs. In some embodiments, the sequencing of steps (a), (c), and (e) comprises performing any of the binding sequencing methods described herein.
[0212] In some embodiments, the density of the plurality of template molecules immobilized on the support is greater than 1 mm 2 10 per 2 ~10 15 In some embodiments, the density of the plurality of template molecules immobilized on the support is less than 1 mm 2 10 per 4 ~10 12 In some embodiments, the density of the plurality of template molecules immobilized on the support is less than 1 mm 2 10 per4 ~10 8 In some embodiments, the density of the plurality of template molecules immobilized on the support is less than 1 mm 2 10 per 2 ~10 15 In some embodiments, the density of the plurality of template molecules immobilized on the support is less than 1 mm 2 10 per 4 ~10 12 In some embodiments, the density of the plurality of template molecules immobilized on the support is less than 1 mm 2 10 per 4 ~10 8 In some embodiments, the plurality of template molecules are immobilized at random locations on the support. In some embodiments, the plurality of template molecules are immobilized on the support in a predetermined pattern.
[0213] In some embodiments, the sequencing order includes: (1) sequencing the right sample index, where the right index includes a k-mer sequence (e.g., NNN) and a universal sample index sequence; (2) sequencing the insertion region; and (3) sequencing the left sample index. In some embodiments, the left sample index includes a left universal sample index sequence. In some embodiments, the left sample index includes a second k-mer sequence (e.g., NNN) and a left universal sample index sequence. In some embodiments, sequencing the right sample index region, which includes a first random sequence (e.g., NNN) and a right universal sample index sequence, may provide sufficient nucleotide diversity for the purpose of generating base calling templates herein, such that sequencing the left sample index may be omitted.
[0214] In some embodiments, a method for sequencing template molecules immobilized on a support, wherein each template molecule comprises (i) a universal binding sequence for a first surface primer (520), (ii) a left sample index sequence having a left universal sample index sequence, (iii) a universal binding sequence for a forward sequencing primer (540), (iv) a sequence of interest (510), (v) a universal binding sequence for a reverse sequencing primer (550), (vi) a right sample index sequence having a k-mer sequence (e.g., NNN) directly linked to the right universal sample index sequence, and (vii) a universal binding sequence for a second surface primer (530). The method includes step (a): hybridizing the template molecule with a first plurality of soluble sequencing primers that hybridize to a universal binding sequence for the reverse sequencing primer (550), sequencing the right sample index sequence including a k-mer sequence (e.g., NNN) and sequencing the right universal sample index sequence, thereby generating a first plurality of sample index extension products hybridized to the immobilized template molecule. The first plurality of sample index extension products are complementary to the right sample index sequence having a k-mer sequence (e.g., NNN) directly linked to the right universal sample index sequence.
[0215] In some embodiments, the method for sequencing further comprises step (b): removing the first plurality of sample index extension products and retaining the immobilized template molecules.
[0216] In some embodiments, the method for sequencing further comprises step (c): hybridizing the retained immobilized template molecules with a second plurality of soluble sequencing primers that hybridize to the universal binding sequence for the forward sequencing primer (540) and sequence the insertion region (510), thereby generating a plurality of insertion extension products hybridized to the immobilized template molecules, the plurality of insertion extension products being complementary to the sequence of interest (510).
[0217] In some embodiments, the method for sequencing further comprises the step (d): removing the multiple insertion extension products and retaining the immobilized template molecule.
[0218] In some embodiments, the method for sequencing further comprises step (e): hybridizing the retained immobilized template molecules with a third plurality of soluble sequencing primers that hybridize to the universal binding sequence (520) for the first surface primer, and sequencing the left sample index sequence, thereby generating a second plurality of sample index extension products hybridized to the immobilized template molecules. The second plurality of sample index extension products are complementary to the left sample index sequence with the left sample index sequence.
[0219] In some embodiments, the method for sequencing further comprises the step (f1): (i) assigning the sequence of the insertion region (510) to (ii) a right sample index sequence having a k-mer sequence (e.g., NNN) directly linked to the right universal sample index sequence, thereby identifying the insertion region as being obtained from the first source.
[0220] In some embodiments, the method for sequencing further comprises the step (f2): (i) assigning the sequence of the insertion region (510) to (ii) a right sample index sequence having a k-mer sequence (e.g., NNN) directly linked to the right universal sample index sequence and (iii) a left sample index sequence, thereby identifying the insertion region as being obtained from the first source.
[0221] In some embodiments, removal of the multiple sequencing extension products in steps (b) and (d) can be performed using a denaturing reagent comprising an SSC (e.g., saline-sodium citrate) buffer with or without formamide at a temperature that promotes nucleic acid denaturation, such as, for example, 50-90 °C.
[0222] In some embodiments, the sequencing of steps (a), (c), and (e) comprises performing any of the sequencing methods described herein using a sequencing polymerase and detectably labeled nucleotide analogs. In some embodiments, the sequencing of steps (a), (c), and (e) comprises performing any of the two-stage sequencing methods described herein using a sequencing polymerase, detectably labeled multivalent molecules, and nucleotide analogs. In some embodiments, the sequencing of steps (a), (c), and (e) comprises performing any of the binding sequencing methods described herein.
[0223] In some embodiments, the density of the plurality of template molecules immobilized on the support is greater than 1 mm 2 10 per 2 ~10 15 In some embodiments, the density of the plurality of template molecules immobilized on the support is less than 1 mm 2 10 per 4 ~10 12 In some embodiments, the density of the plurality of template molecules immobilized on the support is less than 1 mm 2 10 per4 ~10 8 In some embodiments, the density of the plurality of template molecules immobilized on the support is less than 1 mm 2 10 per 2 ~10 15 In some embodiments, the density of the plurality of template molecules immobilized on the support is less than 1 mm 2 10 per 4 ~10 12 In some embodiments, the density of the plurality of template molecules immobilized on the support is less than 1 mm 2 10 per 4 ~10 8 In some embodiments, the plurality of template molecules are immobilized at random locations on the support. In some embodiments, the plurality of template molecules are immobilized on the support in a predetermined pattern.
[0224] In some embodiments, the order of performing the sequencing reaction in multiple cycles corresponding to the template molecules immobilized on the support includes: (1) sequencing the insert region (510); (2) sequencing the right sample index, where the right index comprises a random sequence (e.g., NNN) and a universal sample index sequence; and (3) sequencing the left sample index. In some embodiments, the left sample index comprises a left universal sample index sequence. In some embodiments, the left sample index comprises a second random sequence (e.g., NNN) and a left universal sample index sequence. In some embodiments, sequencing the right sample index region, which comprises a first random sequence (e.g., NNN) and a right universal sample index sequence, may provide sufficient nucleotide diversity for the purpose of generating base calling templates, such that sequencing the left sample index may be omitted.
[0225] In some embodiments, a method for sequencing template molecules immobilized on a support, each template molecule comprising (i) a universal binding sequence for a first surface primer (520), (ii) a left sample index sequence having a left universal sample index sequence, (iii) a universal binding sequence for a forward sequencing primer (540), (iv) a sequence of interest (510), (v) a universal binding sequence for a reverse sequencing primer (550), (vi) a right sample index sequence having a k-mer sequence (e.g., NNN) directly linked to the right universal sample index sequence, and (vii) a universal binding sequence for a second surface primer (530). The method comprises the step (a): hybridizing the template molecule with a first plurality of soluble sequencing primers that hybridize to the universal binding sequence for the forward sequencing primer (540) and sequencing the insertion region (510), thereby generating a plurality of insertion extension products hybridized to the immobilized template molecule. The multiple insertion extension products are complementary to the sequence of interest (510).
[0226] In some embodiments, the method for sequencing further comprises the step (b): removing the multiple insertion extension products and retaining the immobilized template molecule.
[0227] In some embodiments, the method for sequencing further comprises step (c): hybridizing the template molecule with a second plurality of soluble sequencing primers that hybridize to the universal binding sequence for the reverse sequencing primer (550), sequencing the right sample index sequence including a k-mer sequence (e.g., NNN) and sequencing the right universal sample index sequence, thereby generating a first plurality of sample index extension products hybridized to the immobilized template molecule. The first plurality of sample index extension products are complementary to the right sample index sequence having a k-mer sequence (e.g., NNN) directly linked to the right universal sample index sequence.
[0228] In some embodiments, the method for sequencing further comprises the step (d): removing the first plurality of sample index extension products and retaining the immobilized template molecules.
[0229] In some embodiments, the method for sequencing further comprises step (e): hybridizing the retained immobilized template molecules with a third plurality of soluble sequencing primers that hybridize to the universal binding sequence (520) for the first surface primer, and sequencing the left sample index sequence, thereby generating a second plurality of sample index extension products hybridized to the immobilized template molecules. The second plurality of sample index extension products are complementary to the left sample index sequence with the left sample index sequence.
[0230] In some embodiments, the method for sequencing further comprises the step (f1): (i) assigning the sequence of the insertion region (510) to (ii) a right sample index sequence having a k-mer sequence (e.g., NNN) directly linked to the right universal sample index sequence, thereby identifying the insertion region as being obtained from the first source.
[0231] In some embodiments, the method for sequencing further comprises the step (f2): (i) assigning the sequence of the insertion region (510) to (ii) a right sample index sequence having a k-mer sequence (e.g., NNN) directly linked to the right universal sample index sequence and (iii) a left sample index sequence, thereby identifying the insertion region as being obtained from the first source.
[0232] In some embodiments, removal of the multiple sequencing extension products in steps (b) and (d) can be performed using a denaturing reagent comprising an SSC (e.g., saline-sodium citrate) buffer with or without formamide at a temperature that promotes nucleic acid denaturation, such as, for example, 50-90 °C.
[0233] In some embodiments, the sequencing of steps (a), (c), and (e) comprises performing any of the sequencing methods described herein using a sequencing polymerase and detectably labeled nucleotide analogs. In some embodiments, the sequencing of steps (a), (c), and (e) comprises performing any of the two-stage sequencing methods described herein using a sequencing polymerase, detectably labeled multivalent molecules, and nucleotide analogs. In some embodiments, the sequencing of steps (a), (c), and (e) comprises performing any of the binding sequencing methods described herein.
[0234] In some embodiments, the density of the plurality of template molecules immobilized on the support is greater than 1 mm 2 10 per 2 ~10 15 In some embodiments, the density of the plurality of template molecules immobilized on the support is less than 1 mm 2 10 per 4 ~10 12 In some embodiments, the density of the plurality of template molecules immobilized on the support is less than 1 mm2 10 per 4 ~10 8 In some embodiments, the density of the plurality of template molecules immobilized on the support is less than 1 mm 2 10 per 2 ~10 15 In some embodiments, the density of the plurality of template molecules immobilized on the support is less than 1 mm 2 10 per 4 ~10 12 In some embodiments, the density of the plurality of template molecules immobilized on the support is less than 1 mm 2 10 per 4 ~10 8 In some embodiments, the plurality of template molecules are immobilized at random locations on the support. In some embodiments, the plurality of template molecules are immobilized on the support in a predetermined pattern.
[0235] In some embodiments, the order of performing the sequencing reactions in multiple cycles corresponding to the template molecule includes: (1) sequencing the first 3-5 bases of the insertion region (510); (2) sequencing the right sample index, where the right index includes an optional random sequence (e.g., NNN) and a universal sample index sequence; and (3) sequencing the left sample index. In some embodiments, the left sample index includes a left universal sample index sequence. In some embodiments, the left sample index includes a second random sequence (e.g., NNN) and a left universal sample index sequence. In some embodiments, sequencing the first 3-5 bases of the insertion region (510) may provide sufficient sequence diversity such that the right sample index and the left sample index do not include a k-mer sequence (e.g., NNN). In some embodiments, sequencing the right sample index region including a first random sequence (e.g., NNN) and a right universal sample index sequence may provide sufficient nucleotide diversity such that sequencing the left sample index may be omitted.
[0236] In some embodiments, a method for sequencing template molecules immobilized on a support, each template molecule comprising (i) a universal binding sequence for a first surface primer (520), (ii) a left sample index sequence having a left universal sample index sequence, (iii) a universal binding sequence for a forward sequencing primer (540), (iv) a sequence of interest (510), (v) a universal binding sequence for a reverse sequencing primer (550), (vi) a right sample index sequence having a k-mer sequence (e.g., NNN) directly linked to the right universal sample index sequence, and (vii) a universal binding sequence for a second surface primer (530). The method comprises the step (a): hybridizing the template molecule with a first plurality of soluble sequencing primers that hybridize to the universal binding sequence for the forward sequencing primer (540) and sequencing the first 3-5 bases of the insertion region (510), thereby generating a plurality of insertion extension products hybridized to the immobilized template molecule. The multiple insertion extension products are complementary to the sequence of interest (510). The sequence of the first 3-5 bases of the insertion region (510) may provide sufficient sequence diversity and color balance for polony mapping and template alignment.
[0237] In some embodiments, the method for sequencing further comprises the step (b): removing the multiple insertion extension products and retaining the immobilized template molecule.
[0238] In some embodiments, the method for sequencing further comprises step (c): hybridizing the template molecule with a second plurality of soluble sequencing primers that hybridize to the universal binding sequence for the reverse sequencing primer (550) and sequencing the right sample index sequence including the k-mer sequence (e.g., NNN) (if present) and sequencing the right universal sample index sequence, thereby generating a first plurality of sample index extension products hybridized to the immobilized template molecule. The first plurality of sample index extension products are complementary to the right sample index sequence.
[0239] In some embodiments, the method for sequencing further comprises the step (d): removing the first plurality of sample index extension products and retaining the immobilized template molecules.
[0240] In some embodiments, the method for sequencing further comprises step (e): hybridizing the retained immobilized template molecules with a third plurality of soluble sequencing primers that hybridize to the universal binding sequence (520) for the first surface primer, and sequencing the left sample index sequence, thereby generating a second plurality of sample index extension products hybridized to the immobilized template molecules. The second plurality of sample index extension products are complementary to the left sample index sequence with the left sample index sequence.
[0241] In some embodiments, the method for sequencing further comprises the step (f): removing the second plurality of sample index extension products and retaining the immobilized template molecules.
[0242] In some embodiments, the method for sequencing further comprises step (g): hybridizing the template molecule with a fourth plurality of soluble sequencing primers that hybridize to the universal binding sequence for the forward sequencing primer (540) and sequence the entire length of the insertion region (510), thereby generating a plurality of full-length insertion extension products hybridized to the immobilized template molecule. The plurality of full-length insertion extension products are complementary to the sequence of interest (510).
[0243] In some embodiments, the method for sequencing further comprises the step (h1): (i) assigning the full length sequence of the insertion region (510) to (ii) the right sample index sequence, thereby identifying the insertion region as being obtained from the first source.
[0244] In some embodiments, the method for sequencing further comprises the step (h2): (i) assigning the full length sequence of the insertion region (510) to (ii) a right sample index sequence, and (iii) a left sample index sequence, thereby identifying the insertion region as being obtained from the first source.
[0245] In some embodiments, removal of the multiple sequencing extension products in steps (b), (d), and (f) can be performed using a denaturing reagent comprising an SSC (e.g., saline-sodium citrate) buffer with or without formamide at a temperature that promotes nucleic acid denaturation, such as, for example, 50-90 °C.
[0246] In some embodiments, the sequencing of steps (a), (c), (e), and (g) comprises performing any of the sequencing methods described herein using a sequencing polymerase and detectably labeled nucleotide analogs. In some embodiments, the sequencing of steps (a), (c), (e), and (g) comprises performing any of the two-stage sequencing methods described herein using a sequencing polymerase, detectably labeled multivalent molecules, and nucleotide analogs. In some embodiments, the sequencing of steps (a), (c), (e), and (g) comprises performing any of the binding sequencing methods described herein.
[0247] In some embodiments, the density of the plurality of template molecules immobilized on the support is greater than 1 mm 2 10 per 2 ~10 15 In some embodiments, the density of the plurality of template molecules immobilized on the support is less than 1 mm 2 10 per 4 ~10 12 In some embodiments, the density of the plurality of template molecules immobilized on the support is less than 1 mm 2 10 per 4 ~10 8 In some embodiments, the density of the plurality of template molecules immobilized on the support is less than 1 mm 2 10 per 2 ~10 15 In some embodiments, the density of the plurality of template molecules immobilized on the support is less than 1 mm 2 10 per 4 ~10 12 In some embodiments, the density of the plurality of template molecules immobilized on the support is less than 1 mm 2 10 per 4 ~10 8In some embodiments, the plurality of template molecules are immobilized at random locations on the support. In some embodiments, the plurality of template molecules are immobilized on the support in a predetermined pattern.
[0248] In some embodiments, the sequencing order includes: (1) sequencing the first 3-5 bases of the insertion region (510) of the immobilized template molecule (e.g., sequencing in the forward direction); (2) sequencing the right sample index, where the right index comprises a first random sequence (e.g., NNN) and a right universal sample index sequence; (3) sequencing the left sample index; (4) performing a pairwise turn reaction such that the immobilized template molecule is replaced with an immobilized strand that is complementary to the template molecule; and (5) sequencing the entire length of the insertion region (510) of the immobilized complementary strand (e.g., sequencing in the reverse direction). In some embodiments, the sequence of the first 3-5 bases of the insertion region (510) of the population of library molecules may provide sufficient sequence diversity for improved base calling accuracy. In some embodiments, the left sample index comprises a left universal sample index sequence. In some embodiments, the left sample index comprises a second random sequence (e.g., NNN) and a left universal sample index sequence. In some embodiments, sequencing the right sample index region comprising the first random sequence (e.g., NNN) and the right universal sample index sequence may provide sufficient nucleotide diversity such that sequencing the left sample index may be omitted.
[0249] In some embodiments, a method of sequencing template molecules immobilized on a support, wherein each template molecule is covalently attached to an immobilized capture primer lacking uracil bases, and each template molecule comprises randomly distributed uracil bases, and each template molecule comprises (i) a universal binding sequence for a first surface primer (520), (ii) a left sample index sequence having a left universal sample index sequence, (iii) a universal binding sequence for a forward sequencing primer (540), (iv) a sequence of interest (510), (v) a universal binding sequence for a reverse sequencing primer (550), (vi) a right sample index sequence having a k-mer sequence (e.g., NNN) directly linked to the right universal sample index sequence, and (vii) a universal binding sequence for a second surface primer (530). The method includes step (a): hybridizing a template molecule with a first plurality of soluble sequencing primers (e.g., forward sequencing primers) that hybridize to a universal binding sequence for the forward sequencing primer (540) to sequence the first 3-5 bases of the insertion region (510), thereby generating a plurality of insertion extension products hybridized to the immobilized template molecule. The plurality of insertion extension products are complementary to the sequence of interest (510).
[0250] In some embodiments, the method for sequencing further comprises the step (b): removing the multiple insertion extension products and retaining the immobilized template molecule.
[0251] In some embodiments, the method for sequencing further comprises step (c): hybridizing the template molecule with a second plurality of soluble sequencing primers that hybridize to the universal binding sequence for the reverse sequencing primer (550), sequencing the right sample index sequence including a k-mer sequence (e.g., NNN) and sequencing the right universal sample index sequence, thereby generating a first plurality of sample index extension products hybridized to the immobilized template molecule. The first plurality of sample index extension products are complementary to the right sample index sequence having a k-mer sequence (e.g., NNN) directly linked to the right universal sample index sequence.
[0252] In some embodiments, the method for sequencing further comprises the step (d): removing the first plurality of sample index extension products and retaining the immobilized template molecules.
[0253] In some embodiments, the method for sequencing further comprises step (e): hybridizing the retained immobilized template molecules with a third plurality of soluble sequencing primers that hybridize to the universal binding sequence (520) for the first surface primer, and sequencing the left sample index sequence, thereby generating a second plurality of sample index extension products hybridized to the immobilized template molecules. The second plurality of sample index extension products are complementary to the left sample index sequence with the left sample index sequence.
[0254] In some embodiments, the method for sequencing further comprises the step (f): displacing a second plurality of sample index extension products hybridized to the immobilized template molecule by performing a primer extension reaction using a strand displacing polymerase and a plurality of nucleotides to generate extension products hybridized to the immobilized template molecule that comprises an immobilized capture primer.
[0255] In some embodiments, the method for sequencing further comprises step (g): generating abasic sites at uracil sites of the immobilized template molecule by generating gaps at the abasic sites, removing the immobilized template molecule, thereby generating gap-containing template molecules, while retaining the extension products generated in step (f), where each extension product is retained by hybridizing to an immobilized capture primer. In some embodiments, a pairwise turn is achieved by performing steps (g) and (h).
[0256] In some embodiments, the method for sequencing further comprises the step (h): hybridizing the retained extension products with a fourth plurality of soluble sequencing primers (e.g., reverse sequencing primers) that hybridize to a universal binding sequence for the reverse sequencing primer (550) and sequencing the insertion region (510) (e.g., sequencing at least a portion or the entire length of the insertion region (510)).
[0257] In some embodiments, the method for sequencing further comprises the step (i1): (i) assigning the sequence of the insertion region (510) to (ii) a right sample index sequence having a k-mer sequence (e.g., NNN) directly linked to the right universal sample index sequence, thereby identifying the insertion region as being obtained from the first source.
[0258] In some embodiments, the method for sequencing further comprises the step (i2): (i) assigning the sequence of the insertion region (510) to (ii) a right sample index sequence having a k-mer sequence (e.g., NNN) directly linked to the right universal sample index sequence and (iii) a left sample index sequence, thereby identifying the insertion region as being obtained from the first source.
[0259] In some embodiments, removal of the multiple sequencing extension products in steps (b) and (d) can be performed using a denaturing reagent comprising an SSC (e.g., saline-sodium citrate) buffer with or without formamide at a temperature that promotes nucleic acid denaturation, such as, for example, 50-90 °C.
[0260] In some embodiments, the sequencing of steps (a), (c), (e), and (h) comprises performing any of the sequencing methods described herein using a sequencing polymerase and detectably labeled nucleotide analogs. In some embodiments, the sequencing of steps (a), (c), (e), and (h) comprises performing any of the two-stage sequencing methods described herein using a sequencing polymerase, detectably labeled multivalent molecules, and nucleotide analogs. In some embodiments, the sequencing of steps (a), (c), (e), and (h) comprises performing any of the binding sequencing methods described herein.
[0261] In some embodiments, the density of the plurality of template molecules immobilized on the support is greater than 1 mm 2 In some embodiments, the plurality of template molecules are immobilized at random positions on the support. In some embodiments, the plurality of template molecules are immobilized on the support in a predetermined pattern.
[0262] In some embodiments, the disclosure provides a method for sequencing a nucleic acid, the method comprising: (a) providing a plurality of nucleic acid template molecules immobilized (e.g., randomly or at predetermined locations) on a support, each immobilized template molecule comprising an insert sequence region and one sample index, each sample index comprising a k-mer sequence linked to a universal sample index sequence that identifies a sample source of the insert sequence, different immobilized template molecules having different k-mer sequences and the same universal sample index sequence, the immobilized template molecules having different insert sequences; and (b) performing three polymerase-mediated sequencing reaction cycles of the k-mer sequences of the plurality of immobilized template molecules using a plurality of detectably labeled nucleotide reagents comprising a mixture of different types of nucleobases A, G, C, and T / U. (c) performing three polymerase-mediated sequencing reaction cycles in which a balanced diversity of A, G, C, and T / U nucleobases are detected and imaged in each of the first, second, and third sequencing cycles among the plurality of immobilized template molecules, wherein the nucleotide reagents comprise different detectable color labels corresponding to each different type of nucleobase, and the three sequencing cycles include detecting and imaging light color signals emitted from the detectably labeled nucleotide reagents bound to the immobilized template molecules, thereby determining a sequence of the k-mer sequence within each template molecule of the plurality of immobilized template molecules; and (d) performing three polymerase-mediated sequencing reaction cycles in which a balanced diversity of A, G, C, and T / U nucleobases are detected and imaged in each of the first, second, and third sequencing cycles among the plurality of immobilized template molecules, wherein the sequence of the intercalated region is not used to generate the map.
[0263] In some embodiments, in the method for sequencing a nucleic acid, the balanced diversity of step (b) is about 5-85%, or about 5-60%, or about 10-50%, or about 15-55%, or about 25-75% of each of the nucleobases A, G, C, and T / U detected and imaged in each of the first, second, and third sequencing cycles.
[0264] In some embodiments, in the method for sequencing a nucleic acid, the method further comprises: (a) sequencing a universal sample index sequence of a plurality of immobilized template molecules; (b) sequencing an insert sequence region of a plurality of immobilized template molecules; and (c) assigning an insert sequence of a given template molecule obtained in step (b) to a universal sample index sequence from the same given template molecule obtained in step (a), thereby identifying a sample source of the given insert sequence.
[0265] In some embodiments, in the method for sequencing a nucleic acid, the plurality of nucleic acid template molecules further comprises a second sample index comprising a second universal sample index sequence that identifies a sample source of the insert sequence, wherein the second sample index lacks random sequences.
[0266] In some embodiments, in the method for sequencing nucleic acids, the method further comprises: (a) sequencing the k-mer sequences of the plurality of immobilized template molecules to obtain a balanced diversity of A, G, C, and T / U nucleobases that are detected and imaged in each of the first, second, and third sequencing cycles to generate a map of the positions of the plurality of immobilized template molecules; (b) sequencing a first universal sample index sequence of the plurality of immobilized template molecules; (c) sequencing a second universal sample index sequence of the plurality of immobilized template molecules; (d) sequencing an insert sequence region of the plurality of immobilized template molecules; and (e) assigning an insert sequence of a given template molecule to the first and second universal sample index sequences from the same given template molecule obtained in steps (a) and (b), thereby identifying a sample source of the given insert sequence.
[0267] In some embodiments, the disclosure provides a method for sequencing a nucleic acid, the method comprising: (a) providing a plurality of nucleic acid template molecules immobilized (e.g., randomly or at a predetermined location) on a support, each immobilized template molecule comprising an insert sequence region and one sample index, each sample index comprising a k-mer sequence linked to a universal sample index sequence that identifies a sample source of the insert sequence, the universal sample index sequence comprising between 3 and 20 nucleotides, different immobilized template molecules having different k-mer sequences and the same universal sample index sequence, the immobilized template molecules having different insert sequences; and (b) performing four polymerase-mediated sequencing reactions of the k-mer sequences of the plurality of immobilized template molecules and a first base position of the universal sample index sequence using a plurality of detectably labeled nucleotide reagents comprising a mixture of different types of nucleobases A, G, C, and T / U. (c) performing four polymerase-mediated sequencing reaction cycles, wherein the nucleotide reagents comprise different detectable color labels corresponding to each different type of nucleobase, the four sequencing cycles comprising detecting and imaging light color signals emitted from the detectably labeled nucleotide reagents bound to the immobilized template molecules, thereby determining a sequence of a first base position of the k-mer sequence and the universal sample index sequence within each template molecule of the plurality of immobilized template molecules, wherein a balanced diversity of A, G, C, and T / U nucleobases are detected and imaged in each of the first, second, third, and fourth sequencing cycles among the plurality of immobilized template molecules; and (c) generating a map of the positions of the plurality of immobilized template molecules using the images of the four polymerase-mediated sequencing reaction cycles obtained in step (b), wherein the sequence of the insertion region is not used to generate the map.
[0268] In some embodiments, in the method for sequencing a nucleic acid, the balanced diversity of step (b) is about 5-85%, or about 5-60%, or about 10-50%, or about 15-55%, or about 25-75% of each of the nucleobases A, G, C, and T / U detected and imaged in each of the first, second, third, and fourth sequencing cycles.
[0269] In some embodiments, in the method for sequencing a nucleic acid, the method further comprises: (a) sequencing the remaining base positions of the universal sample index sequence of the plurality of immobilized template molecules; (b) sequencing the insert sequence region of the plurality of immobilized template molecules; and (c) assigning the insert sequence of a given template molecule obtained in step (b) to the universal sample index sequence from the same given template molecule obtained in step (a), thereby identifying the sample source of the given insert sequence.
[0270] In some embodiments, in the method for sequencing a nucleic acid, the plurality of nucleic acid template molecules further comprises a second sample index comprising a second universal sample index sequence that identifies a sample source of the insert sequence, wherein the second sample index lacks random sequences.
[0271] In some embodiments, in the method for sequencing nucleic acids, the method further comprises: (a) sequencing a first base position of the k-mer sequence and the universal sample index sequence of the plurality of immobilized template molecules to obtain a balanced diversity of A, G, C, and T / U nucleobases that are detected and imaged in each of the first, second, third, and fourth sequencing cycles to generate a map of the positions of the plurality of immobilized template molecules; (b) sequencing the remaining base positions of the first universal sample index sequence of the plurality of immobilized template molecules; (c) sequencing the second universal sample index sequence of the plurality of immobilized template molecules; (d) sequencing the insert sequence region of the plurality of immobilized template molecules; and (e) assigning the insert sequence of a given template molecule to the first and second universal sample index sequences from the same given template molecule obtained in steps (a) and (b), thereby identifying the sample source of the given insert sequence.
[0272] In some embodiments, in any of the methods for sequencing nucleic acids, the support comprises a glass substrate or a plastic substrate. In some embodiments, the support is configured on a flow cell channel, a flow cell, or a capillary lumen. In some embodiments, the support is passivated with at least one hydrophilic polymer coating having a water contact angle of 45 degrees or less. In some embodiments, the at least one hydrophilic polymer coating comprises a molecule selected from the group consisting of polyethylene glycol (PEG), poly(vinyl alcohol) (PVA), poly(vinylpyridine), poly(vinylpyrrolidone) (PVP), poly(acrylic acid) (PAA), polyacrylamide, poly(N-isopropylacrylamide) (PNIPAM), poly(methyl methacrylate) (PMA), poly(2-hydroxylethyl methacrylate) (PHEMA), poly(oligo(ethylene glycol) methyl ether methacrylate) (POEGMA), polyglutamic acid (PGA), polylysine, polyglucoside, streptavidin, and dextran. In some embodiments, at least one hydrophilic polymer coating comprises a branched hydrophilic polymer molecule having at least 4 branches. In some embodiments, at least one hydrophilic polymer coating comprises a polymer molecule having a molecular weight of at least 1000 Daltons.
[0273] In some embodiments, in any of the methods for sequencing nucleic acids, the immobilized template molecules comprise a plurality of immobilized concatemer molecules having a tandem repeat sequence of an insert sequence and one sample index. In some embodiments, the immobilized template molecules comprise a plurality of different clustered template molecules having one copy of the insert sequence and one copy of one sample index, and the clustered template molecules are generated by bridge amplification. In some embodiments, the density of the immobilized nucleic acid template molecules arranged at random or predetermined positions on the support is less than 1 mm 2 10 per 4 ~10 8In some embodiments, the sample source of the insert sequence is genomic DNA, double-stranded cDNA, or cell-free circulating DNA.
[0274] In some embodiments, in any of the methods for sequencing a nucleic acid, the detectably labeled nucleotide reagent comprises a nucleotide each comprising an aromatic nucleobase, a 5-carbon sugar moiety, 1-10 phosphate groups, and a fluorophore. In some embodiments, the detectably labeled nucleotide reagent comprises a nucleotide each comprising an aromatic nucleobase, a 5-carbon sugar moiety having a chain terminating group at the 3' carbon sugar position, 1-10 phosphate groups, and a fluorophore. In some embodiments, the detectably labeled nucleotide reagent comprises a multivalent molecule each comprising: (1) a core, (2) a plurality of nucleotide arms, and (3) at least one fluorophore, each of the nucleotide arms comprising: (i) a core attachment moiety, (ii) a spacer comprising a PEG moiety, (iii) a linker, and (iv) a nucleotide unit, wherein the core is attached to the plurality of nucleotide arms, the spacer is attached to the linker, and the linker is attached to the nucleotide unit.
[0275] In some embodiments, in any of the methods for sequencing a nucleic acid, the detectably labeled nucleotide reagents that are bound to the immobilized template molecules in step (b) include individual immobilized template molecules that hybridize to a sequencing primer to form a duplex, the duplex is bound to a polymerase to form a hybridized polymerase, and the hybridized polymerase is bound to the detectably labeled nucleotide reagent. In some embodiments, the hybridized polymerase is bound to the detectably labeled nucleotide reagent under conditions suitable for binding the detectably labeled nucleotide reagent to the hybridized polymerase and incorporating the detectably labeled nucleotide into the hybridized sequencing primer, the detectably labeled nucleotide reagent comprising an aromatic nucleobase, a five-carbon sugar moiety, 1-10 phosphate groups, and a fluorophore. In some embodiments, the multiplexed polymerase is coupled to a detectably labeled nucleotide reagent under conditions suitable for binding the detectably labeled nucleotide reagent to the multiplexed polymerase and incorporating the detectably labeled nucleotide into the hybridized sequencing primer, the detectably labeled nucleotide reagent comprising an aromatic nucleobase, a 5-carbon sugar moiety having a chain terminating group at the 3' carbon sugar position, from 1 to 10 phosphate groups, and a fluorophore. In some embodiments, the hybridized polymerase is bound to the detectably labeled nucleotide reagent under conditions suitable for binding the detectably labeled nucleotide reagent to the hybridized polymerase, which conditions are suitable for inhibiting nucleotide incorporation, and the detectably labeled nucleotide reagent comprises a multivalent molecule comprising (1) a core, (2) a plurality of nucleotide arms, and (3) at least one fluorophore, each nucleotide arm comprising (i) a core attachment moiety, (ii) a spacer comprising a PEG moiety, (iii) a linker, and (iv) a nucleotide unit, wherein the core is attached to the plurality of nucleotide arms, the spacer is attached to the linker, and the linker is attached to the nucleotide unit.
[0276] In some embodiments, in any of the methods for sequencing a nucleic acid, the immobilized template molecule comprises immobilized concatemeric molecules hybridized to a plurality of sequencing primers to form at least a first and a second duplex on the same concatemeric molecule, the first and second duplexes being bound to a first polymerase and the second duplex being bound to a second polymerase to form a first and a second hybridized polymerase, and the method further comprises (a) hybridizing the plurality of multivalent molecules to the first and the second hybridized polymerase on the same concatemeric template molecule. wherein each multivalent molecule comprises (1) a core, (2) a plurality of nucleotide arms, and (3) at least one fluorophore, each nucleotide arm comprising (i) a core attachment moiety, (ii) a spacer comprising a PEG moiety, (iii) a linker, and (iv) a nucleotide unit, wherein the core is attached to the plurality of nucleotide arms, the spacer is attached to the linker, and the linker is attached to the nucleotide unit, and the contacting is performed under conditions suitable for binding a single multivalent molecule from the plurality to a first and a second composite polymerase. wherein a first nucleotide unit of the single multivalent molecule binds to a first multiplexed polymerase comprising a first sequencing primer hybridized to a first portion of the concatemeric template molecule, thereby forming a first binding complex, and a second nucleotide unit of the single multivalent molecule binds to a second multiplexed polymerase comprising a second sequencing primer hybridized to a second portion of the concatemeric template molecule, thereby forming a second binding complex, and the first and second binding complexes that bind to the same multivalent molecule form an avidity complex. (b) detecting the first and second binding complexes on the same concatemeric template molecule; (c) imaging light color signals emitted from the detectably labeled multivalent molecules that form the first and second binding complexes on the same concatemeric template molecule; and (d) identifying the first nucleotide unit in the first binding complex, thereby identifying the first nucleotide unit in the first binding complex.determining the sequence of the first portion of the concatemeric template molecule and identifying a second nucleotide unit in the second binding complex, thereby identifying the sequence of the second portion of the concatemeric template molecule;
[0277] In some embodiments, the disclosure provides a method for multiplex sequencing of nucleic acids, the method comprising: (a) providing a first plurality of library molecules among the plurality comprising (i) an insert sequence region derived from a first sample source, (ii) a first sample index having a k-mer sequence linked to a first universal sample index sequence, and (iii) a second sample index having a second universal sample index sequence lacking random sequence, wherein a combination of the first and second universal sample index sequences uniquely identifies the first sample source of the insert sequence, and different first library molecules have different k-mer sequences and have different insert sequences; and (b) providing a second plurality of library molecules among the plurality comprising (i) an insert sequence region derived from a second sample source, (ii) a third sample index having a k-mer sequence linked to a third universal sample index sequence, and (iii) a fourth sample index having a fourth universal sample index sequence lacking random sequence, wherein each molecule of the third and fourth universal sample index sequences uniquely identifies the first sample source of the insert sequence, and different first library molecules have different k-mer sequences and have different insert sequences. (c) pooling the first and second plurality of library molecules; (d) distributing the pooled library molecules on a support and conducting an amplification reaction to generate a plurality of clonally amplified template molecules immobilized on the support (e.g., immobilized at random or predetermined locations); and (e) performing three polymerase-mediated sequencing reaction cycles of the k-mer sequences of the first and third sample indexes using a plurality of detectably labeled nucleotide reagents comprising a mixture of different types of nucleobases A, G, C, and T / U, wherein the nucleotide reagents comprise different detectable color labels corresponding to each of the different types of nucleobases, and the three sequencing cycles include detecting and imaging light color signals emitted from the detectably labeled nucleotide reagents bound to the immobilized amplified template molecules, therebydetermining the sequence of the k-mer sequence in each template molecule of the plurality of immobilized template molecules, performing three polymerase-mediated sequencing reaction cycles in which a balanced diversity of A, G, C, and T / U nucleobases are detected and imaged in each of the first, second, and third sequencing cycles among the plurality of immobilized amplified template molecules; and (f) using the image obtained in step (e) to generate a map of the positions of the plurality of immobilized template molecules, wherein the sequence of the inserted region is not used to generate the map.
[0278] In some embodiments, in the method for multiplex sequencing of nucleic acids, the balanced diversity of step (e) is about 5-85%, or about 5-60%, or about 10-50%, or about 15-55%, or about 25-75% of each of the nucleobases A, G, C, and T / U detected and imaged in each of the first, second, and third sequencing cycles.
[0279] In some embodiments, in the method for multiplex sequencing of nucleic acids, the method further comprises: (a) sequencing a first universal sample index sequence of the plurality of immobilized template molecules; (b) sequencing a second universal sample index sequence of the plurality of immobilized template molecules; (c) sequencing an insert sequence region of the plurality of immobilized template molecules derived from a first library molecule; and (d) assigning an insert sequence of a given template molecule obtained in step (c) to the first and second universal sample index sequences from the same given template molecule, thereby identifying a first sample source of the given insert sequence.
[0280] In some embodiments, in the method for multiplex sequencing of nucleic acids, the method further comprises: (a) sequencing a third universal sample index sequence of the plurality of immobilized template molecules; (b) sequencing a fourth universal sample index sequence of the plurality of immobilized template molecules; (c) sequencing an insert sequence region of the plurality of immobilized template molecules derived from a second library molecule; and (d) assigning the insert sequence of a given template molecule obtained in step (c) to the third and fourth universal sample index sequences from the same given template molecule, thereby identifying a second sample source of the given insert sequence.
[0281] In some embodiments, the disclosure provides a method for multiplex sequencing of nucleic acids, the method comprising: (a) providing a first plurality of library molecules among the plurality, the first plurality of library molecules including (i) an insert sequence region from a first sample source, (ii) a first sample index having a k-mer sequence linked to a first universal sample index sequence, and (iii) a second sample index having a second universal sample index sequence lacking random sequence, wherein a combination of the first and second universal sample index sequences uniquely identifies the first sample source of the insert sequence, the first universal sample index sequence comprising between 3 and 20 nucleotides, and different first library molecules having different k-mer sequences and having different insert sequences; and (b) providing a second plurality of library molecules among the plurality, each molecule including (i) an insert sequence region from a second sample source, (ii) a third sample index having a k-mer sequence linked to a third universal sample index sequence, and (iii) a fourth sample index having a fourth universal sample index sequence lacking random sequence. (c) pooling the first and second plurality of library molecules; (d) distributing the pooled library molecules on a support and conducting an amplification reaction to generate a plurality of clonally amplified template molecules immobilized on the support (e.g., immobilized at random or predetermined locations); and (e) performing four polymerase-mediated sequencing reaction cycles of the k-mer sequences of the first and third sample indexes to sequence the first base positions of the first and third universal sample index sequences using a plurality of detectably labeled nucleotide reagents comprising a mixture of different types of nucleobases A, G, C, and T / U, wherein the nucleotide reagents comprise different detectable color labels corresponding to each of the different types of nucleobases, and the three sequencing cycles aredetecting and imaging light color signals emitted from detectably labeled nucleotide reagents bound to the immobilized amplified template molecules, thereby determining the sequence of the k-mer sequence in each template molecule of the plurality of immobilized template molecules, performing four polymerase-mediated sequencing reaction cycles of the k-mer sequences of the first and third sample indexes, in which a balanced diversity of A, G, C, and T / U nucleobases are detected and imaged in each of the first, second, third, and fourth sequencing cycles among the plurality of immobilized amplified template molecules, to sequence the first base positions of the first and third universal sample index sequences; and (f) generating a map of the positions of the plurality of immobilized template molecules using the image obtained in step (e), wherein the sequence of the insertion region is not used to generate the map.
[0282] In some embodiments, in the method for multiplex sequencing nucleic acids, the balanced diversity of step (e) is about 5-85%, or about 5-60%, or about 10-50%, or about 15-55%, or about 25-75% of each of the nucleobases A, G, C, and T / U detected and imaged in each of the first, second, third, and fourth sequencing cycles.
[0283] In some embodiments, in the method for multiplex sequencing of nucleic acids, the method further comprises: (a) sequencing the remaining base positions of the first universal sample index sequence of the plurality of immobilized template molecules; (b) sequencing the second universal sample index sequence of the plurality of immobilized template molecules; (c) sequencing an insert sequence region of the plurality of immobilized template molecules derived from a first library molecule; and (d) assigning the insert sequence of a given template molecule obtained in step (c) to the first and second universal sample index sequences from the same given template molecule, thereby identifying a first sample source of the given insert sequence.
[0284] In some embodiments, in the method for multiplex sequencing of nucleic acids, the method further comprises: (a) sequencing the remaining base positions of the third universal sample index sequence of the plurality of immobilized template molecules; (b) sequencing the fourth universal sample index sequence of the plurality of immobilized template molecules; (c) sequencing an insert sequence region of the plurality of immobilized template molecules derived from a second library molecule; and (d) assigning the insert sequence of a given template molecule obtained in step (c) to the third and fourth universal sample index sequences from the same given template molecule, thereby identifying a second sample source of the given insert sequence.
[0285] In some embodiments, in any of the methods for multiplex sequencing of nucleic acids, the support comprises a glass substrate or a plastic substrate. In some embodiments, the support is configured on a flow cell channel, a flow cell, or a capillary lumen. In some embodiments, the support is passivated with at least one hydrophilic polymer coating having a water contact angle of 45 degrees or less. In some embodiments, the at least one hydrophilic polymer coating comprises a molecule selected from the group consisting of polyethylene glycol (PEG), poly(vinyl alcohol) (PVA), poly(vinylpyridine), poly(vinylpyrrolidone) (PVP), poly(acrylic acid) (PAA), polyacrylamide, poly(N-isopropylacrylamide) (PNIPAM), poly(methyl methacrylate) (PMA), poly(2-hydroxylethyl methacrylate) (PHEMA), poly(oligo(ethylene glycol) methyl ether methacrylate) (POEGMA), polyglutamic acid (PGA), polylysine, polyglucoside, streptavidin, and dextran. In some embodiments, at least one hydrophilic polymer coating comprises a branched hydrophilic polymer molecule having at least 4 branches. In some embodiments, at least one hydrophilic polymer coating comprises a polymer molecule having a molecular weight of at least 1000 Daltons.
[0286] In some embodiments, in any of the methods for multiplex sequencing of nucleic acids, the immobilized template molecules comprise a plurality of immobilized concatemer molecules having a tandem repeat sequence of an insert sequence and one sample index. In some embodiments, the immobilized template molecules comprise a plurality of different clustered template molecules having one copy of the insert sequence and one copy of one sample index, and the clustered template molecules are generated by bridge amplification. In some embodiments, the density of the immobilized nucleic acid template molecules (e.g., random or fixed at a predetermined location) on the support is greater than 1 mm 2 10 per 4 ~108 In some embodiments, the sample source of the insert sequence is genomic DNA, double-stranded cDNA, or cell-free circulating DNA.
[0287] In some embodiments, in any of the methods for multiplex sequencing of nucleic acids, the detectably labeled nucleotide reagent comprises a nucleotide each comprising an aromatic nucleobase, a 5-carbon sugar moiety, 1-10 phosphate groups, and a fluorophore. In some embodiments, the detectably labeled nucleotide reagent comprises a nucleotide each comprising an aromatic nucleobase, a 5-carbon sugar moiety having a chain terminating group at the 3' carbon sugar position, 1-10 phosphate groups, and a fluorophore. In some embodiments, the detectably labeled nucleotide reagent comprises a multivalent molecule each comprising (1) a core, (2) a plurality of nucleotide arms, and (3) at least one fluorophore, each of the nucleotide arms comprising (i) a core attachment moiety, (ii) a spacer comprising a PEG moiety, (iii) a linker, and (iv) a nucleotide unit, wherein the core is attached to a plurality of nucleotide arms, the spacer is attached to the linker, and the linker is attached to the nucleotide unit.
[0288] In some embodiments, in any of the methods for multiplex sequencing of nucleic acids, the detectably labeled nucleotide reagents that are bound to the immobilized template molecules in step (e) include individual immobilized template molecules that hybridize to a sequencing primer to form a duplex, the duplex is bound to a polymerase to form a hybridized polymerase, and the hybridized polymerase is bound to the detectably labeled nucleotide reagent. In some embodiments, the hybridized polymerase is bound to the detectably labeled nucleotide reagent under conditions suitable for binding the detectably labeled nucleotide reagent to the hybridized polymerase and incorporating the detectably labeled nucleotide into the hybridized sequencing primer, the detectably labeled nucleotide reagent comprising an aromatic nucleobase, a five-carbon sugar moiety, 1-10 phosphate groups, and a fluorophore. In some embodiments, the multiplexed polymerase is coupled to a detectably labeled nucleotide reagent under conditions suitable for binding the detectably labeled nucleotide reagent to the multiplexed polymerase and incorporating the detectably labeled nucleotide into the hybridized sequencing primer, the detectably labeled nucleotide reagent comprising an aromatic nucleobase, a 5-carbon sugar moiety having a chain terminating group at the 3' carbon sugar position, from 1 to 10 phosphate groups, and a fluorophore. In some embodiments, the hybridized polymerase is bound to the detectably labeled nucleotide reagent under conditions suitable for binding the detectably labeled nucleotide reagent to the hybridized polymerase, which conditions are suitable for inhibiting nucleotide incorporation, and the detectably labeled nucleotide reagent comprises a multivalent molecule comprising (1) a core, (2) a plurality of nucleotide arms, and (3) at least one fluorophore, each nucleotide arm comprising (i) a core attachment moiety, (ii) a spacer comprising a PEG moiety, (iii) a linker, and (iv) a nucleotide unit, wherein the core is attached to the plurality of nucleotide arms, the spacer is attached to the linker, and the linker is attached to the nucleotide unit.
[0289] In some embodiments, in any of the methods for multiplex sequencing of nucleic acids, the immobilized template molecule comprises immobilized concatemeric molecules hybridized to a plurality of sequencing primers to form at least first and second duplexes on the same concatemeric molecule, the first and second duplexes bind to a first polymerase and the second duplexes bind to a second polymerase to form first and second hybridized polymerases, and the method further comprises (a) hybridizing the plurality of multivalent molecules to the first and second hybridized polymerases on the same concatemeric template molecule. and contacting a first and second hybrid polymerase with a first polymerase, wherein each multivalent molecule comprises (1) a core, (2) a plurality of nucleotide arms, and (3) at least one fluorophore, each nucleotide arm comprising (i) a core attachment moiety, (ii) a spacer comprising a PEG moiety, (iii) a linker, and (iv) a nucleotide unit, wherein the core is attached to the plurality of nucleotide arms, the spacer is attached to the linker, and the linker is attached to the nucleotide unit, and the contacting is suitable for binding a single multivalent molecule from the plurality to a first and second hybrid polymerase. The method is carried out under conditions in which a first nucleotide unit of the single multivalent molecule binds to a first multiplexed polymerase comprising a first sequencing primer hybridized to a first portion of the concatemeric template molecule, thereby forming a first binding complex, a second nucleotide unit of the single multivalent molecule binds to a second multiplexed polymerase comprising a second sequencing primer hybridized to a second portion of the concatemeric template molecule, thereby forming a second binding complex, and the first and second binding complexes that bind to the same multivalent molecule are avidin- forming and contacting a first and second binding complex under conditions suitable for inhibiting polymerase-catalyzed incorporation of the bound first and second nucleotide units in the first and second binding complexes; (b) detecting the first and second binding complexes on the same concatemeric template molecule; (c) imaging light color signals emitted from the detectably labeled multivalent molecules that form the first and second binding complexes on the same concatemeric template molecule; and (d) identifying the first nucleotide unit in the first binding complex;and identifying a second nucleotide unit in the second binding complex, thereby determining the sequence of the first portion of the concatemeric template molecule and thereby identifying the sequence of the second portion of the concatemeric template molecule.
[0290] Exemplary Computer System Various aspects may be implemented using one or more computer systems, such as, for example, computer system 400 shown in Figure 4. One or more computer systems 400 may be used, for example, to implement any of the aspects discussed herein, as well as combinations and subcombinations thereof.
[0291] Computer system 400 may include one or more processors (also referred to as central processing units, or CPUs), such as processor 404. Processor 404 may be connected to a bus or communication infrastructure 406.
[0292] The computer system 400 may also include user input / output device(s) 403, such as a monitor, keyboard, pointing device, etc., that may communicate with a communications infrastructure 406 via user input / output interface(s) 402. The user input / output device 403 may be coupled to the user interface 124 of FIG.
[0293] One or more of the processors 404 may be a graphics processing unit (GPU). In one aspect, a GPU may be a processor that is a specialized electronic circuit designed to process mathematically intensive applications. GPUs may have a parallel structure that is efficient for parallel processing of large blocks of data, such as, for example, mathematically intensive data common in computer graphics applications, images, videos, vector processing, array processing, and the like, as well as encryption (including brute force cracking), generating cryptographic hashes or hash arrays, solving partial hash reversal problems, and / or generating results for other proof-of-work calculations for some blockchain-based applications. Due to the power of general-purpose computing on a graphics processing unit (GPGPU), GPUs may be particularly useful in at least the image recognition and machine learning aspects described herein.
[0294] In addition, one or more of the processors 404 may include a coprocessor or other implementation of logic for accelerating cryptographic calculations or other specialized mathematical functions, including hardware accelerated cryptographic coprocessors. Such accelerated processors may further include an instruction set(s) for acceleration using the coprocessor and / or other logic to facilitate such acceleration.
[0295] Computer system 400 may also include a main or primary memory 408, such as a random access memory (RAM). Main memory 408 may include one or more levels of cache. Main memory 408 may store control logic (i.e., computer software) and / or data therein.
[0296] Computer system 400 may also include one or more secondary storage or secondary memories 410. Secondary memory 410 may include, for example, a main storage drive 412 and / or a removable storage device or drive 414. Main storage drive 412 may be, for example, a hard disk drive or a solid state drive. Removable storage drive 414 may be a floppy disk drive, a magnetic tape drive, a compact disk drive, an optical storage device, a tape backup device, and / or any other storage device / drive.
[0297] The removable storage drive 414 may interface with a removable storage unit 418 .
[0298] The removable storage unit 418 may include a computer usable or readable storage device that stores computer software (control logic) and / or data. The removable storage unit 418 may be a floppy disk, magnetic tape, compact disk, DVD, optical storage disk, and / or any other computer data storage device. The removable storage drive 414 may read from and / or write to the removable storage unit 418.
[0299] Secondary memory 410 may include other means, devices, components, instruments, or other approaches for allowing computer programs and / or other instructions and / or data to be accessed by computer system 400. Such means, devices, components, equipment, or other approaches may include, for example, removable storage unit 422 and interface 420. Examples of removable storage unit 422 and interface 420 may include program cartridges and cartridge interfaces (such as those found in video game devices), removable memory chips (such as EPROMs or PROMs) and associated sockets, memory sticks and USB ports, memory cards and associated memory card slots, and / or any other removable storage unit and associated interface.
[0300] The computer system 400 may further include a communication or network interface 424. The communication interface 424 may enable the computer system 400 to communicate and interact with any combination of external devices, external networks, external entities, etc. (individually and collectively referenced by reference number 428). For example, the communication interface 424 may enable the computer system 400 to communicate with external or remote devices 428 via communication paths 426, which may be wired and / or wireless (or a combination thereof) and may include any combination of a LAN, a WAN, the Internet, etc. Control logic and / or data may be transmitted to and from the computer system 400 via the communication paths 426. In some aspects, the communication paths 426 are connections to the cloud 130, as shown in FIG. 1. The external devices, etc. referenced by reference number 428 may be devices, networks, entities, etc. within the cloud 130.
[0301] Computer system 400 may also be any of a personal digital assistant (PDA), a desktop workstation, a laptop or notebook computer, a netbook, a tablet, a smartphone, a smart watch or other wearable, an appliance, a part of the Internet of Things (IoT), and / or an embedded system, to name a few non-limiting examples, or any combination thereof.
[0302] It should be understood that the framework described herein may be implemented as a method, process, apparatus, system, or product, such as a non-transitory computer-readable medium or device. For illustrative purposes, the framework may be described in the context of a distributed ledger being publicly available, or at least available to untrusted third parties. A current use case includes a blockchain-based system. However, it should be understood that the framework may also be applied to other settings where sensitive or confidential information may need to pass through the hands of untrusted third parties, and the technology is in no way limited to the use of distributed ledgers or blockchains.
[0303] Computer system 400 may be a client or server accessing or hosting any application and / or data through any delivery paradigm, including but not limited to remote or distributed cloud computing solutions, local or on-premise software (e.g., an “on-premise” cloud-based solution), an “as a service” model (e.g., Content as a Service (CaaS), Digital Content as a Service (DCaaS), Software as a Service (SaaS), Managed Software as a Service (MSaaS), Platform as a Service (PaaS), Desktop as a Service (DaaS), Framework as a Service (FaaS), Backend as a Service (BaaS), Mobile Backend as a Service (MBaaS), Infrastructure as a Service (IaaS), Database as a Service (DBaaS), etc.), and / or a hybrid model including any combination of the above examples or other service or delivery paradigms.
[0304] Any applicable data structures, file formats, and schemas may be derived from standards, including, but not limited to, JavaScript Object Notation (JSON), Extensible Markup Language (XML), Yet Another Markup Language (YAML), Extensible Hypertext Markup Language (XHTML), Wireless Markup Language (WML), MessagePack, XML User Interface Language (XUL), or other functionally similar representations, alone or in combination. Alternatively, proprietary data structures, formats, or schemas may be used exclusively or in combination with known or open standards.
[0305] Any associated data, files, and / or databases may be stored, retrieved, accessed, and / or transmitted in a human readable format, such as a numerical, textual, graphical, or multimedia format, further including various types of markup languages, among other possible formats. Alternatively, or in combination with the above formats, the data, files, and / or databases may be stored, retrieved, accessed, and / or transmitted in binary, encoded, compressed, and / or encrypted format, or any other machine-readable format.
[0306] The interfaces or interconnections between the various systems and layers may use any number of mechanisms, such as any number of protocols, programmatic frameworks, floorplans, or application programming interfaces (APIs), including, but not limited to, Document Object Model (DOM), Discovery Service (DS), NSUserDefaults, Web Services Description Language (WSDL), Message Exchange Patterns (MEP), Web Distributed Data Exchange (WDDX), Web Hypertext Applications Technology Working Group (WHATWG), HTML5 Web Messaging, Representational State Transfer (REST or RESTful Web Services), Extensible User Interface Protocol (XUP), Simple Object Access Protocol (SOAP), XML Schema Definition (XSD), XML Remote Procedure Call (XML-RPC), or any other mechanism that may achieve similar functionality and results.
[0307] Such interfaces or interconnections may also utilize Uniform Resource Identifiers (URIs), which may further include Uniform Resource Locators (URLs) or Uniform Resource Names (URNs). Other forms of uniform and / or unique identifiers, locators, or names may be used, either exclusively or in combination with forms such as those described above.
[0308] Any of the above protocols or APIs may interface with or be implemented in, and compiled or interpreted by, any programming language, procedural, functional, or object-oriented language, including, but not limited to, C, C++, C#, Objective-C, Java, Scala, Clojure, Elixir, Swift, Go, Perl, PHP, Python, Ruby, JavaScript, WebAssembly, or virtually any other language, with any other libraries or schema, with any kind of framework, runtime environment, virtual machine, interpreter, stack, engine, or similar mechanism (including, but not limited to, Node.js, V8, Knockout, jQuery, Dojo, Dijit, OpenUI5, AngularJS, Expressjs, Backbone.js, Ember.js, DHTMLX, Vue, React, Electron, etc., among many other non-limiting examples).
[0309] In some aspects, a tangible, non-transitory apparatus or article of manufacture that includes a tangible, non-transitory computer usable or readable medium having control logic (software) stored thereon may be referred to herein as a computer program product or program storage device. This includes, but is not limited to, computer system 400, main memory 408, secondary memory 410, and removable storage units 418 and 422, as well as tangible articles of manufacture embodying any combination of the foregoing. Such control logic, when executed by one or more data processing devices (such as computer system 400), may cause such data processing devices to operate as described herein.
[0310] Based on the teachings contained herein, it will be apparent to one of ordinary skill in the relevant art(s) how to make and use aspects of the present disclosure using data processing devices, computer systems, and / or computer architectures other than those shown in Figure 4. In particular, aspects may operate with software, hardware, and / or operating system implementations other than those described herein.
[0311] The imager 116 of FIG. 1 can include one or more optical systems. Further disclosed herein are optical system design guidelines and high performance fluorescence imaging methods and systems that provide improved optical resolution and image quality for fluorescence imaging-based genomics applications. The disclosed optical imaging system designs provide larger fields of view, increased spatial resolution, improved modulation transfer, contrast-to-noise ratio, and image quality, higher spatial sampling frequencies, faster transitions between image captures when repositioning the sample plane to capture a series of images (e.g., of different fields of view), and improved imaging system duty cycles, thus enabling higher throughput image acquisition and analysis.
[0312] In some cases, for example, improved imaging performance for dual-sided (flow cell) imaging applications can be achieved by using an electro-optic phase plate in combination with the objective lens to compensate for optical aberrations caused by a fluid layer separating the upper (near) and lower (far) inner surfaces of the flow cell. In some cases, this design approach can also compensate for vibrations introduced by a motion-activated compensator that is moved in or out of the optical path, for example, depending on which surface of the flow cell is the image.
[0313] In some cases, for example, improved imaging performance for dual-sided (flow cell) imaging applications involving the use of thick flow cell walls (e.g., wall (or coverslip) thickness greater than 700 µm) and fluidic channels (e.g., fluidic channel height or thickness between 50-200 µm) can be achieved even when using commercially available off-the-shelf objectives by using tube lens designs that, in combination with the objective, correct the optical aberrations caused by the thick flow cell walls and / or intermediate fluid layers.
[0314] In some cases, for example, improved imaging performance for multi-channel (e.g., two-color or four-color) imaging applications can be achieved by using multiple tube lenses (one for each imaging channel), with each tube lens design optimized for the particular wavelength range used in that imaging channel.
[0315] Exemplary aspects disclosed herein may include a fluorescence imaging system comprising: a) at least one light source configured to provide excitation light within one or more specified wavelength ranges; and b) an objective configured to collect fluorescence arising from within a specified field of view of the sample surface upon exposure of the sample surface to the excitation light, wherein the objective has a numerical aperture of at least 0.1, at least 0.2, at least 0.3, at least 0.4, at least 0.5, at least 0.6, at least 0.7, at least 0.8, at least 0.9, or a numerical aperture within a range defined by any two of the foregoing; a working distance of the objective is at least 400 μm, at least 500 μm, at least 600 μm, at least 700 μm, at least 800 μm, at least 900 μm, at least 1000 μm, or a working distance within a range defined by any two of the foregoing; and a field of view of at least 0.1 mm. 2 , at least 0.2 mm 2 , at least 0.5 mm 2 , at least 0.7 mm 2 , at least 1 mm 2, at least 2 mm 2 , at least 3 mm 2 , at least 5 mm 2 , or at least 10 mm 2 or any two of the foregoing; and c) at least one image sensor, where fluorescence collected by the objective lens is imaged onto the image sensor, and pixel dimensions of the image sensor are selected such that a spatial sampling frequency of the fluorescence imaging system is at least twice the optical resolution of the fluorescence imaging system.
[0316] In some embodiments, the numerical aperture may be at least 0.75. In some embodiments, the numerical aperture is at least 1.0. In some embodiments, the working distance is at least 850 μm. In some embodiments, the working distance is at least 1,000 μm. In some embodiments, the field of view is at least 2.5 mm. 2 In some embodiments, the field of view may have an area of at least 3 mm 2In some embodiments, the spatial sampling frequency may be at least 2.5 times the optical resolution of the fluorescence imaging system. In some embodiments, the spatial sampling frequency may be at least 3 times the optical resolution of the fluorescence imaging system. In some embodiments, the system may further comprise an XYZ translation stage such that the system is configured to acquire a series of two or more fluorescence images in an automated manner, each image of the series of images being acquired or may be acquired for a different field of view. In some embodiments, the position of the sample plane may be simultaneously adjusted in the X, Y, and Z directions to coincide with the position of the objective focal plane during acquisition of the images of the different fields of view. In some embodiments, the time required for simultaneous adjustment in the X, Y, and Z directions may be less than 0.3 seconds, less than 0.4 seconds, less than 0.5 seconds, less than 0.7 seconds, or less than 1 second, or a time within a range defined by any two of the foregoing. In some embodiments, the system further comprises an autofocus mechanism configured to adjust the position of the focal plane before acquiring the images of the different fields of view if the error signal indicates that the difference in the positions of the focal plane and the sample plane in the Z direction is greater than a specified error threshold. In some embodiments, the specified error threshold is 100 nm or more. In some embodiments, the specified error threshold is 50 nm or less. In some embodiments, the system comprises three or more image sensors, and the system is configured to image the fluorescence in each of the three or more wavelength ranges onto a different image sensor. In some embodiments, the difference in the position of the focal plane between each of the three or more image sensors and the sample plane is less than 100 nm. In some embodiments, the difference in the position of the focal plane between each of the three or more image sensors and the sample plane is less than 50 nm. In some embodiments, the total time required to reposition the sample plane, adjust focus if necessary, and acquire an image is less than 0.4 seconds per field of view. In some embodiments, the total time required to reposition the sample plane, adjust focus if necessary, and acquire an image is less than 0.3 seconds per field of view.
[0317] Also disclosed herein is a fluorescence imaging system for dual-sided imaging of a flow cell, the fluorescence imaging system including: a) an objective lens configured to collect fluorescence originating from within a particular field of view of a sample surface within the flow cell; and b) at least one tube lens disposed between the objective lens and at least one image sensor, the at least one tube lens configured to compensate an imaging performance metric for the combination of the objective lens, the at least one tube lens, and the at least one image sensor when imaging an inner surface of the flow cell, the flow cell having a wall thickness of at least 700 μm and a gap between the upper and lower inner surfaces of at least 50 μm, the imaging performance metric being substantially the same for imaging the upper or lower inner surface of the flow cell without moving an optical compensator in or out of the optical path between the flow cell and the at least one image sensor, without moving one or more optical elements of the tube lens along the optical path, and without moving one or more optical elements of the tube lens in or out of the optical path.
[0318] In some embodiments, the objective lens may be a commercially available microscope objective lens. In some embodiments, the commercially available microscope objective lens may have a numerical aperture of at least 0.3. In some embodiments, the objective lens may have a working distance of at least 700 μm. In some embodiments, the objective lens may be corrected to compensate for a coverslip thickness (or flow cell wall thickness) of 0.17 mm or greater or less than 0.17 mm. In some embodiments, the optical system may be corrected to compensate for a coverslip thickness, a flow cell thickness, or a distance between the desired focal planes. In some embodiments, the correction may be made by inserting a corrective optical system, such as a lens or optical assembly, into the optical path of the optical system. In some embodiments, the correction may be made without inserting a corrective optical system, such as a lens or optical assembly, into the optical path of the optical system. In some embodiments, the fluorescence imaging system may further comprise an electro-optical phase plate disposed adjacent to the objective lens and between the objective lens and the tube lens, the electro-optical phase plate may provide correction for optical aberrations caused by a fluid filling a gap between the upper and lower inner surfaces of the flow cell. In some embodiments, the at least one tube lens can be a compound lens that includes three or more optical components. In some embodiments, the at least one tube lens is a compound lens that includes four optical components, which can include one or more of a first asymmetric convex-convex lens, a second convex-flat lens, a third asymmetric concave-concave lens, and a fourth asymmetric convex-concave lens, and can be in the above order or any alternative order. In some embodiments, the at least one tube lens is configured to correct the imaging performance metric for the combination of the objective lens, the at least one tube lens, and the at least one imaging element when imaging the inner surface of a flow cell having a wall thickness of at least 1 mm.In some embodiments, the at least one tube lens is configured to correct an imaging performance metric for the combination of the objective lens, the at least one tube lens, and the at least one imaging element when imaging an inner surface of a flow cell having a gap of at least 100 μm. In some embodiments, the at least one tube lens is configured to correct an imaging performance metric for the combination of the objective lens, the at least one tube lens, and the at least one imaging element when imaging an inner surface of a flow cell having a gap of at least 200 μm. In some embodiments, the system includes a single objective lens, two tube lenses, and two image sensors, each of the two tube lenses designed to provide optimal imaging performance at a different fluorescent wavelength. In some embodiments, the system includes a single objective lens, three tube lenses, and three image sensors, each of the three tube lenses designed to provide optimal imaging performance at a different fluorescent wavelength. In some embodiments, the system includes a single objective lens, four tube lenses, and four image sensors, each of the four tube lenses designed to provide optimal imaging performance at a different fluorescent wavelength. In some embodiments, the design of the objective lens or at least one tube lens is configured to optimize the modulation transfer function in the mid-high spatial frequency range. In some embodiments, the imaging performance metric includes a measurement of the modulation transfer function (MTF) at one or more specified spatial frequencies, defocus, spherical aberration, chromatic aberration, coma, astigmatism, field curvature, image distortion, contrast-to-noise ratio (CNR), or any combination thereof. In some embodiments, the difference in the imaging performance metric for imaging the upper and lower inner surfaces of the flow cell is less than 10%. In some embodiments, the difference in the imaging performance metric for imaging the upper and lower inner surfaces of the flow cell is less than 5%.In some embodiments, the use of at least one tube lens improves imaging performance metrics for dual-sided imaging by at least the same or better than that for a conventional system with an objective lens, a motion-activated compensator, and an image sensor. In some embodiments, the use of at least one tube lens improves imaging performance metrics for dual-sided imaging by at least 10% compared to that for a conventional system with an objective lens, a motion-activated compensator, and an image sensor.
[0319] Disclosed herein is an illumination system for use in imaging-based solid-phase genotyping and sequencing applications, the illumination system comprising: a) a light source; and b) a liquid light guide configured to collect light emitted by the light source and deliver it to a designated illumination field on a support surface containing a tethered biological macromolecule.
[0320] In some embodiments, the illumination system further comprises a focusing lens. 2 In some embodiments, the light delivered to the designated illumination field is of uniform intensity across a field of view designated for an imaging system used to acquire an image of the support surface. In some embodiments, the designated field of view is at least 2 mm 2 In some embodiments, the light delivered to the designated illumination field is of uniform intensity across the designated field of view when the coefficient of variation (CV) of the light intensity is less than 10%. In some embodiments, the light delivered to the designated illumination field is of uniform intensity across the designated field of view when the coefficient of variation (CV) of the light intensity is less than 5%. In some embodiments, the light delivered to the designated illumination field has a speckle contrast value of less than 0.1. In some embodiments, the light delivered to the designated illumination field has a speckle contrast value of less than 0.05.
[0321] Imaging Modules and Systems: It will be understood by those skilled in the art that the disclosed optical systems, imaging systems, or modules may, in some cases, be standalone optical systems designed to image a sample or substrate surface. In some cases, they may include one or more processors or computers. In some cases, they may include one or more software packages that provide instrument control and / or image processing functions. In some cases, in addition to optical components such as light sources (e.g., solid-state lasers, dye lasers, diode lasers, arc lamps, tungsten halogen lamps, etc.), lenses, prisms, mirrors, dichroic reflectors, optical filters, optical bandpass filters, apertures, and image sensors (e.g., complementary metal oxide semiconductor (CMOS) image sensors and cameras, charge-coupled device (CCD) image sensors and cameras, etc.), they may also include mechanical and / or opto-mechanical components such as XY translation stages, XYZ translation stages, piezoelectric focusing mechanisms, etc. In some cases, they may function as modules, components, subassemblies, or subsystems of a larger system designed for genomics applications (e.g., genetic testing and / or nucleic acid sequencing applications). For example, in some cases they may function as modules, components, subassemblies, or subsystems of a larger system that further includes light-tight and / or other environmental control housings, temperature control modules, fluid control modules, fluid dispensing robotics, pick-and-place robotics, one or more processors or computers, one or more local and / or cloud-based software packages (e.g., instrument / system control software packages, image processing software packages, data analysis software packages), data storage modules, data communications modules (e.g., Bluetooth, WiFi, intranet, or internet communications hardware and associated software), display modules, or any combination thereof.
[0322] Methods for Sequencing Some embodiments of the present disclosure provide a method for sequencing immobilized or non-immobilized template molecules. The method can be operated in the system 100, for example, in the sequencer 114. In some embodiments, the immobilized template molecules include a plurality of nucleic acid template molecules having one copy of a target sequence of interest. In some embodiments, the nucleic acid template molecules having one copy of a target sequence of interest can be generated by performing bridge amplification using linear library molecules. In some embodiments, the immobilized template molecules include a plurality of nucleic acid template molecules (e.g., concatemers) each having two or more tandem copies of a target sequence of interest. In some embodiments, the nucleic acid template molecules including concatemer molecules can be generated by performing rolling circle amplification of circularized linear library molecules. In some embodiments, the non-immobilized template molecules include circular molecules. In some embodiments, the method for sequencing uses a soluble (e.g., non-immobilized) sequencing polymerase or a sequencing polymerase immobilized on a support.
[0323] In some embodiments, the sequencing reaction uses detectably labeled nucleotide analogues. In some embodiments, the sequencing reaction uses a two-step sequencing reaction that includes binding to detectably labeled multivalent molecules and incorporating nucleotide analogues. In some embodiments, the sequencing reaction uses unlabeled nucleotide analogues. In some embodiments, the sequencing reaction uses phosphate-chain labeled nucleotides.
[0324] In some embodiments, the immobilized concatemers each include a tandem repeat unit of a sequence of interest (e.g., an insert region) and an optional adapter sequence. For example, the tandem repeat unit includes (i) a left universal adapter sequence (e.g., a surface pinning primer) having a binding sequence for a first surface primer (720), (ii) a left universal adapter sequence having a binding sequence for a first sequencing primer (740) (e.g., a forward sequencing primer), (iii) a sequence of interest (710), (iv) a right universal adapter sequence having a binding sequence for a second sequencing primer (750) (e.g., a reverse sequencing primer), (v) a right universal adapter sequence having a binding sequence for a second surface primer (730) (e.g., a surface capture primer), and (vii) a left sample index sequence (760) and / or a right sample index sequence (770). In some embodiments, the tandem repeat unit further comprises a left unique identifier sequence (780) and / or a right unique identifier sequence (790). In some embodiments, the tandem repeat unit further comprises at least one binding sequence for a compaction oligonucleotide. In some embodiments, Figures 6 and 7 show units of a linear library molecule or a concatemer molecule.
[0325] The immobilized concatemers can self-disintegrate into compact nucleic acid nanoballs. The inclusion of one or more compaction oligonucleotides in the RCA reaction can further compact the size and / or shape of the nanoballs. An increase in the number of tandem repeat units in a given concatemer increases the number of sites along the concatemer for hybridizing to multiple sequencing primers (e.g., sequencing primers with universal sequences) that serve as multiple initiation sites for polymerase-catalyzed sequencing reactions. When the sequencing reaction uses detectably labeled nucleotides and / or detectably labeled multivalent molecules (e.g., with nucleotide units), the signals emitted by the nucleotides or nucleotide units involved in parallel sequencing reactions along the concatemers result in increased signal intensity for each concatemer. Multiple portions of a given concatemer can be sequenced simultaneously. Furthermore, multiple binding complexes can be formed along a particular concatemer molecule, each binding complex comprising a sequencing polymerase bound to a template / primer duplex and bound to a multivalent molecule, and the multiple binding complexes remain stable without dissociation, resulting in an increased duration, which increases signal intensity and reduces imaging time.
[0326] Methods for sequencing using nucleotide analogues Some aspects of the present disclosure provide methods for sequencing any of the immobilized template molecules described herein, comprising contacting (a) a sequencing polymerase with (i) a nucleic acid template molecule and (ii) a nucleic acid sequencing primer, wherein the contacting is performed under conditions suitable for binding the sequencing polymerase to the nucleic acid template molecule that hybridizes to the nucleic acid primer, and the nucleic acid template molecule that hybridizes to the nucleic acid primer forms a nucleic acid duplex. In some aspects, the sequencing polymerase comprises a recombinant mutant sequencing polymerase capable of binding and incorporating nucleotide analogs.
[0327] In some embodiments, in the method for sequencing the template molecule, the sequencing primer comprises a 3' extendable end or a 3' non-extendable end. In some embodiments, the plurality of nucleic acid template molecules comprises amplified template molecules (e.g., clonally amplified template molecules). In some embodiments, the plurality of nucleic acid template molecules comprises one copy of a target sequence of interest. In some embodiments, the plurality of nucleic acid molecules comprises two or more tandem copies of a target sequence of interest (e.g., concatemers). In some embodiments, the plurality of nucleic acid template molecules comprises the same target sequence of interest or different target sequences of interest. In some embodiments, the plurality of nucleic acid primers are in solution or immobilized on a support. In some embodiments, when the plurality of nucleic acid template molecules and / or the plurality of nucleic acid primers are immobilized on a support, binding with the first sequencing polymerase generates a plurality of immobilized first hybridized polymerases. In some embodiments, the plurality of nucleic acid template molecules and / or the nucleic acid primers are immobilized at 102 to 1015 different sites on the support. In some embodiments, binding of the plurality of template molecules and nucleic acid primers with the plurality of first sequencing polymerases generates a plurality of first multiplexed polymerases immobilized at 102 to 1015 different sites on the support. In some embodiments, the plurality of immobilized first multiplexed polymerases on the support are immobilized at pre-determined sites or random sites on the support. In some embodiments, the plurality of immobilized first multiplexed polymerases are in fluid communication with each other to allow a solution of reagents (e.g., sequencing polymerases, enzymes including multivalent molecules, nucleotides, and / or divalent cations) to be flowed over the support, thereby allowing the plurality of immobilized multiplexed polymerases on the support to react with the solution of reagents in a massively parallel manner.
[0328] In some embodiments, the method for sequencing further comprises step (b): contacting a sequencing polymerase with the plurality of nucleotides under conditions suitable for binding at least one nucleotide to the sequencing polymerase that binds to the nucleic acid duplex and suitable for polymerase-catalyzed nucleotide incorporation that extends the sequencing primer by one nucleotide. In some embodiments, the sequencing polymerase is contacted with the plurality of nucleotides in the presence of at least one catalytic cation comprising magnesium and / or manganese. In some embodiments, the plurality of nucleotides comprises at least one nucleotide analog having a chain-terminating moiety at the sugar 2' or 3' position. In some embodiments, the chain-terminating moiety is removable from the sugar 2' or 3' position to convert the chain-terminating moiety to an OH or H group. In some embodiments, the plurality of nucleotides comprises at least one nucleotide lacking a chain-terminating moiety. In some embodiments, at least the nucleotide is labeled with a detectable reporter moiety (e.g., a fluorophore) that emits a detectable signal. The detectable reporter moiety comprises a fluorophore. In some embodiments, the fluorophore is attached to a rheobase that is eluted. In some embodiments, the fluorophore is attached to the nucleobase with a linker, and the linker is cleavable / removable from the base. In some embodiments, at least one of the nucleotides in the plurality of nucleotides is not labeled with a detectable reporter moiety. In some embodiments, the specific detectable reporter moiety (e.g., fluorophore) attached to the nucleotide can correspond to the nucleotide base (e.g., dATP, dGTP, dCTP, dTTP, or dUTP) to allow for detection and identification of the nucleobase. When the incorporated chain-terminating nucleotide is detectably labeled, step (b) further comprises detecting a signal emitted from the incorporated chain-terminating nucleotide. In some embodiments, step (b) further comprises identifying the nucleobase of the incorporated chain-terminating nucleotide.
[0329] In some embodiments, the method for sequencing further comprises step (c): removing a chain-terminating moiety from the incorporated chain-terminating nucleotide to generate an extendable 3'OH group. In some embodiments, step (c) further comprises removing a detectable label from the incorporated chain-terminating nucleotide. In some embodiments, the sequencing polymerase remains bound to the template molecule hybridized to the sequencing primer that is extended by one nucleobase.
[0330] In some embodiments, the method for sequencing further comprises step (d): repeating steps (b) and (c) at least one time.
[0331] A two-step method for nucleic acid sequencing Some embodiments of the present disclosure provide a two-step method for sequencing any of the immobilized template molecules described herein.In some embodiments, the first step generally includes binding a multivalent molecule to a complex polymerase to form a multivalent complex polymerase and detecting the multivalent complex polymerase.
[0332] In some embodiments, the first step comprises (a) contacting a plurality of first sequencing polymerases with (i) a plurality of nucleic acid template molecules and (ii) a plurality of nucleic acid sequencing primers, the contacting being performed under conditions suitable for binding the plurality of first sequencing polymerases to the plurality of nucleic acid template molecules and the plurality of nucleic acid primers, thereby forming a plurality of first multiplexed polymerases, each comprising a first sequencing polymerase that binds to a nucleic acid duplex, the nucleic acid duplex comprising a nucleic acid template molecule that hybridizes to the nucleic acid primer. In some embodiments, the first polymerase comprises a recombinant mutant sequencing polymerase.
[0333] In some embodiments, in the method for sequencing the template molecule, the sequencing primer comprises an oligonucleotide having a 3' extendable end or a 3' non-extendable end. In some embodiments, the plurality of nucleic acid template molecules comprises amplified template molecules (e.g., clonally amplified template molecules). In some embodiments, the plurality of nucleic acid template molecules comprises one copy of a target sequence of interest. In some embodiments, the plurality of nucleic acid molecules comprises two or more tandem copies of a target sequence of interest (e.g., concatemers). In some embodiments, the nucleic acid template molecules in the plurality of nucleic acid template molecules comprise the same target sequence of interest or different target sequences of interest. In some embodiments, the plurality of nucleic acid template molecules and / or the plurality of nucleic acid primers are in solution or immobilized on a support. In some embodiments, when the plurality of nucleic acid template molecules and / or the plurality of nucleic acid primers are immobilized on a support, binding with the first sequencing polymerase generates a plurality of immobilized first hybridized polymerases. In some embodiments, the plurality of nucleic acid template molecules and / or the nucleic acid primers are immobilized at 102 to 1015 different sites on the support. In some embodiments, binding of the plurality of template molecules and nucleic acid primers with the plurality of first sequencing polymerases generates a plurality of first multiplexed polymerases immobilized at 102 to 1015 different sites on the support. In some embodiments, the plurality of immobilized first multiplexed polymerases on the support are immobilized at pre-determined sites or random sites on the support. In some embodiments, the plurality of immobilized first multiplexed polymerases are in fluid communication with each other to allow a solution of reagents (e.g., sequencing polymerases, enzymes including multivalent molecules, nucleotides, and / or divalent cations) to be flowed over the support, thereby allowing the plurality of immobilized multiplexed polymerases on the support to react with the solution of reagents in a massively parallel manner.
[0334] In some embodiments, the method for sequencing further comprises step (b): contacting a plurality of first multiplexed polymerases with a plurality of multivalent molecules to form a plurality of multivalent multiplexed polymerases (e.g., bound complexes). In some embodiments, each multivalent molecule in the plurality of multivalent molecules comprises a core attached to a plurality of nucleotide arms, each of the nucleotide arms being attached to a nucleotide (e.g., a nucleotide unit) (e.g., Figures 9-13). In some embodiments, the contacting of step (b) is performed under conditions suitable for binding a complementary nucleotide unit of the multivalent molecule to at least two of the plurality of first multiplexed polymerases, thereby forming a plurality of multivalent multiplexed polymerases. In some embodiments, the conditions are suitable for inhibiting polymerase-catalyzed incorporation of the complementary nucleotide unit into the primer of the plurality of multivalent multiplexed polymerases. In some embodiments, the plurality of multivalent molecules includes at least one multivalent molecule having multiple nucleotide arms (e.g., Figures 9-12), each of which includes at least one multivalent molecule attached with a nucleotide analog (e.g., a nucleotide analog unit), the nucleotide analog comprising a chain-terminating moiety at the sugar 2' and / or 3' position. In some embodiments, the plurality of multivalent molecules includes at least one multivalent molecule comprising multiple nucleotide arms, each of which is attached with a nucleotide unit lacking a chain-terminating moiety. In some embodiments, at least one of the multivalent molecules in the plurality of multivalent molecules is labeled with a detectable reporter moiety that emits a signal. In some embodiments, the detectable reporter moiety comprises a fluorophore. In some embodiments, the contacting of step (b) is performed in the presence of at least one non-catalytic cation comprising strontium, barium, and / or calcium.
[0335] In some embodiments, the method for sequencing further comprises step (c): detecting the multiple multivalent complex polymerases. In some embodiments, the detecting comprises detecting a signal emitted by the multivalent molecule that binds to the complex polymerase, where the complementary nucleotide units of the multivalent molecule bind to the primers, but incorporation of the complementary nucleotide units is inhibited. In some embodiments, the multivalent molecule is labeled with a detectable reporter moiety to enable detection. In some embodiments, the labeled multivalent molecule comprises a fluorophore attached to the core, linker, and / or nucleotide units of the multivalent molecule.
[0336] In some embodiments, the method for sequencing further comprises step (d): identifying complementary bases of the complementary nucleotide units bound to the plurality of first multiplexed polymerases, thereby determining the sequence of the nucleic acid template. In some embodiments, the multivalent molecule is labeled with a detectable reporter moiety that corresponds to a particular nucleotide unit attached to the nucleotide arm, so as to enable identification of the complementary nucleotide units (e.g., the nucleotide bases adenine, guanine, cytosine, thymine, or uracil) bound to the plurality of first multiplexed polymerases.
[0337] In some embodiments, the method for sequencing further comprises the step (e): dissociating the plurality of multivalent composite polymerases and removing the plurality of first sequencing polymerases and their bound multivalent molecules, retaining the plurality of nucleic acid duplexes.
[0338] In some embodiments, the second step of the two-step sequencing method generally comprises nucleotide incorporation. In some embodiments, the method for sequencing further comprises step (f): contacting the plurality of retained nucleic acid duplexes of step (e) with a plurality of second sequencing polymerases, wherein the contacting is performed under conditions suitable for the plurality of second sequencing polymerases to bind to the plurality of retained nucleic acid duplexes, thereby forming a plurality of second hybrid polymerases, each of the plurality of second hybrid polymerases comprising a second sequencing polymerase bound to a nucleic acid duplex. In some embodiments, the second sequencing polymerase comprises a recombinant mutant polymerase.
[0339] In some embodiments, the plurality of first sequencing polymerases in step (a) have an amino acid sequence that is 100% identical to the amino acid sequence of the plurality of second sequencing polymerases in step (f). In some embodiments, the plurality of first sequencing polymerases in step (a) have an amino acid sequence that is different from the amino acid sequence of the plurality of second sequencing polymerases in step (f).
[0340] In some embodiments, the method for sequencing further comprises step (g): contacting a plurality of second hybrid polymerases with a plurality of nucleotides, the contacting being performed under conditions suitable for binding complementary nucleotides from the plurality of nucleotides to at least two of the second hybrid polymerases, thereby forming a plurality of nucleotide hybrid polymerases. In some embodiments, the contacting in step (g) is performed under conditions suitable for promoting polymerase-catalyzed incorporation of the bound complementary nucleotides into the primer by the nucleotide hybrid polymerase, thereby extending the sequencing primer by one nucleobase. In some embodiments, the incorporation of nucleotides into the 3' end of the sequencing primer in step (g) comprises a primer extension reaction. In some embodiments, the contacting in step (g) is performed in the presence of at least one catalytic cation comprising magnesium and / or manganese. In some embodiments, the plurality of nucleotides comprises natural nucleotides (e.g., non-analog nucleotides) or nucleotide analogs. In some embodiments, the plurality of nucleotides comprises removable or non-removable 2' and / or 3' chain termination moieties. In some embodiments, at least one of the nucleotides in the plurality of nucleotides is not labeled with a detectable reporter moiety. In some embodiments, the plurality of nucleotides is unlabeled. In some embodiments, the plurality of nucleotides comprises a plurality of nucleotides labeled with a detectable reporter moiety. The detectable reporter moiety comprises a fluorophore. In some embodiments, the fluorophore is attached to the nucleotide base. In some embodiments, the fluorophore is attached to the nucleotide base with a linker, the linker being cleavable / removable from the base or not removable from the base.In some embodiments, a specific detectable reporter moiety (e.g., a fluorophore) attached to a nucleotide can correspond to a nucleotide base (e.g., dATP, dGTP, dCTP, dTTP, or dUTP) to allow for detection and discrimination of the nucleotide base.
[0341] In some embodiments, when the plurality of nucleotides in step (g) are detectably labeled, the method for sequencing further comprises step (h): detecting the complementary nucleotide incorporated into the primer of the nucleotide-complexed polymerase. In some embodiments, the plurality of nucleotides are labeled with a detectable reporter moiety to allow detection. In some embodiments, when the plurality of nucleotides in step (g) are unlabeled, the detection of step (h) is omitted.
[0342] In some embodiments, when the plurality of nucleotides in step (g) are detectably labeled, the method for sequencing further comprises step (i): identifying the base of the complementary nucleotide incorporated into the primer of the nucleotide-complexed polymerase. In some embodiments, the identification of the incorporated complementary nucleotide in step (i) can be used to confirm the identity of the complementary nucleotide of the multivalent molecule bound to the plurality of first composite polymerases in step (d). In some embodiments, the identifying in step (i) can be used to determine the sequence of the nucleic acid template molecule. In some embodiments, when the plurality of nucleotides in step (g) are unlabeled, the identifying in step (i) is omitted.
[0343] In some embodiments, the method for sequencing further comprises step (j): removing a chain terminating portion from the incorporated nucleotide when step (g) is performed by contacting a plurality of second hybrid polymerases with a plurality of nucleotides comprising at least one nucleotide having a 2' and / or 3' chain terminating portion.
[0344] In some embodiments, the method for sequencing further comprises step (k): repeating steps (a) through (j) at least once. In some embodiments, the sequence of the nucleic acid template molecule may be determined in steps (c) and (d) by detecting and identifying multivalent molecules that bind to the sequencing polymerase but are not incorporated into the 3' end of the primer. In some embodiments, the sequence of the nucleic acid template molecule may be determined (or confirmed) in steps (h) and (i) by detecting and identifying nucleotides that are incorporated into the 3' end of the primer.
[0345] In some embodiments, in any of the methods for sequencing a nucleic acid molecule, combining a plurality of first multiplexed polymerases with a plurality of multivalent molecules forms at least one avidity complex, the method comprising the steps of: (a) combining a first nucleic acid primer, a first sequencing polymerase, and a first multivalent molecule with a first portion of a concatemeric template molecule, thereby forming a first binding complex, where a first nucleotide unit of the first multivalent molecule binds to the first sequencing polymerase; and (b) combining a second nucleic acid primer, a second sequencing polymerase, and a first multivalent molecule with a second portion of the same concatemeric template molecule, thereby forming a second binding complex, where a second nucleotide unit of the first multivalent molecule binds to the second sequencing polymerase, and the first and second binding complexes comprising the same multivalent molecule form an avidity complex. In some embodiments, the first sequencing polymerase comprises any wild-type or mutant polymerase described herein. In some embodiments, the second sequencing polymerase comprises any wild-type or mutant polymerase described herein. The concatemer template molecule comprises a tandem repeat sequence of a sequence of interest and at least one universal sequencing primer binding site. The first and second nucleic acid primers can bind to the sequencing primer binding sites along the concatemer template molecule. Exemplary multivalent molecules are shown in Figures 9-12.
[0346] In some embodiments, in any of the methods for sequencing a nucleic acid molecule, the method comprises combining a plurality of first composite polymerases with a plurality of multivalent molecules to form at least one avidity complex, the method comprising the steps of: (a) contacting a plurality of sequencing polymerases and a plurality of nucleic acid primers with different portions of a concatemeric nucleic acid concatemeric molecule to form at least a first and a second composite polymerase on the same concatemeric molecule; and (b) contacting the plurality of multivalent molecules with at least a first and a second composite polymerase on the same concatemeric template molecule under conditions suitable for binding a single multivalent molecule from the plurality of multivalent molecules to the first and second composite polymerases, wherein at least a first nucleotide unit of the single multivalent molecule comprises a first primer that hybridizes to a first portion of the concatemeric template molecule, thereby forming a first binding complex (e.g., a first ternary complex). (c) contacting a second composite polymerase comprising a second primer that binds to the first composite polymerase and where at least a second nucleotide unit of the single multivalent molecule hybridizes to a second portion of the concatemeric template molecule, thereby forming a second binding complex (e.g., a second ternary complex), under conditions suitable to inhibit polymerase-catalyzed incorporation of the bound first and second nucleotide units in the first and second binding complexes, whereby the first and second binding complexes that bind to the same multivalent molecule form an avidity complex; (d) identifying the first nucleotide unit in the first binding complex, thereby determining the sequence of the first portion of the concatemeric template molecule, and identifying the second nucleotide unit in the second binding complex, thereby determining the sequence of the second portion of the concatemeric template molecule. In some embodiments, the plurality of sequencing polymerases comprises any wild-type or mutant sequencing polymerase described herein.The concatemer template molecule comprises a tandem repeat sequence of a sequence of interest and at least one universal sequencing primer binding site. A plurality of nucleic acid primers can bind to the sequencing primer binding sites along the concatemer template molecule. Exemplary multivalent molecules are shown in Figures 10-13.
[0347] Sequencing by ligation Some aspects of the present disclosure provide a method for sequencing any of the immobilized template molecules described herein, the sequencing method comprising a sequencing by binding (SBB) procedure using non-labeled chain terminating nucleotides. In some aspects, the sequencing by binding (SBB) comprises: (a) sequentially contacting a primed template nucleic acid with at least two separate mixtures under ternary complex stabilizing conditions, each of the at least two separate mixtures comprising a polymerase and a nucleotide, whereby the sequential contacting results in the primed template nucleic acid being contacted with nucleotide analogs of the first, second, and third base types in the template under ternary complex stabilizing conditions; and (b) examining the at least two separate mixtures to determine whether a ternary complex has formed. (c) identifying the next correct nucleotide of the primed template nucleic acid molecule, where the next correct nucleotide is identified as a first, second, or third base type cognate if a ternary complex is detected in step (b), and the next correct nucleotide is presumed to be a fourth base type nucleotide cognate based on the absence of a ternary complex in step (b), (d) adding the next correct nucleotide to the primed template nucleic acid after step (b), thereby producing an extension primer, and (e) repeating steps (a)-(d) on the primed template nucleic acid containing the extension primer. Exemplary sequencing-by-binding methods are described in U.S. Patent Nos. 10,246,744 and 10,731,141, the entire contents of both of which are incorporated herein by reference.
[0348] Methods for sequencing using phosphate-chain-labeled nucleotides Some embodiments of the present disclosure provide a method for sequencing using an immobilized sequencing polymerase that binds a non-immobilized template molecule, where the sequencing reaction is performed with phosphate-chain-labeled nucleotides. In some embodiments, the sequencing method includes step (a): providing a support on which a plurality of sequencing polymerases are immobilized. In some embodiments, the sequencing polymerase includes a processive DNA polymerase. In some embodiments, the sequencing polymerase includes a wild-type or mutant DNA polymerase, including, for example, Phi29 DNA polymerase. In some embodiments, the support includes a plurality of separate compartments, and the sequencing polymerase is immobilized at the bottom of the compartment. In some embodiments, the separate compartment includes a silica bottom that allows light to penetrate. In some embodiments, the separate compartment includes a silica bottom that is composed of a nanophotonic confinement structure that includes a hole in a metal cladding film (e.g., an aluminum cladding film). In some embodiments, the hole in the metal cladding has a small opening, for example, approximately 70 nm. In some embodiments, the height of the nanophotonic confinement structure is approximately 100 nm. In some embodiments, the nanophotonic confinement structure comprises a zero mode waveguide (ZMW). In some embodiments, the nanophotonic confinement structure comprises a liquid.
[0349] In some embodiments, the sequencing method further comprises step (b): contacting a plurality of immobilized sequencing polymerases with a plurality of single-stranded circular nucleic acid template molecules and a plurality of oligonucleotide sequencing primers under conditions in which each immobilized sequencing polymerase is suitable for binding to the single-stranded circular template molecules and each sequencing primer is suitable for hybridizing to each single-stranded circular template molecule, thereby generating a plurality of polymerase / template / primer complexes. In some embodiments, each sequencing primer hybridizes to a universal sequencing primer binding site on the single-stranded circular template molecule.
[0350] In some embodiments, the sequencing method further comprises step (c): contacting the plurality of polymerase / template / primer complexes with a plurality of phosphate-strand-labeled nucleotides, each of which comprises an aromatic base, a five-carbon sugar (e.g., ribose or deoxyribose), and a phosphate strand comprising 3-20 phosphate groups, wherein the terminal phosphate group is attached to a detectable reporter moiety (e.g., a fluorophore). The first, second, and third phosphate groups may be referred to as alpha, beta, and gamma phosphate groups. In some embodiments, the particular detectable reporter moiety attached to the terminal phosphate group corresponds to a nucleobase (e.g., dATP, dGTP, dCTP, dTTP, or dUTP) to allow for detection and discrimination of the nucleobase. In some embodiments, the plurality of polymerase / template / primer complexes are contacted with the plurality of phosphate-strand-labeled nucleotides under conditions suitable for polymerase-catalyzed nucleotide incorporation. In some embodiments, the sequencing polymerase is capable of binding a complementary phosphate strand-labeled nucleotide and incorporating a complementary nucleotide opposite the nucleotide in the template molecule. In some embodiments, the polymerase-catalyzed nucleotide incorporation reaction cleaves between the alpha and beta phosphate groups, thereby releasing a polyphosphate strand linked to a fluorophore.
[0351] In some embodiments, the sequencing method further comprises step (d): detecting a fluorescent signal emitted by a phosphate-chain-labeled nucleotide that is bound by the sequencing polymerase and incorporated onto the end of the sequencing primer. In some embodiments, step (d) further comprises identifying the phosphate-chain-labeled nucleotide that is bound by the sequencing polymerase and incorporated onto the end of the sequencing primer.
[0352] In some embodiments, the sequencing method further comprises step (d): repeating steps (c) through (d) at least once. In some embodiments, the sequencing method using phosphate-chain-labeled nucleotides can be performed according to the methods described in U.S. Patent Nos. 7,170,050, 7,302,146, and / or 7,405,281.
[0353] Sequencing Polymerase Some aspects of the present disclosure provide a method for sequencing a nucleic acid molecule, wherein any of the sequencing methods described herein use at least one type of sequencing polymerase and a plurality of nucleotides, or at least one type of sequencing polymerase and a plurality of nucleotides and a plurality of multivalent molecules. In some embodiments, the sequencing polymerase(s) can incorporate a complementary nucleotide opposite the nucleotide in the template molecule. In some embodiments, the sequencing polymerase(s) can bind a complementary nucleotide unit of the multivalent molecule opposite the nucleotide in the template molecule. In some embodiments, the plurality of sequencing polymerases comprises recombinant mutant polymerases.
[0354] Examples of polymerases suitable for use in sequencing with nucleotides and / or polyvalent molecules include Klenow DNA polymerase; Thermus aquaticus DNA polymerase I (Taq polymerase); KlenTaq polymerase; Candidatus altiarchaeales archaea; Candidatus Hadarchaeum Yellowstonense; Hadesarchaea archaea; Euryarchaeota archaea; Thermoplasmata archaea; Thermococcus polymerases, such as Thermococcus litoralis, bacteriophage T7 DNA polymerase; human alpha, delta, and epsilon DNA polymerases; bacteriophage polymerases, such as T4, RB69, and phi29 bacteriophage DNA polymerases; Pyrococcus furiosus DNA polymerase (Pfu polymerase); Bacillus subtilis DNA polymerase III; E. coli DNA polymerase III alpha and epsilon; 9 degree Examples of DNA polymerases include, but are not limited to, N polymerase; reverse transcriptases, such as HIV type M or O reverse transcriptase; avian myeloblastosis virus reverse transcriptase; Moloney murine leukemia virus (MMLV) reverse transcriptase; or telomerase. Further non-limiting examples of DNA polymerases include those from various archaeal genera, such as Aeropyrum, Archaeglobus, Desulfurococcus, Pyrobaculum, Pyrococcus, Pyrolobus, Pyrodictium, Staphylothermus, Stetteria, Sulfolobus, Thermococcus, and Vulcanisaeta, or variants thereof, including such polymerases known in the art, such as 9 degrees N, VENT, DEEP VENT, THERMINATOR, Pfu, KOD, Pfx, Tgo, and RB69 polymerases.
[0355] nucleotide Some aspects of the present disclosure provide a method for sequencing a nucleic acid molecule, in which any of the sequencing methods described herein uses at least one nucleotide. The nucleotide comprises a base, a sugar, and at least one phosphate group. In some embodiments, at least one nucleotide in the plurality of nucleotides comprises an aromatic base, a five-carbon sugar (e.g., ribose or deoxyribose), and one or more phosphate groups (e.g., 1-10 phosphate groups). The plurality of nucleotides can comprise at least one type of nucleotide selected from the group consisting of dATP, dGTP, dCTP, dTTP, and dUTP. The plurality of nucleotides can comprise in any combination mixture of two or more types of nucleotides selected from the group consisting of dATP, dGTP, dCTP, dTTP, and / or dUTP. In some embodiments, at least one nucleotide in the plurality of nucleotides is not a nucleotide analog. In some embodiments, at least one nucleotide in the plurality of nucleotides comprises a nucleotide analog.
[0356] In some embodiments, in any of the methods for sequencing nucleic acid molecules described herein, at least one nucleotide of the plurality of nucleotides comprises a chain of 1, 2, or 3 phosphorus atoms, the chain typically being attached to the 5' carbon of the sugar moiety via an ester or phosphoramide bond. In some embodiments, at least one nucleotide in the plurality of nucleotides is an analog having a phosphorus chain, in which the phosphorus atoms are linked together with an intervening O, S, NH, methylene, or ethylene. In some embodiments, the phosphorus atoms in the chain comprise a substituted side chain group comprising O, S, or BH3. In some embodiments, the chain comprises a phosphate group substituted with an analog comprising phosphoramidate, phosphorothioate, phosphorodithioate, and O-methyl phosphoramidate groups.
[0357] In some embodiments, in any of the methods for sequencing a nucleic acid molecule described herein, at least one nucleotide in the plurality of nucleotides comprises a terminator nucleotide analog, the terminator nucleotide analog having a chain terminating moiety (e.g., a blocking moiety) at the sugar 2' position, the sugar 3' position, or the sugar 2' and 3' positions. In some embodiments, the chain terminating moiety can inhibit polymerase-catalyzed incorporation of a subsequent nucleotide unit or free nucleotide in the nascent chain during a primer extension reaction. In some embodiments, the chain terminating moiety is attached to the 3' sugar hydroxyl position, where the sugar comprises a ribose or deoxyribose sugar moiety. In some embodiments, the chain terminating moiety is removable / cleavable from the 3' sugar position to generate a nucleotide having a 3' OH sugar group that is extendable with a subsequent nucleotide in a polymerase-catalyzed nucleotide incorporation reaction. In some embodiments, the chain terminating moiety comprises an alkyl group, an alkenyl group, an alkynyl group, an allyl group, an aryl group, a benzyl group, an azide group, an amine group, an amide group, a keto group, an isocyanate group, a phosphate group, a thio group, a disulfide group, a carbonate group, a urea group, a silyl group, or an acetal group. In some embodiments, the chain terminating moiety is cleavable / removable from the nucleotide, for example, by reacting the chain terminating moiety with a chemical agent, a pH change, light, or heat. In some embodiments, the chain terminating moieties alkyl, alkenyl, alkynyl, and allyl are cleavable with tetrakis(triphenylphosphine)palladium(0) (Pd(PPh3)4) with piperidine or with 2,3-dichloro-5,6-dicyano-1,4-benzo-quinone (DDQ). In some embodiments, the chain terminating moieties aryl and benzyl are cleavable with HPd / C. In some embodiments, the chain terminating moieties amine, amide, keto, isocyanate, phosphate, thio, disulfide are cleavable with phosphines or thiol groups, including beta mercaptoethanol or dithiothritol (DTT).In some embodiments, the carbonate chain terminating moiety can be cleaved with potassium carbonate (K2CO3) in MeOH, triethylamine in pyridine, or Zn(AcOH) in acetic acid. In some embodiments, the urea and silyl chain terminating moieties can be cleaved with tetrabutylammonium fluoride, pyridine-HF, ammonium fluoride, or triethylamine trihydrofluoride. In some embodiments, the chain terminating moiety can be cleaved / removed with nitrous acid. In some embodiments, the chain terminating moiety can be cleaved / removed using a solution containing nitrite, for example, a combination of nitrite and an acid, for example, acetic acid, sulfuric acid, or nitric acid. In some further embodiments, the solution can include an organic acid.
[0358] In some embodiments, in any of the methods for sequencing a nucleic acid molecule described herein, at least one nucleotide in the plurality of nucleotides comprises a terminator nucleotide analog, the terminator nucleotide analog having a chain terminating moiety (e.g., a blocking moiety) at the sugar 2' position, the sugar 3' position, or the sugar 2' and 3' positions. In some embodiments, the chain terminating moiety comprises an azide, azido, and azidomethyl group. In some embodiments, the chain terminating moiety comprises a 3'-O-azido or a 3'-O-azidomethyl group. In some embodiments, the chain terminating moiety azide, azido, and azidomethyl group are cleavable / removable with a phosphine compound. In some embodiments, the phosphine compound comprises a derivatized trialkylphosphine moiety or a derivatized triarylphosphine moiety. In some embodiments, the phosphine compound comprises tris(2-carboxyethyl)phosphine (TCEP), or bis-sulfotriphenylphosphine (BS-TPP), or tris(hydroxypropyl)phosphine (THPP). In some embodiments, the cleavage agent comprises 4-dimethylaminopyridine (4-DMAP). In some embodiments, chain terminating moieties comprising one or more of a 3'-O-amino group, a 3'-O-aminomethyl group, a 3'-O-methylamino group, or derivatives thereof can be cleaved with nitrous acid via a mechanism utilizing nitrous acid or using a solution comprising nitrous acid. In some embodiments, chain terminating moieties comprising one or more of a 3'-O-amino group, a 3'-O-aminomethyl group, a 3'-O-methylamino group, or derivatives thereof can be cleaved using a solution comprising a nitrite. In some embodiments, for example, the nitrite can be combined with or contacted with an acid such as acetic acid, sulfuric acid, or nitric acid. In some further embodiments, for example, the nitrite may be combined with or contacted with an organic acid, such as, for example, formic acid, acetic acid, propionic acid, butyric acid, isobutyric acid, etc. In some embodiments, the chain terminating moiety comprises a 3'-acetal moiety that can be cleaved with a palladium deblocking reagent (e.g., Pd(0)).
[0359] In some embodiments, in any of the methods for sequencing a nucleic acid molecule described herein, the nucleotide analogs are 3'-deoxynucleotides, 2',3'-dideoxynucleotides, 3'-methyl, 3'-azido, 3'-azidomethyl, 3'-O-azidoalkyl, 3'-O-ethynyl, 3'-O-aminoalkyl, 3'-O-fluoroalkyl, 3'-fluoromethyl, 3'-difluoromethyl, 3'-trifluoromethyl, 3'-di ... and 3'-sulfonyl, 3'-malonyl, 3'-amino, 3'-O-amino, 3'-sulfhydral, 3'-aminomethyl, 3'-ethyl, 3'butyl, 3'-tertbutyl, 3'-fluorenylmethyloxycarbonyl, 3'tert-butyloxycarbonyl, 3'-O-alkylhydroxylamino groups, 3'-phosphorothioate, 3-O-benzyl, and 3'-O-benzyl, 3-acetal moieties, or derivatives thereof.
[0360] In some embodiments, in any of the methods for sequencing a nucleic acid molecule described herein, the plurality of nucleotides comprises a plurality of nucleotides labeled with a detectable reporter moiety. The detectable reporter moiety comprises a fluorophore. In some embodiments, the fluorophore is attached to the nucleotide base. In some embodiments, the fluorophore is attached to the nucleotide base with a linker, the linker being cleavable / removable from the base. In some embodiments, at least one of the nucleotides in the plurality of nucleotides is not labeled with a detectable reporter moiety. In some embodiments, the particular detectable reporter moiety (e.g., fluorophore) attached to the nucleotide can correspond to the nucleotide base (e.g., dATP, dGTP, dCTP, dTTP, or dUTP) to allow detection and discrimination of the nucleotide base.
[0361] In some embodiments, in any of the methods for sequencing a nucleic acid molecule described herein, the cleavable linker on the nucleotide base comprises a cleavable moiety comprising an alkyl group, an alkenyl group, an alkynyl group, an allyl group, an aryl group, a benzyl group, an azide group, an amine group, an amide group, a keto group, an isocyanate group, a phosphate group, a thio group, a disulfide group, a carbonate group, a urea group, or a silyl group. In some embodiments, the cleavable linker on the base is cleavable / removable from the base by reacting the cleavable moiety with a chemical agent, a pH change, light, or heat. In some embodiments, the cleavable moieties alkyl, alkenyl, alkynyl, and allyl are cleavable with tetrakis(triphenylphosphine)palladium(0) (Pd(PPh3)4) with piperidine or with 2,3-dichloro-5,6-dicyano-1,4-benzo-quinone (DDQ). In some embodiments, the cleavable moieties aryl and benzyl are cleavable with HPd / C. In some embodiments, the cleavable moieties amine, amide, keto, isocyanate, phosphate, thio, disulfide are cleavable with phosphines or thiol groups, including beta-mercaptoethanol or dithiothritol (DTT). In some embodiments, the cleavable moiety carbonate is cleavable with potassium carbonate (K2CO3) in MeOH, triethylamine in pyridine, or Zn(AcOH) in acetic acid. In some embodiments, the cleavable moieties urea and silyl are cleavable with tetrabutylammonium fluoride, pyridine-HF, ammonium fluoride, or triethylamine trihydrofluoride.
[0362] In some embodiments, in any of the methods for sequencing a nucleic acid molecule described herein, the cleavable linker on the nucleotide base comprises a cleavable moiety comprising azide, azido, and azidomethyl groups. In some embodiments, the cleavable moieties azide, azido, and azidomethyl groups are cleavable / removable with a phosphine compound. In some embodiments, the phosphine compound comprises a derivatized trialkylphosphine moiety or a derivatized triarylphosphine moiety. In some embodiments, the phosphine compound comprises tris(2-carboxyethyl)phosphine (TCEP), or bis-sulfotriphenylphosphine (BS-TPP), or tris(hydroxypropyl)phosphine (THPP). In some embodiments, the cleavage agent comprises 4-dimethylaminopyridine (4-DMAP).
[0363] In some embodiments, in any of the methods for sequencing a nucleic acid molecule described herein, the chain terminating moiety (e.g., at the sugar 2' and / or sugar 3' positions) and the cleavable linker on the nucleotide base have the same or different cleavable moieties. In some embodiments, the chain terminating moiety (e.g., at the sugar 2' and / or sugar 3' positions) and the detectable reporter moiety attached to the base are chemically cleavable / removable with the same chemical agent. In some embodiments, the chain terminating moiety (e.g., at the sugar 2' and / or sugar 3' positions) and the detectable reporter moiety attached to the base are chemically cleavable / removable with different chemical agents.
[0364] Multivalent molecules Some embodiments of the present disclosure provide a method for sequencing a nucleic acid molecule, where any of the sequencing methods described herein employs at least one multivalent molecule. In some embodiments, the multivalent molecule comprises a plurality of nucleotide arms attached to a core and having any configuration, including a starburst, helter skelter, or bottle brush configuration (e.g., FIG. 9). The multivalent molecule comprises (1) a core and (2) a plurality of nucleotide arms, the plurality of nucleotide arms comprising (i) a core attachment moiety, (ii) a spacer comprising a PEG moiety, (iii) a linker, and (iv) a nucleotide unit, where the core is attached to the plurality of nucleotide arms, the spacer is attached to the linker, and the linker is attached to the nucleotide unit. In some embodiments, the nucleotide unit comprises a base, a sugar, and at least one phosphate group, and the linker is attached to the nucleotide unit via the base. In some embodiments, the linker comprises an aliphatic chain or an oligoethylene glycol chain, and both linker chains have 2-6 subunits. In some embodiments, the linker also comprises an aromatic moiety. Exemplary nucleotide arms are shown in FIG. 13. Exemplary multivalent molecules are shown in Figures 9-12. An exemplary spacer is shown in Figure 14 (top) and an exemplary linker is shown in Figure 15 (bottom) and Figure 15. Exemplary nucleotides attached to linkers are shown in Figures 16-19. An exemplary biotinylated nucleotide arm is shown in Figure 20.
[0365] In some embodiments, the multivalent molecule comprises a core bound to a plurality of nucleotide arms, the plurality of nucleotide arms having the same type of nucleotide unit selected from the group consisting of dATP, dGTP, dCTP, dTTP, and dUTP.
[0366] In some embodiments, the multivalent molecule comprises a core to which are attached a plurality of nucleotide arms, each arm comprising a nucleotide unit. In some embodiments, the nucleotide unit comprises an aromatic base, a five-carbon sugar (e.g., ribose or deoxyribose), and one or more phosphate groups (e.g., 1-10 phosphate groups). The plurality of multivalent molecules can comprise one type of multivalent molecule having one type of nucleotide unit selected from the group consisting of dATP, dGTP, dCTP, dTTP, and dUTP. The plurality of multivalent molecules can be comprised in any combination mixture of two or more types of multivalent molecules, with each individual multivalent molecule in the mixture comprising a nucleotide unit selected from the group consisting of dATP, dGTP, dCTP, dTTP, and / or dUTP.
[0367] In some embodiments, the nucleotide unit comprises a chain of one, two, or three phosphorus atoms, typically attached to the 5' carbon of the sugar moiety via an ester or phosphoramido bond. In some embodiments, at least one nucleotide unit is a nucleotide analogue with a phosphorus chain, in which the phosphorus atoms are linked together with intervening O, S, NH, methylene, or ethylene. In some embodiments, the phosphorus atoms in the chain comprise substituted side groups, including O, S, or BH3. In some embodiments, the chain comprises phosphate groups substituted with analogues, including phosphoramidate, phosphorothioate, phosphorodithioate, and O-methyl phosphoramidite groups.
[0368] In some embodiments, the multivalent molecule comprises a core attached to multiple nucleotide arms, each nucleotide arm comprising a nucleotide unit that is a nucleotide analog with a chain terminating moiety (e.g., a blocking moiety) at the sugar 2' position, the sugar 3' position, or the sugar 2' and 3' positions. In some embodiments, the nucleotide unit comprises a chain terminating moiety (e.g., a blocking moiety) at the sugar 2' position, the sugar 3' position, or the sugar 2' and 3' positions. In some embodiments, the chain terminating moiety can inhibit polymerase-catalyzed incorporation of a subsequent nucleotide unit or free nucleotide in a nascent chain during a primer extension reaction. In some embodiments, the chain terminating moiety is attached to the 3' sugar hydroxyl position, where the sugar comprises a ribose or deoxyribose sugar moiety. In some embodiments, the chain terminating moiety is removable / cleavable from the 3' sugar position to generate a nucleotide with a 3' OH sugar group that is extendable with a subsequent nucleotide in a polymerase-catalyzed nucleotide incorporation reaction. In some embodiments, the chain terminating moiety comprises an alkyl group, an alkenyl group, an alkynyl group, an allyl group, an aryl group, a benzyl group, an azide group, an amine group, an amide group, a keto group, an isocyanate group, a phosphate group, a thio group, a disulfide group, a carbonate group, a urea group, or a silyl group. In some embodiments, the chain terminating moiety is cleavable / removable from the nucleotide unit, for example, by reacting the chain terminating moiety with a chemical agent, a pH change, light, or heat. In some embodiments, the chain terminating moieties alkyl, alkenyl, alkynyl, and allyl are cleavable with tetrakis(triphenylphosphine)palladium(0) (Pd(PPh3)4) with piperidine or with 2,3-dichloro-5,6-dicyano-1,4-benzo-quinone (DDQ). In some embodiments, the chain terminating moieties aryl and benzyl are cleavable with HPd / C. In some embodiments, the chain terminating moieties amine, amide, keto, isocyanate, phosphate, thio, disulfide are cleavable with phosphines or thiol groups, including beta mercaptoethanol or dithiothritol (DTT).In some embodiments, the carbonate chain terminating moiety can be cleaved with potassium carbonate (K2CO3) in MeOH, triethylamine in pyridine, or Zn(AcOH) in acetic acid. In some embodiments, the urea and silyl chain terminating moieties can be cleaved with tetrabutylammonium fluoride, pyridine-HF, ammonium fluoride, or triethylamine trihydrofluoride.
[0369] In some embodiments, the nucleotide units comprise a chain-terminating moiety (e.g., a blocking moiety) at the sugar 2' position, the sugar 3' position, or the sugar 2' and 3' positions. In some embodiments, the chain-terminating moiety comprises an azide, azido, and azidomethyl group. In some embodiments, the chain-terminating moiety comprises a 3'-O-azido or a 3'-O-azidomethyl group. In some embodiments, the chain-terminating moiety azide, azido, and azidomethyl groups are cleavable / removable with a phosphine compound. In some embodiments, the phosphine compound comprises a derivatized trialkylphosphine moiety or a derivatized triarylphosphine moiety. In some embodiments, the phosphine compound comprises tris(2-carboxyethyl)phosphine (TCEP), or bis-sulfotriphenylphosphine (BS-TPP), or tris(hydroxypropyl)phosphine (THPP). In some embodiments, the cleaving agent comprises 4-dimethylaminopyridine (4-DMAP).
[0370] In some embodiments, the nucleotide units comprise a chain terminating moiety selected from the group consisting of 3'-deoxynucleotides, 2',3'-dideoxynucleotides, 3'-methyl, 3'-azido, 3'-azidomethyl, 3'-O-azidoalkyl, 3'-O-ethynyl, 3'-O-aminoalkyl, 3'-O-fluoroalkyl, 3'-fluoromethyl, 3'-difluoromethyl, 3'-trifluoromethyl, 3'-sulfonyl, 3'-malonyl, 3'-amino, 3'-O-amino, 3'-sulfhydral, 3'-aminomethyl, 3'-ethyl, 3'butyl, 3'-tertbutyl, 3'-fluorenylmethyloxycarbonyl, 3'tert-butyloxycarbonyl, 3'-O-alkylhydroxylamino groups, 3'-phosphorothioates, and 3-O-benzyl, or derivatives thereof.
[0371] In some embodiments, the multivalent molecule comprises a core attached to a plurality of nucleotide arms, the nucleotide arms comprising spacers, linkers and nucleotide units, and the core, linkers and / or nucleotide units are labeled with a detectable reporter moiety. In some embodiments, the detectable reporter moiety comprises a fluorophore. In some embodiments, a particular detectable reporter moiety (e.g., a fluorophore) attached to the multivalent molecule can correspond to a base of a nucleotide unit (e.g., dATP, dGTP, dCTP, dTTP, or dUTP) to allow for detection and discrimination of the nucleotide base.
[0372] In some embodiments, at least one nucleotide arm of the multivalent molecule has a nucleotide unit attached to a detectable reporter moiety. In some embodiments, the detectable reporter moiety is attached to a nucleotide base. In some embodiments, the detectable reporter moiety comprises a fluorophore. In some embodiments, the particular detectable reporter moiety (e.g., a fluorophore) attached to the multivalent molecule can correspond to the base of a nucleotide unit (e.g., dATP, dGTP, dCTP, dTTP, or dUTP) to allow for detection and discrimination of the nucleotide base.
[0373] In some embodiments, the core of the multivalent molecule comprises an avidin-like or streptavidin-like moiety and the core-attached moiety comprises biotin. In some embodiments, the core comprises a streptavidin- or avidin-type moiety, including an avidin protein, as well as any derivatives, analogs, and other non-natural forms of avidin that can bind to at least one biotin moiety. Other forms of avidin moieties include natural and recombinant avidin and streptavidin, as well as derivatized molecules, such as non-glycosylated avidin and truncated streptavidin. For example, avidin moieties include deglycosylated forms of avidin, bacterial streptavidin produced by Streptomyces (e.g., Streptomyces avidinii), as well as derivatized forms, such as N-acylavidins, e.g., N-acetyl, N-phthalyl, and N-succinyl avidin, and the commercially available products EXTRAVIDIN, CAPTAVIDIN, NEUTRAVIDIN, and NEUTRALITE AVIDIN.
[0374] In some embodiments, any of the methods for sequencing a nucleic acid molecule described herein may include forming a binding complex, where the binding complex comprises (i) a polymerase, a nucleic acid template molecule duplexed with a primer, and a nucleotide, or the binding complex comprises (ii) a polymerase, a nucleic acid template molecule duplexed with a primer, and a nucleotide unit of a multivalent molecule. In some embodiments, the binding complex has a duration of greater than about 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, or 1 second. The binding complex has a duration of greater than about 0.1-0.25 seconds, or about 0.25-0.5 seconds, or about 0.5-0.75 seconds, or about 0.75-1 seconds, or about 1-2 seconds, or about 2-3 seconds, or about 3-4 seconds, or about 4-5 seconds, and / or the method is or can be performed at 15° C. or higher, 20° C. or higher, 25° C. or higher, 35° C. or higher, 37° C. or higher, 42° C. or higher, 55° C. or higher, 60° C. or higher, 72° C. or higher, or 80° C. or higher, or within a range defined by any of the foregoing. The binding complex (e.g., ternary complex) remains stable until subjected to conditions that cause dissociation of interactions between the polymerase, the template molecule, the primer, and / or any of the nucleotide units or nucleotides. For example, dissociation conditions include contacting the binding complex with any one of detergent, EDTA, and / or water, or any combination thereof. In some embodiments, the disclosure provides such methods, wherein the binding complex is deposited on, attached to, or hybridized to a surface that in the detection step exhibits a contrast to noise ratio of greater than 20. In some embodiments, the disclosure provides such methods, wherein the contacting is performed under conditions that stabilize the binding complex when the nucleotide or nucleotide unit is complementary to the next base of the template nucleic acid and destabilize the binding complex when the nucleotide or nucleotide unit is not complementary to the next base of the template nucleic acid.
[0375] Compaction Oligonucleotides The compaction oligonucleotide comprises a single-stranded linear oligonucleotide having a 5' region capable of hybridizing to a first portion of a concatemer molecule, and a compaction oligonucleotide having a 3' region capable of hybridizing to a second portion of a concatemer molecule (e.g., the same concatemer molecule). In some embodiments, hybridization of the compaction oligonucleotide to individual concatemer molecules causes the concatemer molecules to collapse or fold into DNA nanoballs that are more compact in shape and size compared to non-collapsed DNA molecules. The spot image of the DNA nanoballs can be represented as a Gaussian spot, and the size can be measured as a full width at half maximum (FWHM). Smaller spot size, indicated by a smaller FWHM, typically correlates with improved image of the spot. In some embodiments, the FWHM of the DNA nanoball spot can be about 10 um or less. The DNA nanoballs can be compact nucleic acid structures with smaller full width at half maximum (FWHM) compared to concatemers that are not collapsed / folded into DNA nanoballs.
[0376] In some embodiments, the compaction oligonucleotide comprises a single stranded oligonucleotide comprising DNA, RNA, or a combination of DNA and RNA. The compaction oligonucleotide can be any length including 20-150 nucleotides, or 30-100 nucleotides, or 40-80 nucleotides in length.
[0377] In some embodiments, the compaction oligonucleotide comprises a 5' region and a 3' region, and optionally an intermediate region between the 5' region and the 3' region. The intervening region can be any length, for example, 2 to 20 nucleotides. The intervening region comprises a homopolymer having consecutive identical bases (e.g., AAA, GGG, CCC, TTT, or UUU). The intervening region comprises a non-homopolymer sequence.
[0378] The 5' region of the compaction oligonucleotide may be fully or partially complementary along its length to a first portion of the concatemer molecule. The 3' region of the compaction oligonucleotide may be fully or partially complementary along its length to a second portion of the concatemer molecule. The 5' region of the compaction oligonucleotide may hybridize to a first universal sequence portion of the concatemer molecule. The 3' region of the compaction oligonucleotide may hybridize to a second universal sequence portion of the concatemer molecule. The 5' and 3' regions of the compaction oligonucleotide may hybridize to the concatemer to bring the distal portions of the concatemer together and cause compaction of the concatemer to form a DNA nanoball.
[0379] The 5' region of the compaction oligonucleotide can have the same sequence as the 3' region. The 5' region of the compaction oligonucleotide can have a sequence that is different from the 3' region. The 3' region of the compaction oligonucleotide can have a sequence that is the reverse sequence of the 5' region.
[0380] In some embodiments, sequence data can be derived via nanopore sequencing, which involves sequencing a nucleic acid by translocating the nucleic acid across a membrane, e.g., through a pore, and sequence reads or base calls are made by measuring one or more signals, such as impedance, current, voltage, or capacitance, during the translocation event. In some embodiments, the identity of a nucleotide can be determined by a unique electrical signature, such as the timing, duration, range, or linearity of a current block, impedance change, voltage change, or capacitance change. Sequencing a nucleic acid by translocation across a membrane and / or through a pore does not exclude alternative detection methods, such as optical, chemical, biochemical, fluorescent, luminescent, magnetic, electromagnetic, acoustic, or electroacoustic detection.
[0381] Supports and low non-specific coatings In some embodiments, the flow cell 112 of FIG. 1 can include a support, e.g., a solid support, as disclosed herein. The present disclosure provides pairwise sequencing compositions and methods that use a support having a plurality of oligonucleotide surface primers immobilized thereon. In some embodiments, the support is passivated with a low non-specific binding coating. The surface coatings described herein exhibit very low non-specific binding to reagents, e.g., dyes, nucleotides, enzymes, and nucleic acid primers, typically used for nucleic acid capture, amplification, and sequencing workflows. The surface coatings exhibit low background fluorescent signals or high contrast-to-noise (CNR) ratios compared to conventional surface coatings.
[0382] The low non-specific binding coating comprises one or more layers (FIG. 20). In some embodiments, multiple surface primers are immobilized in the low non-specific binding coating. In some embodiments, at least one surface primer is embedded in the low non-specific binding coating. The low non-specific binding coating allows for improved nucleic acid hybridization and amplification performance. Generally, the support comprises a substrate (or support structure), one or more layers of covalently or non-covalently attached low binding chemically modified layers, e.g., silane layers, polymer films, and one or more covalently or non-covalently attached surface primers that can be used to design single-stranded nucleic acid library molecules to the support. In some embodiments, the formulation of the coating, e.g., the chemical composition of one or more layers, the coupling chemistry used to crosslink one or more layers to the support and / or to each other, and the total number of layers, can be varied such that non-specific binding of proteins, nucleic acid molecules, and other hybridization and amplification reaction components to the coating is minimized or reduced relative to a comparable monolayer. The coating formulations described herein can be varied to minimize or reduce non-specific hybridization on the coating relative to a comparable monolayer. The coating formulations can be varied to minimize or reduce non-specific amplification on the coating relative to a comparable monolayer. The coating formulations can be varied to maximize specific amplification rate and / or yield on the coating. Amplification levels suitable for detection are achieved in some cases disclosed herein within 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30 or less amplification cycles, or more than 30.
[0383] The support structure comprising one or more chemically modified layers, e.g., layers of low non-specific binding polymers, may be freestanding or integrated within another structure or assembly. For example, in some embodiments, the support structure may comprise one or more surfaces within an integrated or assembled microfluidic flow cell. The support structure may comprise one or more surfaces within a microplate format, e.g., the bottom surface of a well in a microplate. In some embodiments, the support structure comprises an inner surface (e.g., lumen surface) of a capillary. In some embodiments, the support structure comprises an inner surface (e.g., lumen surface) of a capillary etched into a planar chip.
[0384] The attachment chemistry used to graft the first chemically modified layer onto the surface of the support generally depends on both the material from which the surface is made and the chemical nature of the layer. In some embodiments, the first layer may be covalently attached to the surface. In some embodiments, the first layer may be non-covalently attached, e.g., adsorbed, to the support via non-covalent interactions between the support and the molecular components of the first layer, e.g., electrostatic interactions, hydrogen bonds, or van der Waals interactions. In either case, the support may be treated prior to attachment or deposition of the first layer. Any of a variety of surface preparation techniques known to those skilled in the art may be used to clean or treat the surface. For example, glass or silicon surfaces may be acid-cleaned using piranha solution (a mixture of sulfuric acid (H2SO4) and hydrogen peroxide (H2O2)), base treatment in KOH and NaOH, and / or cleaned using oxygen plasma treatment methods.
[0385] Silane chemistry constitutes a non-limiting method for covalently modifying silanol groups on glass or silicon surfaces to attach more reactive functional groups (e.g., amine or carboxyl groups), which can then be used in coupling linker molecules (e.g., linear hydrocarbon molecules of various lengths, e.g., C6, C12, C18 hydrocarbons, or linear polyethylene glycol (PEG) molecules) or layer molecules (e.g., branched PEG molecules or other polymers) to the surface. Examples of suitable silanes that can be used in making any of the disclosed low-binding coatings include, but are not limited to, (3-aminopropyl)trimethoxysilane (APTMS), (3-aminopropyl)triethoxysilane (APTES), any of the various PEG silanes (e.g., those with molecular weights of 1K, 2K, 5K, 10K, 20K, etc.), amino-PEG silanes (i.e., those with free amino functional groups), maleimide-PEG silanes, and biotin-PEG silanes.
[0386] Any of a variety of molecules known to those of skill in the art, including but not limited to amino acids, peptides, nucleotides, oligonucleotides, other monomers, or polymers, or combinations thereof, may be used in creating one or more chemically modified layers on a support, and the selection of components used may be varied to modify one or more properties of the layer, such as the surface density of functional groups and / or tethered oligonucleotide primers, the hydrophilicity / hydrophobicity of the layer, or the three-dimensional nature (i.e., "thickness") of the layer. Examples of polymers that may be used to create one or more layers of low nonspecific binding materials in any of the disclosed coatings include, but are not limited to, polyethylene glycol (PEG) of various molecular weights and branched structures, streptavidin, polyacrylamide, polyester, dextran, poly-lysine, and poly-lysine copolymers, or any combination thereof. Examples of conjugation chemistries that may be used to graft one or more layers of material (e.g., polymer layers) to a surface and / or crosslink layers to one another include, but are not limited to, biotin-streptavidin interactions (or variations thereof), his-tag-Ni / NTA conjugation chemistry, methoxy ether conjugation chemistry, carboxylate conjugation chemistry, amine conjugation chemistry, NHS esters, maleimides, thiols, epoxies, azides, hydrazides, alkynes, isocyanates, and silanes.
[0387] The low non-specific binding surface coating can be applied uniformly across the entire support. Alternatively, the surface coating can be patterned so that the chemically modified layer is restricted to one or more distinct regions of the support. For example, the coating can be patterned using photolithography techniques to create an ordered array or random pattern of chemically modified regions on the support. Alternatively or in combination, the coating can be patterned using, for example, contact printing and / or inkjet printing techniques. In some embodiments, the ordered array or random pattern of chemically modified regions may comprise at least 1, 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, or 10,000 or more individual regions.
[0388] In some embodiments, the low non-specific binding coating comprises a hydrophilic polymer that is non-specifically adsorbed or covalently grafted to the support. Typically, passivation is performed using poly(ethylene glycol) (PEG, also known as polyethylene oxide (PEO) or polyoxyethylene) or other hydrophilic polymers with different molecular weights and end groups attached to the support, for example, using silane chemistry. End groups distal to the surface can include, but are not limited to, biotin, methoxy ether, carboxylate, amine, NHS ester, maleimide, and bis-silane. In some embodiments, two or more layers of hydrophilic polymers, for example linear, branched, or hyperbranched polymers, can be deposited on the surface. In some embodiments, the two or more layers can be covalently coupled to each other or internally crosslinked to improve the stability of the resulting coating. In some embodiments, surface primers with different nucleotide sequences and / or base modifications (or other biomolecules, for example enzymes or antibodies) can be tethered to the resulting layer at various surface densities. In some embodiments, for example, both the surface functional group density and the surface primer concentration can be varied to achieve a desired surface primer density range. In addition, the surface primer density can be controlled by diluting the surface primer with other molecules bearing the same functional group. For example, amine-labeled surface primers can be diluted with amine-labeled polyethylene glycol in reaction with the NHS-ester coated surface to reduce the final primer density. Also, surface primers with different lengths of linkers between the hybridization region and the surface-attached functional group can be applied to control the surface density. Examples of suitable linkers include poly-T and poly-A chains (e.g., 0-20 bases) at the 5' end of the primer, PEG linkers (e.g., 3-20 monomer units), and carbon chains (e.g., C6, C12, C18, etc.). To measure primer density, fluorescently labeled primers can be tethered to the surface and the fluorescent readings can then be compared to those for dye solutions of known concentration.
[0389] In some embodiments, the low non-specific binding coating comprises a functionalized polymer coating layer covalently bonded to at least a portion of the substrate via chemical groups on the substrate, a primer grafted to the functionalized polymer coating, and a water-soluble protective coating on the primer and the functionalized polymer coating. In some embodiments, the functionalized polymer coating comprises poly(N-(5-azidoacetamidylpentyl)acrylamide-co-acrylamide (PAZAM).
[0390] To adjust primer surface density and add additional dimensionality to hydrophilic or amphoteric coatings, supports containing multi-layer coatings of PEG and other hydrophilic polymers have been developed. By using hydrophilic and amphoteric surface layering techniques, including but not limited to the polymer / copolymer materials described below, it is possible to significantly increase the primer loading density on the support. Conventional PEG coating techniques use monolayer primer deposition, which has generally been reported for single molecule applications but does not result in high copy numbers for nucleic acid amplification applications. As described herein, "layering" can be accomplished with any compatible polymer or monomer subunit, such that a surface containing two or more highly crosslinked layers can be constructed in sequence using conventional crosslinking techniques. Examples of suitable polymers include, but are not limited to, streptavidin, polyacrylamide, polyester, dextran, poly-lysine, and copolymers of poly-lysine and PEG. In some embodiments, the different layers may be attached to one another via any of a variety of conjugation reactions, including, but not limited to, biotin-streptavidin binding, azide-alkyne click reactions, amine-NHS ester reactions, thiol-maleimide reactions, and ionic interactions between positively and negatively charged polymers. In some embodiments, the high primer density material may be constructed in solution and then layered onto a surface in multiple steps.
[0391] Examples of materials from which the support structure may be fabricated include, but are not limited to, glass, fused silica, silicon, polymers (e.g., polystyrene (PS), macroporous polystyrene (MPPS), polymethyl methacrylate (PMMA), polycarbonate (PC), polypropylene (PP), polyethylene (PE), high density polyethylene (HDPE), cyclic olefin polymer (COP), cyclic olefin copolymer (COC), polyethylene terephthalate (PET)), or any combination thereof. Various compositions of both glass and plastic support structures are contemplated.
[0392] The support structure may be in any of a variety of geometries and dimensions know...
Claims
1. 1. A method comprising: providing a first plurality of library molecules immobilized on a support, each of the first plurality of library molecules comprising a first insert sequence and a first sample index sequence derived from a first sample source, the first sample index sequence comprising a first k-mer sequence and a first universal sample index sequence, and the first universal sample index sequence identifying the first sample source of the first insert sequence; providing a second plurality of library molecules immobilized on the support, each of the second plurality of library molecules comprising a second insert sequence and a second sample index sequence derived from a second sample source, the second sample index sequence comprising a second k-mer sequence and a second universal sample index sequence, and a second universal sample index identifying the second sample source of the second insert sequence; performing k sequencing reaction cycles of the first and second k-mer sequences with a sequencing system, thereby generating a first plurality of flow cell images; determining, by a processor, pixel intensities and respective color purity of each of said pixel intensities for pixels of said first plurality of flow cell images; determining, by the processor and prior to performing one or more sequencing reaction cycles of the first or second insert sequence, a base calling template including base calling positions based on the pixel intensities and the respective color purities of the pixel intensities; determining that the base calling template is configured to align a second plurality of flow cell images of the support at one or more cycles subsequent to the k cycles.
2. pooling the first and second plurality of library molecules; distributing the pooled library molecules on the support and performing an amplification reaction to generate a plurality of nucleic acid template molecules immobilized on the support; 10. The method of claim 1, wherein the plurality of nucleic acid template molecules are clonally amplified from the first library molecule and the second library molecule.
3. determining, by the processor, for the pixels of the first plurality of flow cell images, the pixel intensities and the respective color purity of each of the pixel intensities; determining, by the processor, a base-calling template including base-calling positions based on the pixel intensities and the respective color purity of the pixel intensities; the first sample index array; the first universal sample index sequence; the second sample index array; and 2. The method of claim 1, prior to performing any cycles of sequencing reactions with the second universal sample index sequence.
4. 2. The method of claim 1, wherein performing the k sequencing reaction cycles of the k-mer sequences and of the base positions of the first universal sample index sequence is based on a sequencing order of a sequencing run.
5. 5. The method of claim 4, wherein the order of sequencing comprises sequencing the k-mer sequence, then sequencing the first and second universal sample index sequences, and then sequencing the first and second insert sequences.
6. The method of claim 1 , wherein the first or second plurality of flow cell images are from two, three, or four different color channels.
7. 2. The method of claim 1, wherein the first plurality of flow cell images from the k cycles comprises a balanced diversity of A, G, C, and T / U nucleotide bases among the plurality of nucleic acid template molecules immobilized on the support in each of the k cycles.
8. 2. The method of claim 1, wherein the k-mer sequence comprises a random sequence of at least two or three nucleotide bases of A, G, C, and T / U.
9. The method of claim 1 , wherein the support is comprised in a flow cell device.
10. The density of the nucleic acid template molecules on the support is 1 mm 2 10 per 4 ~10 12 The method of claim 2, wherein
11. cycling the sequencing reaction of the k-mer sequence k times; contacting a polony of nucleotide acid template molecules with a plurality of sequencing primers, a plurality of polymerases, and a mixture of different types of avidites; 10. The method of claim 1, wherein each of the plurality of nucleic acid template molecules immobilized on the support corresponds to a polony.
12. cycling the sequencing reaction k times for the k-mer sequence; 2. The method of claim 1, further comprising acquiring, by an optical system, the first plurality of flow cell images comprising optical color signals emitted from nucleotide reagents bound to template molecules during each of the k cycles.
13. The method of claim 1 , wherein k is an integer greater than 0 and less than 10.
14. 3. The method of claim 2, wherein each of the base calling positions corresponds to a position of the plurality of immobilized template molecules.
15. 2. The method of claim 1, wherein the second plurality of flow cell images comprises light signals emitted from nucleotide reagents bound to A, G, C, and T / U nucleotide bases of unbalanced diversity among the plurality of nucleic acid template molecules immobilized on the support in the one or more cycles following the k cycles.
16. 16. The method of claim 15, wherein the unbalanced diversity of A, G, C, and T / U nucleotide bases among the plurality of nucleic acid template molecules comprises a percentage of the number of at least one type of nucleotide base (1) relative to the total number of bases (2) in the one or more cycles that is less than 20%, 15%, 10%, or 5%.
17. registering, by the processor, the second plurality of flow cell images from one or more subsequent flow cycles to the base-calling template; and 2. The method of claim 1, further comprising: performing, by the processor, base calling of the second plurality of flow cell images at the base calling locations within the base calling template using signals from the aligned second plurality of flow cell images.
18. registering the second plurality of flow cell images from the one or more subsequent flow cycles to the base-calling template; 18. The method of claim 17, comprising generating coordinates of polonies in the second plurality of flow cell images in a common coordinate system as the base calling template.
19. 1. A system comprising: one or more hardware processors; one or more data storage devices storing instructions executable by the one or more hardware processors, the instructions, when executed, causing the one or more hardware processors to perform operations, the operations including: providing a first plurality of library molecules immobilized on a support, each of the first plurality of library molecules comprising a first insert sequence and a first sample index sequence derived from a first sample source, the first sample index sequence comprising a first k-mer sequence and a first universal sample index sequence, and the first universal sample index sequence identifying the first sample source of the first insert sequence; providing a second plurality of library molecules immobilized on the support, each of the second plurality of library molecules comprising a second insert sequence and a second sample index sequence derived from a second sample source, the second sample index sequence comprising a second k-mer sequence and a second universal sample index sequence, and a second universal sample index identifying the second sample source of the second insert sequence; performing k sequencing reaction cycles of the first and second k-mer sequences with a sequencing system, thereby generating a first plurality of flow cell images; determining, by a processor, pixel intensities and respective color purity of each of said pixel intensities for pixels of said first plurality of flow cell images; determining, by the processor and prior to performing one or more sequencing reaction cycles of the first or second insert sequence, a base calling template including base calling positions based on the pixel intensities and the respective color purities of the pixel intensities; determining that the base calling template is configured to align a second plurality of flow cell images of the support at one or more cycles subsequent to the k cycles.
20. one or more non-transitory computer storage media encoded with instructions executable by one or more hardware processors, the instructions, when executed, causing the one or more hardware processors to perform operations in sequencing data analysis, the operations including: providing a first plurality of library molecules immobilized on a support, each of the first plurality of library molecules comprising a first insert sequence and a first sample index sequence derived from a first sample source, the first sample index sequence comprising a first k-mer sequence and a first universal sample index sequence, and the first universal sample index sequence identifying the first sample source of the first insert sequence; providing a second plurality of library molecules immobilized on the support, each of the second plurality of library molecules comprising a second insert sequence and a second sample index sequence derived from a second sample source, the second sample index sequence comprising a second k-mer sequence and a second universal sample index sequence, and a second universal sample index identifying the second sample source of the second insert sequence; performing k sequencing reaction cycles of the first and second k-mer sequences with a sequencing system, thereby generating a first plurality of flow cell images; determining, by a processor, pixel intensities and respective color purity of each of said pixel intensities for pixels of said first plurality of flow cell images; determining, by the processor and prior to performing one or more sequencing reaction cycles of the first or second insert sequence, a base calling template including base calling positions based on the pixel intensities and the respective color purities of the pixel intensities; determining that the base calling template is configured to align a second plurality of flow cell images of the support at one or more cycles subsequent to the k cycles.